Construction method of a system for automatically correcting the retention time of a group of ions to be measured in liquid mass spectrometry based on a labeling system

By adding spiked exogenous ion pairs to the liquid phase-tandem mass spectrometry detection method, an automatic retention time correction system was constructed, which solved the problem of retention time shift of the measured ion pair group, improved the integration efficiency and accuracy, and could early warning of abnormal phenomena.

CN115963188BActive Publication Date: 2025-05-30PRECOGIFY PHARM CHINA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111175592.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-09
Publication Date
2025-05-30
Estimated Expiration
2041-10-09

AI Technical Summary

Technical Problem

In the liquid phase-tandem mass spectrometry detection method, the retention time of the pair of ions to be measured is easily offset, resulting in manual correction when integrating, which is inefficient.

Method used

By adding spiked exogenous ion pairs to the sample, using liquid phase-tandem mass spectrometry technology, the correlation between spiked working liquid ion pairs and the measured ion pairs was compared, and a retention time automatic correction system was constructed.

Benefits of technology

Automatic correction of the retention time of the measured ions pair group is achieved, the efficiency and accuracy of the integration are improved, and abnormal phenomena in the experiment can be detected early.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115963188B_ABST
    Figure CN115963188B_ABST
Patent Text Reader

Abstract

The present invention discloses a construction method for an automatic retention time correction system for a liquid chromatography tandem mass spectrometry detected ion pair group based on a spiking system, and establishes a high performance liquid chromatography tandem mass spectrometry detection method for simultaneously detecting the detected ion pair group and the spiking working solution in human serum; obtains the retention times of the detected ion pair group and the ion pairs of the spiking working solution in the serum; constructs a prediction model using a training set, and tests and verifies the reliability of the prediction model using a test set; and then verifies it on a validation set. The present invention is independent of the dependence on reference standards, and realizes automatic correction of the retention time by spiking exogenous ion pairs; provides an effective, reliable and convenient correction method for the retention times of the detected ion pair groups in the same batch of experiments, improves the efficiency and accuracy of integration; through the deviation analysis of the actual retention time and the corrected retention time of the detected ion pair group, effectively warns of abnormalities in the experiment, and can early detect various abnormal phenomena in the experiment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of liquid chromatography - mass spectrometry, and particularly relates to a method for constructing a system for automatically correcting the retention time of a group of ions to be measured in liquid chromatography - mass spectrometry based on a spiking system. Background Art

[0002] The liquid chromatography - tandem mass spectrometry detection method combines liquid chromatography separation and mass spectrometry ion pair detection, which has the characteristics of high sensitivity and strong specificity. Metabolomics is one of the most active sub - disciplines in the field of systems biology research in recent years. The statistical analysis technology of mass spectrometry data using metabolomics is widely applied in various research fields. In metabolomics analysis, the accuracy of the peak area as basic data is crucial, and the integration accuracy of the peak area depends on whether the retention time of the group of ions to be measured in liquid chromatography - mass spectrometry is correctly set.

[0003] When the liquid chromatography - tandem mass spectrometry detection method combines liquid chromatography separation and mass spectrometry ion pair detection, during the use of the instrument, the retention time of the group of ions to be measured may shift. Some are partial shifts of the group of ions to be measured, and some are shifts of the entire batch of the group of ions to be measured, which causes the retention time not to peak at the expected time point. When integrating, it is necessary to manually correct the retention time of the group of ions to be measured one by one, resulting in low efficiency.

[0004] Endogenous substances may show peaks in some individual samples and not in others according to different individuals, while exogenous substances are relatively stable in expression among the population, and the retention time of their peaks is also relatively stable.

[0005] After adding exogenous substances to individual samples by spiking, if there is a deviation in the retention time of the spiked ion pair within the batch, the retention time of the group of ions to be measured should also deviate. Therefore, according to the deviation in the retention time of the spiked ion pair, a large amount of data is used to establish a model to find the relationship between them, so as to realize the automatic correction of the retention time of the group of ions to be measured.

[0006] This study uses liquid chromatography - mass spectrometry technology to find the variation law by comparing the correlation between the ion pairs of the spiking working solution and the group of ions to be measured, and constructs an effective system for automatically correcting the retention time.

[0007] No existing patent closest to the technology of the present invention was found. However, the commonly used retention time correction tool in the industry, MacCossLab Software (https: / / skyline.ms / project / home / software / Skyline / begin.view), is somewhat similar to the present invention. Its calculation method is to directly use the retention times of several detected endogenous metabolites specified manually to correct the retention times of the measured ion pair groups. However, since endogenous metabolites may show peaks in some individual samples and not in others according to different individuals, whether appropriate metabolites can be selected for different individuals determines whether the retention time can be corrected correctly. And because the individual samples and the metabolites collected in different experiments are all different, it is very difficult to select appropriate metabolites for correction.

[0008] In addition, the article "Development of a plasma pseudotargeted metabolomics method based on ultra-high-performance liquid chromatography-mass spectrometry" DOI: 10.1038

[0009] / s41596-020-0341-5 also mentions the correction of retention time, but it is used to correct the retention time in the second-order mass spectrometry data based on the first-order mass spectrometry data. It requires both first-order and second-order mass spectrometry data, with high experimental costs and different uses from the present invention. Summary of the Invention

[0010] To overcome the defects of the prior art, the present invention provides a method for constructing a system for automatically correcting the retention time of a measured ion pair group in liquid mass spectrometry based on a spiking system, which is independent of the dependence on standard substances and realizes the automatic correction of the retention time by spiking exogenous ion pairs; it provides an effective, reliable and convenient correction method for the retention times of the measured ion pair groups in the same batch of experiments, improving the efficiency and accuracy of integration; through the deviation analysis of the actual retention time and the corrected retention time of the measured ion pair group, it effectively warns of abnormalities in the experiment and can detect various abnormal phenomena in the experiment at an early stage.

[0011] The present invention is implemented by adopting the following technical solutions:

[0012] First, establish a liquid chromatography-tandem mass spectrometry (LC-MS / MS) detection method for simultaneously detecting the analyte ion pairs and spiked working solution ion pairs in human serum; second, obtain the retention times of the analyte ion pairs and spiked working solution ion pairs in the serum; use a part of the downloaded data as training data and another part as test data. Train the model with the training data and then use the test data to check the prediction performance. That is, judge the quality of the model through the test data, then continuously modify the model, perform cross-validation, and then verify it on the validation set.

[0013] As a further description of the present invention: A method for constructing a system for automatically correcting the retention time of analyte ion pairs in liquid mass spectrometry based on a spiking system, comprising the following steps:

[0014] Establish a high performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS) system detection method for simultaneously detecting the analyte ion pairs and spiked working solution in human serum;

[0015] Obtain the retention times of the analyte ion pairs and spiked working solution ion pairs in the serum, and establish a database with the detection sample data;

[0016] Preprocess the detection sample data, and divide the detection sample data in the database into a training set and a test set according to a ratio; select a machine learning algorithm to construct a retention time prediction model for the analyte ion pairs using the training set, use the test set to test and verify the reliability of the prediction model and evaluate it with evaluation indicators;

[0017] Perform N-fold cross-validation on the prediction model, then verify it on the validation set, calculate the prediction accuracy evaluation indicators, and evaluate the prediction results.

[0018] As a further description of the present invention: Preferably, the high performance liquid chromatography-tandem mass spectrometry (HPLC-MS / MS) system detection method for simultaneously detecting the analyte ion pairs and spiked working solution comprises the following steps: First, chromatographically separate the ion pairs in the serum using a liquid phase system; second, use a tandem mass spectrometry system to configure the spiked substance as a mixed working solution, take an appropriate amount of the mixed working solution and add it to the sample to be detected, and perform pretreatment with the sample to be detected; after pretreatment, inject the sample into the liquid chromatography-tandem mass spectrometry system for detection, and determine the retention times of each standard substance based on the chromatographic peaks generated by the characteristic MRM ion pairs of each standard substance.

[0019] As a further illustration of the present invention: Preferably, the pretreatment includes the following steps: After the serum sample to be tested is taken out of the ultra-low temperature refrigerator, it is thawed in an ice bath. After the sample is completely thawed, it is vortexed for 60 seconds. Then, 80 μL of the serum sample is accurately measured with the lid open and placed in a 1.5 mL centrifuge tube; 5 μL of the spiked working solution is added to the 80 μL serum sample and vortexed for 60 seconds; 150 μL of the sample precipitant is added and vortexed for 120 seconds; 50 μL of the ammonium formate buffer solution is added and vortexed for 120 seconds; it is centrifuged at 13000 rpm for 5 minutes at room temperature; 60 μL of the supernatant is accurately measured into an EP tube containing 150 μL of water, mixed evenly and then loaded into the sample vial, waiting for sample injection and detection; the sample precipitant is acetonitrile: isopropanol with a volume ratio of 8:2; the ammonium formate buffer solution is prepared by mixing ammonium formate: water at 0.6 g: 1 ml.

[0020] As a further illustration of the present invention: Preferably, the preparation method of the spiked working solution: Accurately measure 10 μL of the spiked intermediate working solution of the mixed drug reference substance, add it to 990 μL of the methanol solution, and use it after vortexing for 60 seconds; the spiked working solution needs to be prepared and used immediately and cannot be stored;

[0021] The preparation method of the spiked intermediate working solution of the mixed drug reference substance: Take out the stock solutions of each drug reference substance, thaw them to room temperature and vortex for 60 seconds; accurately measure the stock solutions into another separate container, add methanol respectively and vortex for 60 seconds; accurately measure the stock solutions or diluted stock solutions into the same bottle; add methanol to the vial of the mixed stock solution and vortex for 60 seconds; accurately measure the stock solutions of each drug reference substance into the same brown glass bottle respectively;

[0022] The preparation of the stock solutions of each drug reference substance: Weigh an appropriate amount of each drug reference substance accurately and place it in a brown glass bottle, add an appropriate amount of methanol to dissolve the drug reference substance, ensuring that the mass concentration of each drug reference substance stock solution is 1.0 mg / mL; the prepared drug reference substance stock solutions are stored in an ultra-low temperature refrigerator (-80 °C) with a validity period of 6 months; when using, it needs to be thawed to room temperature at room temperature and vortexed for 60 seconds before opening the lid for use.

[0023] As a further illustration of the present invention: Preferably, the detection and analysis conditions of the high performance liquid chromatography-tandem mass spectrometry system are as follows:

[0024] The high performance liquid chromatography conditions are: injection volume: 15 μL; column temperature: 35 °C; sample chamber temperature: 15 °C; flow rate: 0.4 ml / min; gradient elution: mobile phase A: aqueous solution of 0.1% formic acid, mobile phase B: acetonitrile solution of 0.1% formic acid

[0025] The mobile phase gradient elution program is:

[0026]

[0027] As a further illustration of the present invention: Preferably, the training process of the prediction model for the retention time of the ion pair group to be measured based on the labeled working fluid ion pair is specifically as follows: The detection sample data is divided into a training set and a test set according to a ratio of 4:1, and a model is established through decision tree regression of the training set samples; Finally, by verifying the effect of the model on the test set, the final model is determined, the retention time of the ion pair group to be measured on the verification set is predicted through the model, and the evaluation index of the corresponding model prediction effect is calculated.

[0028] As a further illustration of the present invention: Preferably, the preprocessing is to clean and sort the data, correct missing values, normalize / standardize the values to make them comparable, and transform the data.

[0029] As a further illustration of the present invention: Preferably, the mean absolute error MAE is used as the evaluation index for the model prediction effect, and the calculation formula of the mean absolute error MAE is:

[0030]

[0031] In the formula, N is the total number of samples, y ai and y pi represent the true value and the predicted value respectively.

[0032] As a further illustration of the present invention: Preferably, the ten-fold cross-validation method is used to verify the regression decision tree model, and the verification steps include: randomly dividing the data set into 10 parts, taking 9 of them as training data and 1 as verification data in turn to verify the model. The corresponding mean absolute error value will be obtained each time of verification, and the average value of the 10 mean absolute error values is used as the output of the model. The model is verified and the model parameters are adjusted through this output.

[0033] As a further illustration of the present invention: Preferably, the prediction accuracy evaluation indexes include the threshold of the ion pair, the sample threshold, and the batch MAE value;

[0034] Determination of the threshold of the ion pair: Through model prediction, the predicted retention time and the calculated MAE are obtained, and the threshold of the ion pair is determined to be 0.02. If it is less than 0.02, it is considered an ion pair that is relatively easy to predict;

[0035] Determination of the sample threshold: When the threshold is 80% of the number of samples, the sample threshold MAE should be less than 0.03. When the threshold of the sample exceeds 0.03, the ion pairs of this sample should be focused on;

[0036] Determination of the MAE value of each batch: According to the predicted sample results, calculate the average MAE value (mean_MAE) of all samples in each batch, and it can be known whether there is an offset in a certain batch and the degree of the offset. Make a judgment based on the mean_MAE value of the batch. The threshold is set to 0.03. If it is greater than this threshold, a large retention time offset has occurred. When a large retention time offset occurs, corresponding processing is carried out in the experiment and calculation.

[0037] The present invention provides a system and a construction method for automatically correcting the retention time of a liquid chromatography-mass spectrometry detected ion pair group based on a spiking system. Compared with the prior art, the present invention has the following beneficial technical effects:

[0038] 1) Traditionally, a standard substance is used to correct the retention time, while the present invention is independent of the dependence on the standard substance and realizes automatic correction of the retention time by spiking exogenous ion pairs.

[0039] 2) While the automatic correction system uses liquid chromatography-tandem mass spectrometry to detect the content of metabolites in human serum, it also detects the content of the spiking working solution. The pretreatment process of the spiking working solution is simple and the analysis time is short, which is suitable for high-throughput analysis and testing.

[0040] 3) The retention time automatic correction system constructed by the present invention provides an effective, reliable and convenient correction method for the retention time of the detected ion pair group in the same batch of experiments, improving the efficiency and accuracy of integration.

[0041] 4) The retention time correction system constructed by the present invention can effectively give an early warning of abnormalities in the experiment by analyzing the deviation between the actual retention time and the corrected retention time of the detected ion pair group, and can detect various abnormal phenomena in the experiment at an early stage. Description of the Drawings

[0042] Figure 1 It is a schematic flow chart of a construction method for a system for automatically correcting the retention time of a liquid chromatography-mass spectrometry detected ion pair group based on a spiking system.

[0043] Figure 2 It is a statistical result graph of the retention time prediction of an ion pair in the test set. Among them, the X-axis is the absolute value of the difference between the true retention time and the predicted retention time, the left Y-axis is the statistical number that meets the conditions, and the right Y-axis is the percentage of the statistical number.

[0044] Figure 3 It is a verification result graph of an ion pair in the validation set. Among them, the X-axis is the true retention time and the Y-axis is the predicted retention time.

[0045] Figure 4Scatter plot of negative ion pairs of the two samples with the highest MAE in the test results based on the MAE evaluation index, where x is the true retention time of the ion pair and y is the predicted retention time of the ion pair.

[0046] Figure 5 Scatter plot of negative ion pairs of the two samples with the lowest MAE, where x is the true retention time of the ion pair and y is the predicted retention time of the ion pair.

[0047] Figure 6 Graphs of the two samples with the highest test results based on the MAE evaluation index.

[0048] Figure 7 Graphs of the two samples with the lowest test results based on the MAE evaluation index.

[0049] Figure 8 Performance of a predicted ion pair in the test set and validation set.

[0050] Figure 9 Statistical graph of the proportion of MAE values of the samples in the validation set. The X-axis is the value of MAE, the left Y-axis is the number of statistics that meet the conditions, and the right Y-axis is the percentage of the number of statistics. Detailed implementation method

[0051] In order to more clearly understand the above-mentioned objects, features and advantages of the present invention, the present invention will be described in detail below with reference to specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0052] Embodiment 1:

[0053] Establishment of a liquid chromatography-tandem mass spectrometry detection method for simultaneously detecting the ion pair group to be detected in human serum and 11 kinds of spiked working solutions

[0054] 1 Purpose

[0055] By spiking samples, establish a liquid chromatography-tandem mass spectrometry detection method for simultaneously detecting the ion pair group to be detected in human serum and spiked working solutions

[0056] 2 Experimental instruments and materials

[0057] 2.1 Instruments

[0058] SCIEXTripleQuad4500 tandem mass spectrometry system; Analyst Software data processing workstation, both are products of AB Sciex Company, USA. Shimadzu Prominence UFLC high performance liquid chromatography system, including binary pump, vacuum degasser, auto sampler, column oven, all are products of Shimadzu (SHIMADZU) Company, Japan. Electronic balance, Mettler Toledo AB104 type (maximum load 101 g, division value 0.1 mg), product of Mettler Company, Switzerland. Pipette, single-channel adjustable range 20 μL, 100 μL, 200 μL, 1000 μL, product of Eppendorf (Shanghai) International Trading Co., Ltd. Vortex mixer, VORTEX-GENIE2 type, product of Scientific Industries Company, USA. High-speed centrifuge, Centrifuge 5415R type, product of Eppendorf Company, Germany. Refrigerator-freezer, model BCD-182NE, product of Aucma Co., Ltd. Ultra-low temperature freezer, DW-86L338J type, product of Qingdao Haier Biomedical Co., Ltd.

[0059] 2.2 Reagents and Consumables

[0060] Methanol is LC-MS Grade, acetonitrile is LC-MS Grade, formic acid is LC-MS Grade, isopropanol is LC-MS Grade, water is HPLC grade, all are produced by Thermo Fisher Scientific Company, USA. Ammonium acetate is of analytical grade, produced by Sinopharm Chemical Reagent Co., Ltd.

[0061] Disposable centrifuge tubes, 1.5 mL, 2 mL, all are produced by Axygen Biotechnology Co., Ltd. Disposable pipette tips, 10 μL, 200 μL, 1000 μL, produced by Axygen Biotechnology Co., Ltd. Disposable sample vials, 300 μL, produced by Thermo Fisher Scientific Company, USA.

[0062] Shimadzu Shim-pack Velox C 18 2.7 μm 2.1×100 mm chromatographic column, product of Shimadzu Company.

[0063] The spiked drug reference substances randomly added in this experiment (see Table 1) are all purchased from the National Drug Reference Substances Inquiry and Ordering Platform of the National Institutes for Food and Drug Control.

[0064] Table 1: Table of spiked drug reference substances randomly added in this experiment

[0065] Variety Number Chinese Name Batch Number Specification Standard Classification 101008 Ropivacaine Mesylate 101008-201302 20mg Chemical Reference Substance 101276 Pirfenidone 101276-201301 100mg / vial Chemical Reference Substance 420018 Capecitabine 420018-201802 100mg Chemical Reference Substance 100460 Chlorpromazine Hydrochloride 100460-201302 100mg Chemical Reference Substance 510011 Deflazacort 510011-201301 50mg / vial Chemical Reference Substance 130496 Rifampicin 130496-201403 100mg / vial Chemical Reference Substance 100364 Chlorzoxazone 100364-201302 100mg Chemical Reference Substance 100257 Indapamide 100257-201605 100mg / vial Chemical Reference Substance 100555 Nimesulide 100555-201803 100mg / vial Chemical Reference Substance 100190 Mefenamic Acid 100190-201104 100mg Chemical Reference Substance 100677 Benzbromarone 100677-201802 100mg / vial Chemical Reference Substance

[0066] 3 Liquid chromatography-tandem mass spectrometry detection method

[0067] The liquid chromatography-tandem mass spectrometry detection method includes the following steps: Configure the above drug reference substances into a mixed working solution, take an appropriate amount of the working solution and add it to the sample to be detected, and perform pretreatment along with the sample to be detected. After pretreatment, the sample is injected into the liquid chromatography-tandem mass spectrometry system for detection. According to the chromatographic peaks generated by the characteristic MRM ion pairs of each reference substance, determine the retention time of each reference substance.

[0068] 3.1 Preparation and pretreatment of related solutions

[0069] 3.1.1 Preparation of stock solutions of each drug reference substance

[0070] Accurately weigh appropriate amounts (about 1 mg) of each drug reference substance and place them in suitable brown glass bottles. According to the accurately weighed mass, accurately add appropriate amounts of methanol to dissolve the drug reference substances, ensuring that the mass concentration of each drug reference substance stock solution is 1.0 mg / mL. The prepared drug reference substance stock solutions are stored in an ultra-low temperature refrigerator (-80 °C) with a validity period of 6 months. When in use, thaw to room temperature at room temperature and vortex for 60 seconds before opening the lid for use.

[0071] 3.1.2 Preparation of spike-in intermediate working solution of mixed drug reference substances.

[0072] Take out the stock solutions of each drug reference substance, thaw to room temperature and vortex for 60 seconds, and then accurately measure the stock solutions of each drug reference substance into the same brown glass bottle according to the steps in Table 2.

[0073] Table 2: Configuration table of spike-in intermediate working solution of mixed drug reference substances

[0074]

[0075] The spike-in intermediate working solution is stored frozen in the refrigerator (-20 °C) with a validity period of 60 days. When in use, thaw to room temperature at room temperature and vortex for 60 seconds before opening the lid for use.

[0076] 3.1.3 Preparation of spike-in working solution

[0077] Accurately measure 10 μL of the spike-in intermediate working solution and add it to 990 μL of methanol solution, and use it after vortexing for 60 seconds. The spike-in working solution needs to be prepared and used immediately and cannot be stored.

[0078] 3.1.4 Preparation of sample precipitant

[0079] Acetonitrile: Isopropanol (8:2 by volume)

[0080] 3.1.5 Preparation of ammonium formate buffer

[0081] Ammonium formate: Water (0.6 g: 1 ml)

[0082] 3.1.6 Preparation of mobile phase and needle washing solution for liquid chromatography

[0083] Mobile phase A: Aqueous solution of 0.1% formic acid

[0084] Mobile phase B: Acetonitrile solution of 0.1% formic acid

[0085] Needle washing solution: 0.1% methanol: Water (1:1)

[0086] 3.2 Sample pretreatment

[0087] 3.2.1 After the tested serum sample is taken out of the ultra-low temperature refrigerator, it is thawed in an ice bath. After the sample is completely thawed, it is vortexed for 60 seconds. Open the lid and accurately measure 80 μL of the serum sample and place it in a 1.5 mL centrifuge tube;

[0088] 3.2.2 Add 5 μL of the spike-in working solution to 80 μL of the serum sample and vortex for 60 seconds;

[0089] 3.2.3 Add 150 μL of the sample precipitant and vortex for 120 seconds;

[0090] 3.2.4 Add 50 μL of the ammonium formate buffer and vortex for 120 seconds;

[0091] 3.2.5 Centrifuge at 13000 rpm for 5 minutes at room temperature;

[0092] 3.2.6 Accurately measure 60 μL of the supernatant into an EP tube containing 150 μL of water, mix well and transfer it to an injection vial, waiting for injection and detection

[0093] 3.3 Analytical conditions of high performance liquid chromatography-tandem mass spectrometry system

[0094] 3.3.1 High performance liquid chromatography conditions

[0095] Injection volume: 15 μL Column temperature: 35 °C

[0096] Needle washing mode: Before and after aspiration, 200 μL, 2 Sec

[0097] Sample chamber temperature: 15 °C Flow rate: 0.4 ml / min

[0098] The mobile phase gradient elution program is shown in the following table:

[0099] Table 3: Mobile Phase Gradient Elution Program Table

[0100]

[0101]

[0102] 3.3.2 Tandem Mass Spectrometry Analysis Conditions

[0103] The ion source parameters are shown in the following table:

[0104] Table 4: Ion Source Parameter Table

[0105]

[0106] MRM Parameter Table:

[0107] Table 5: MRM Parameter Table

[0108]

[0109]

[0110] 3.4 Mass Spectrometry Data Analysis

[0111] The Analyst Software data processing workstation is used for mass spectrometry data processing, and the detection results are presented in the form of csv for the next step of data analysis.

[0112] 4 Results

[0113] For the concentration detection deviation of the spiked working solution, the inter-day precision RSD is "chlorzoxazone" = 5.29%, "indapamide" = 4.67%, "nimesulide" = 5.69%, "mefenamic acid" = 3.71%, "benzbromarone" = 4.51%, "R-ropivacaine mesylate" = 9.19%, "pirfenidone" = 4.65%, "capecitabine" = 5.77%, "deflazacort" = 7.51%, "chlorpromazine hydrochloride" = 10.08%, "rifampicin" = 13.54%, indicating the stability of the concentration of the spiked working solution in serum. It can be seen that during the detection process of the entire batch of samples, the performance of the chromatographic column and the mass spectrometer is relatively stable, and the quality of the experimental data this time is reliable.

[0114] 5 Summary

[0115] This experiment conducted targeted quantitative analysis on the spiked working solution randomly added to the serum sample and the ion pair group to be measured, and established a liquid chromatography-tandem mass spectrometry detection method for simultaneously detecting the ion pair group to be measured and the spiked working solution in human serum, both of which had good chromatographic separation and mass spectrometry response signals. In terms of day-to-day precision, during the detection process of the entire batch of samples, the performance of the chromatographic column and the mass spectrometer was relatively stable.

[0116] Example 2:

[0117] Study on the retention time of the ion pair of the spiked working solution

[0118] 1 Purpose

[0119] Based on the retention time of the ion pair of the spiked working solution, predict the retention time of the randomly measured ion pair group.

[0120] 2 Data processing and statistical methods

[0121] Write a program in Python language for multi-dimensional data processing and model training.

[0122] 3 Preprocessing of sample data

[0123] First, the data was cleaned, sorted, missing values were corrected, the values were normalized / standardized to make them comparable, and the data was transformed, etc. For example: the 0 in the output set was processed, and the value of the previous row was used to fill this 0 value. The quality of the data will have a great impact on the quality of the generated model. During the development process of a machine learning model, it is hoped that the trained model can perform well on new, unseen data. In order to simulate new, unseen data, the available data was divided, and thus it was divided into 2 parts (training set and test set). In particular, the first part is a larger data subset, used as the training set (such as accounting for 80% of the original data), and the second part is usually a smaller subset, used as the test set (the remaining 20% of the data).

[0124] Therefore, the training set is 80%, and the test set is 20%. The purpose of using the training set and the test set is to make model selection, and the validation set is for model evaluation. Different selections are made according to the sample size, but one principle is that the validation set needs to remain unknown and independent of the training set and the test set. Model selection and model evaluation both rely on the validation set for adjustment. Therefore, from the perspective of the independence of model evaluation, the best approach is to replace several groups of unknown data sets as the "validation set" for model evaluation. Here, 23 batches of data are used as the training set (80%) and the test set (20%), and 26-34 batches of data are used as the validation set.

[0125] According to the data type (qualitative or quantitative) of the target variable (usually referred to as the Y variable), a classification (if Y is qualitative) or regression (if Y is quantitative) model needs to be established. Here, a regression model is established to make accurate retention time predictions for each ion pair information. There are many types of regression models, including linear regression, polynomial regression, ridge regression, Lasso regression, decision tree regression, random forest, and SVR, etc. After comparing several methods, the decision tree regression method is finally selected. It includes three steps: feature selection, decision tree generation, and decision tree pruning.

[0126] Model comparison and selection:

[0127] Observing the characteristics of the data, the retention times of each ion pair basically conform to the normal distribution, and the probability of the predicted data also conforms to the normal distribution. Judging from the calculated coefficient of determination (R2), negative values are more likely to appear, and the data belongs to R2_score < 0. Therefore, this kind of data is not very suitable for algorithms such as linear regression. Linear regression, ridge regression, lasso regression, and lowess locally weighted linear regression are excluded. After comparing the random forest and decision tree regression, the random forest is inferior to the decision tree regression in the 10-fold cross-validation of the test set and validation set. The random forest model is also more complex and the training time is longer. Therefore, the decision tree regression is finally selected as the model for training.

[0128] 4 Summary

[0129] The study on the retention times of the spiked working fluid ion pairs in Example 2 shows that the preprocessing of the data in the early stage is very important. If it is not handled well, it may directly affect the final effect and evaluation indicators of the model. Secondly, the selection of the model is also a key step. By comparing the differences and effects of different models, the establishment of the prediction model can be carried out in the next step.

[0130] Example 3:

[0131] Establishment of a prediction model based on the retention times of the spiked working fluid ion pairs

[0132] 1 Purpose

[0133] According to the retention times of the spiked working fluid ion pairs, establish a prediction model to predict the retention times of all its ion pairs and conduct model verification.

[0134] 2 Data processing and statistical methods

[0135] Use the Python language for programming analysis, and calculate the MAE values in three dimensions of ion pairs, samples, and batches through the prediction results of the model.

[0136] 3 Establishment of the retention time prediction model and model verification

[0137] The evaluation index is MAE. It is required that the distance between the true value and the prediction result is the smallest. We can directly subtract them, take the absolute value, add them m times and then divide by m to obtain the average distance, which is called the Mean Absolute Error (MAE). The larger the mean absolute error value, the greater the prediction deviation.

[0138] The performance of the model on the test set: The evaluation index is MAE. The average absolute deviation of the retention time is arranged from large to small, and 33 / 48 of them are less than 0.02, indicating that the prediction accuracy of the retention time is relatively high.

[0139] 10-fold cross-validation:

[0140] To most economically utilize the existing data, N-fold cross-validation (CV) is usually used, and the data set is divided into N folds (i.e., usually 5-fold or 10-fold CV). In such N-fold CV, one fold is reserved as the test data, and the remaining folds are used as the training data for model building. For example, in 10-fold CV, 1 fold is omitted as the test data, and the remaining 9 folds are grouped together as the training data for model building. Then, the trained model is applied to the above test data. This process is repeated until all folds have the opportunity to be set aside as test data. Therefore, 10 models will be built (i.e., each of the 10 folds is set aside as the test set), and each of the 10 models contains relevant performance metrics. Finally, the metric value is the average performance calculated based on 5 models.

[0141] The MAE of the 10-fold cross-validation results shows that, from the evaluation index data, the performance of the model on the training set is relatively stable.

[0142] Validation set data and preprocessing

[0143] The validation set is divided into 2 batches. One batch is batches 26 - 34 (a total of 776 samples), and the other batch is batches 24 - 25 with relatively large retention time deviations (a total of 165 samples). The data is preprocessed and optimized: Regarding the processing of 0, use the non-zero value in the previous row to replace 0 to avoid large interference to the results caused by 0.

[0144] The results of the average MAE for each batch from 24 to 34 of the validation set are shown in Table 6:

[0145] Table 6: Results table of the average MAE for each batch from 24 to 34 of the validation set

[0146]

[0147]

[0148] Figure 3For the verification result of an ion pair in the validation set, the X-axis represents the true retention time, and the Y-axis represents the predicted retention time. Since there will be overlapping points, the more overlapping there is, the darker the color is set.

[0149] Figure 4 These are the two samples with the highest MAE. Among them, x is the true retention time of the ion pair, and y is the predicted retention time of the ion pair. The whole graph is a scatter plot distribution of negative ion pairs, and it can be intuitively seen that the retention time offset of the ion pair is relatively obvious.

[0150] Figure 5 These are the two samples with the lowest MAE. The scatter points of this ion pair are basically on this 45° diagonal line, indicating that the prediction of the retention time is basically accurate.

[0151] Conclusion 1: Determine the threshold of the ion pair

[0152] The predicted retention time and the calculated MAE can be obtained through model prediction. The MAE threshold is set to 0.02, and if it is less than 0.02, it is considered an ion pair that is relatively easy to predict.

[0153] Conclusion 2: Determine the sample threshold

[0154] This summarizes 941 samples in the validation set. If the threshold is determined according to 80% of the number of samples, the sample threshold MAE should be less than 0.03. When the threshold of the sample exceeds 0.03, the ion pair of this sample should be focused on.

[0155] Conclusion 3: The degree of offset can be judged by the batch MAE value

[0156] According to the predicted sample results, calculate the average MAE value (mean_MAE) of all samples in each batch. It can be known whether there is an offset in which batch and the degree of offset. We can judge according to the mean_MAE value of the batch. The threshold is set to 0.03. If it is greater than this threshold, a relatively large retention time offset has occurred. When a relatively large retention time offset occurs, corresponding treatments should be carried out in the experiment and calculation.

[0157] 4 Summary

[0158] Based on the research of the retention time prediction model in Example 3, through decision tree regression, a prediction model is established based on 5 spiked working solution ion pairs in negative ions ("chlorzoxazone", "indapamide", "nimesulide", "mefenamic acid", "benzbromarone") and 6 spiked working solution ion pairs in positive ions ("R-ropivacaine mesylate", "pirfenidone", "capecitabine", "deflazacort", "chlorpromazine hydrochloride (SP55)", "rifampicin"); a total of 11 spiked working solution ion pairs.

[0159] In actual application, the predicted results will correctly predict the ion pairs that can be well predicted, quickly adjust their retention times, focus more on the ion pairs that are difficult to predict, manually confirm their correct retention times, and judge whether there is an offset and the degree of offset according to the average absolute deviation of the calculated batches, and determine what adjustments need to be made in experiments, experimental instruments, and data processing.

[0160] Applied to the integration confirmation stage before QC (Quality Control), it can help detect the ion pairs of samples that are prone to problems in integration. It has good practical application value.

[0161] The above is only an example display of the embodiments of the present invention, and does not impose any formal restrictions on the present invention. The protection scope of the present invention is subject to the claims and is not limited by the above specific embodiments. Any simple modification, equivalent change, or modification made to the above embodiments based on the technical essence of the present invention belongs to the protection scope of the present invention.

Claims

1. A method for constructing a system for automatically correcting the retention time of a group of ions to be measured in liquid chromatography-mass spectrometry based on a spiking system, characterized in that it includes the following steps: establishing a detection method for a high-performance liquid chromatography-tandem mass spectrometry system that simultaneously detects a group of ions to be measured in human serum and a spiking working solution; obtaining the retention times of the group of ions to be measured in serum and the ions of the spiking working solution, and establishing a database with the detected sample data; preprocessing the detected sample data, and dividing the detected sample data in the database into a training set and a test set according to a ratio; selecting a machine learning algorithm to construct a retention time prediction model based on the ions of the spiking working solution using the training set, testing and verifying the reliability of the retention time prediction model using the test set and evaluating it with evaluation indicators; performing N-fold cross-validation on the prediction model, then verifying it on the validation set, calculating the prediction accuracy evaluation indicators, and evaluating the prediction results; The detection method for the high-performance liquid chromatography-tandem mass spectrometry system that simultaneously detects a group of ions to be measured in human serum and a spiking working solution includes the following steps: configuring a drug reference substance into a mixed working solution, taking an appropriate amount of the working solution and adding it to the sample to be detected, and performing pretreatment with the sample to be detected; injecting the pretreated sample into the high-performance liquid chromatography-tandem mass spectrometry system for detection, and determining the retention time of each reference substance based on the chromatographic peaks generated by the characteristic MRM ion pairs of each reference substance; The MRM parameter table is: The drug reference substances are R-ropivacaine mesylate, pirfenidone, capecitabine, chlorpromazine hydrochloride, deflazacort, rifampicin, chlorzoxazone, indapamide, nimesulide, mefenamic acid, and benzbromarone.

2. The method for constructing a system for automatically correcting the retention time of a group of ions to be measured in liquid chromatography-mass spectrometry based on a spiking system according to claim 1, characterized in that: The pretreatment includes the following steps: after taking out the serum sample to be detected from the ultra-low temperature refrigerator, thawing it in an ice bath, and after the sample is completely thawed, vortexing for 60 seconds, precisely measuring 80 μL of the serum sample with the lid open and placing it in a 1.5 mL centrifuge tube; adding 5 μL of the spiking working solution to the serum sample and vortexing for 60 seconds; adding 150 μL of the sample precipitant and vortexing for 120 seconds; adding 50 μL of the ammonium formate buffer solution and vortexing for 120 seconds; centrifuging at 13000 rpm for 5 minutes at room temperature; precisely measuring 60 μL of the supernatant into an EP tube containing 150 μL of water, mixing evenly and loading it into an injection vial, waiting for injection and detection; the sample precipitant is acetonitrile: isopropanol with a volume ratio of 8:2; the ammonium formate buffer solution is prepared by mixing ammonium formate: water at 0.6 g: 1 ml.

3. The method for constructing a system for automatically correcting the retention time of a group of ions to be measured in liquid chromatography-mass spectrometry based on a spiking system according to claim 2, characterized in that: The method for preparing the spiking working solution: precisely measuring 10 μL of the mixed drug reference substance spiking intermediate working solution, adding it to 990 μL of methanol solution, and using it after vortexing for 60 seconds; The spiking working solution needs to be prepared and used immediately and cannot be stored. Method for preparing the spiked intermediate working solution of the mixed drug reference substance: Take out the stock solutions of each drug reference substance, thaw them to room temperature and vortex for 60 seconds; accurately measure the stock solutions into another separate container, add methanol respectively and vortex for 60 seconds; accurately measure the stock solutions or diluted stock solutions into the same bottle; add methanol to the vial of the mixed stock solution and vortex for 60 seconds; accurately measure the stock solutions of each drug reference substance into the same brown glass bottle respectively; Method for preparing the stock solutions of each drug reference substance: Weigh an appropriate amount of each drug reference substance accurately and place it in a brown glass bottle, add an appropriate amount of methanol to dissolve the drug reference substance, ensuring that the mass concentration of each drug reference substance stock solution is 1.0 mg / mL; the prepared drug reference substance stock solutions are stored in an ultra-low temperature refrigerator at -80 °C, with a validity period of 6 months; when in use, it needs to be thawed to room temperature at room temperature and vortexed for 60 seconds before opening the lid for use.

4. The method for constructing a system for automatically correcting the retention time of the ion pairs to be measured in liquid chromatography-mass spectrometry based on the spiked system according to claim 1, characterized in that, the detection and analysis conditions of the high performance liquid chromatography-tandem mass spectrometry system are: High performance liquid chromatography conditions are: Injection volume: 15 μL; Column temperature: 35 °C; Sample chamber temperature: 15 °C; Flow rate: 0.4 ml / min; Gradient elution: Mobile phase A: Aqueous solution of 0.1% formic acid, Mobile phase B: Acetonitrile solution of 0.1% formic acid The mobile phase gradient elution program is: 。 5. The method for constructing a system for automatically correcting the retention time of the ion pairs to be measured in liquid chromatography-mass spectrometry based on the spiked system according to claim 1, characterized in that, the training process of the retention time prediction model for the ion pairs to be measured is specifically as follows: Divide the detection sample data into a training set and a test set according to a ratio of 4:1, establish a model through decision tree regression of the training set samples; finally, verify the effect of the model on the test set, determine the final model, predict the retention time of the ion pairs to be measured on the validation set through the model, and calculate the evaluation index of the corresponding model prediction effect.

6. The method for constructing a system for automatically correcting the retention time of the ion pairs to be measured in liquid chromatography-mass spectrometry based on the spiked system according to claim 1, characterized in that, the preprocessing is to clean and sort the data, correct missing values, normalize / standardize the values to make them comparable, and transform the data.

7. The method for constructing a system for automatically correcting the retention time of the ion pairs to be measured in liquid chromatography-mass spectrometry based on the spiked system according to claim 5, characterized in that, the mean absolute error MAE is used as the evaluation index for the model prediction effect, and the calculation formula of the mean absolute error MAE is: where N is the total number of samples, y ai and y pi represent the true value and the predicted value, respectively.

8. The method for constructing a system for automatically correcting the retention time of the ion pairs to be measured in liquid chromatography-mass spectrometry based on the spiked system according to claim 1, characterized in that, The ten-fold cross-validation method is adopted to verify the regression decision tree model. The verification steps include: randomly dividing the data set into 10 parts, taking 9 of them as training data and 1 as verification data in turn to verify the model. The corresponding mean absolute error value will be obtained for each verification. The average value of the mean absolute error values for 10 times is used as the output of the model, and the model and its parameters are verified through this output.

9. The construction method of the retention time automatic correction system for the liquid-phase mass spectrometry measured ion pair group based on the labeling system according to claim 1, characterized in that, the prediction accuracy evaluation indexes include the threshold of the ion pair, the sample threshold, and the batch MAE value; determination of the threshold of the ion pair: the predicted retention time and the calculated MAE are obtained through model prediction, and the threshold of the ion pair is determined to be 0.

02. An ion pair with a value less than 0.02 is considered to be an ion pair that is relatively easy to predict; determination of the sample threshold: when the threshold is 80% of the number of samples, the sample threshold MAE should be less than 0.

03. When the threshold of the sample exceeds 0.03, the ion pairs of this sample should be focused on; determination of the batch MAE value: according to the predicted sample results, calculate the mean MAE value mean_MAE of all samples in each batch, and it can be known whether a shift has occurred in which batch and the degree of the shift. Judgment is made according to the mean_MAE value of the batch, and the threshold is set to 0.

03. If it is greater than this threshold, a large retention time shift has occurred. When a large retention time shift occurs, corresponding processing is carried out in experiments and calculations.

Citation Information

Patent Citations

  • High-coverage lipidomics analysis method based on liquid chromatography-mass spectrometer

    CN109870536A

  • Automated Spectral Library Retention Time Correction

    US20210293764A1