Non-targeted mass spectrometric identification of carbonyl compounds in tobacco

The method uses chemical derivatization and multi-stage data filtering to enhance the detection and identification of aldehyde and ketone compounds in tobacco by overcoming sensitivity and interference challenges in LC-MS data, achieving efficient and accurate structural annotation.

JP7793645B2Active Publication Date: 2026-01-05CHINA TOBACCO YUNNAN IND
View PDF 11 Cites 0 Cited by

Patent Information

Application Number
JP2023566919
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-25
Filing Date
2023-06-27
Publication Date
2026-01-05
Estimated Expiration
2043-06-27

AI Technical Summary

Technical Problem

Current methods for analyzing carbonyl compounds in tobacco face challenges in achieving high-sensitivity detection due to their low ionization efficiency, and a wide concentration range, and complex LC-MS data sets are overwhelmed by interfering components, making it difficult to identify low-abundance aldehyde and ketone compounds.

Method used

A method involving chemical derivatization with 2,4-dinitrophenylhydrazine (DNPH) and isotope-labeled DNPH-d3, combined with UPLC-Orbitrap-HRMS analysis and multi-stage data filtering based on statistical characteristics, mass defects, and multi-ion mass spectrometry, to efficiently identify and distinguish aldehyde and ketone compounds in tobacco.

Benefits of technology

The method effectively removes noise and interference, enabling highly sensitive detection and identification of aldehyde and ketone compounds in tobacco by reducing complex LC-MS data through automated processing, achieving a high reduction in chromatographic peaks and improving the accuracy of structural annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007793645000012
    Figure 0007793645000012
  • Figure 0007793645000013
    Figure 0007793645000013
  • Figure 0007793645000014
    Figure 0007793645000014
Patent Text Reader

Abstract

The method for non-targeted mass spectrometric identification of tobacco carbonyl components includes the steps of: (1) preparing a tobacco sample extract; (2) derivatizing the extract with 2,4-dinitrophenylhydrazine (DNPH) to obtain a first tobacco sample; (3) derivatizing the extract with DNPH-d3 instead of DNPH to obtain a second tobacco sample; (4) preparing a mixed tobacco sample; (5) preparing a blank sample; (6) performing UPLC-Orbitrap-HRMS analysis on each of the first tobacco sample, the second tobacco sample, the mixed tobacco sample, and the blank sample to obtain LC-MS raw data; (7) processing the LC-MS raw data to obtain raw mass spectrometry characteristic data; (8) performing multistage filtering on the raw mass spectrometry characteristic data; and (9) performing structural annotation or identification on the finally retained chromatographic peaks to obtain tobacco carbonyl components. By performing multi-stage filtering on the mass spectrometry profile data, the chromatographic-mass spectrometry information of noise or interference components in the raw data set can be rapidly removed to obtain compositional information of aldehyde and ketone chemical components in cigarettes.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to the technical field of tobacco, and more particularly to a method for non-targeted mass spectrometry identification of tobacco carbonyl compounds. The carbonyl compounds in tobacco can be obtained by filtering mass spectrometry profile data and performing structural identification on the retained chromatographic peaks. [Background technology]

[0002] Aldehydes and ketones are important carbonyl compounds commonly present in biological organisms and the human environment. Low-molecular-weight (<10 carbon chain length) aliphatic aldehydes or ketones can react with biological molecules such as DNA, proteins, and enzymes and are highly cytotoxic and genotoxic. Formaldehyde, acetaldehyde, and crotonaldehyde are classified as Category 1, Category 2B, and Category 3 human carcinogens by the International Agency for Research on Cancer (IARC), respectively. Acrolein is classified as a "hazardous air pollutant" by the US Environmental Protection Agency (US EPA), and formaldehyde, acetaldehyde, acetone, acrolein, propionaldehyde, crotonaldehyde, 2-butanone, and butyraldehyde are all on the Hoffmann list.

[0003] Aldehyde and ketone compounds are the most abundant aroma compounds in tobacco, with over 500 identified in cigarette smoke to date. Given their significant impact on aroma and safety in cigarettes, comprehensive qualitative and quantitative analysis of these compounds is crucial. Researchers have conducted extensive qualitative and quantitative analysis of aldehyde and ketone compounds in different systems using gas chromatography-mass spectrometry (GC-MS) or liquid chromatography-mass spectrometry (LC-MS). LC-MS combines the powerful separation capabilities of LC with the high specificity and sensitivity of MS, making it one of the most prominent techniques in the field of qualitative and quantitative analysis of small molecular weight compounds. Depending on the research objective, targeted or non-targeted data acquisition methods can be used to analyze the target compounds. Targeted LC-MS methods primarily involve quantitative analysis of known target compounds using methods such as external standards to provide accurate content information. However, targeted methods typically focus only on small amounts of target components for which chemical standards are available, and therefore many potentially important components may exist in complex sample systems, which cannot be targeted for analysis due to the lack of suitable chemical standards.

[0004] Compared to targeted analysis, non-targeted LC-MS techniques focus on analyzing and detecting all possible targets in a sample system. High-resolution mass spectrometry (HRMS), such as ion trap (orbitrap) mass spectrometers or time-of-flight (TOF) mass spectrometers, is widely used for non-targeted identification and detection of unknown components due to its extremely high resolution and mass accuracy. LC-HRMS can provide a wealth of data information, including accurate molecular weight, isotope distribution, and multi-stage mass analysis. It is currently widely used in the fields of metabolomics, environmental analysis, and food safety analysis, where it can obtain a large number of chromatographic-mass spectrometry signals, including known and unknown components, to comprehensively understand the chemical composition of a test sample.

[0005] Low-molecular-weight aldehyde and ketone compounds have diverse structures, large polarity differences, high volatility, low ionization efficiency, and a wide concentration range. Therefore, achieving highly sensitive, untargeted analysis of these compounds using LC-MS remains challenging. Chemical derivatization (chemical labeling) is an important method for improving the LC behavior of target compounds and enhancing MS detection sensitivity and specificity. Researchers have previously developed detection methods for aldehyde and ketone compounds using a series of chemical derivatization reactions combined with LC-QqQ-MS or LC-HRMS, using p-nitrophenylhydrazine, 2,4-dinitrophenylhydrazine, Girard's reagent T, and ammonium acetate-phenanthraquinone as derivatization reagents. These results suggest that combining derivatization reactions with LC-HRMS significantly improves the detection performance of LC-MS for low-molecular-weight aldehyde and ketone compounds.

[0006] However, data sets generated by LC-HRMS experiments are extremely large and complex, and processing these complex raw data to rapidly identify and determine target components in complex matrices has always been a challenging task. First, complex raw data contain tens of thousands of mass spectrometric features, including a large amount of redundant chromatographic-mass spectrometric information. Even blank samples containing no target components can contain thousands of chromatographic peaks with good peak shapes. Furthermore, complex samples have a wide concentration distribution of compounds, and the mass spectrometric signals of low-abundance target components can be overwhelmed by signals from background noise ions or matrix interference ions due to their low intensity. Therefore, it is difficult to intuitively grasp the chromatographic-mass spectrometric information of low-abundance target components from the raw data, increasing the difficulty of their qualitative and quantitative analysis.

[0007] Data cleaning and mass spectrometry information filtering can help efficiently identify and differentiate target components, especially those with very low abundance, from complex LC-HRMS datasets. Based on the chromatographic behavior and mass spectrometry characteristics of target compounds, different data cleaning methods and mass spectrometry information filtering strategies can be designed to efficiently filter interfering ions in raw data and target components. For example, in metabolomics analysis, the coefficient of variation of chromatographic peaks (mass spectrometry characteristics) in quality control samples can be used to evaluate stability and remove detected, abundant interfering chromatographic peaks. In drug metabolism analysis, setting a specific mass defect window can filter the chromatographic-mass spectrometry signals of many interfering components to reveal drug metabolites of the same type with a specific structure. In natural product analysis, neutral loss filtering and diagnostic fragment ion filtering can be used to rapidly identify chemical components with similar chemical structures. However, many currently proposed LC-MS filtering methods still rely on one-dimensional chromatographic or mass spectrometry information to set filters, resulting in low information filtering efficiency and excessive false positives. In particular, there is currently no mass spectrometry information filtering method that combines labeling with chemical derivatization. Therefore, by comprehensively utilizing information such as the statistical rules, chromatographic behavior, and mass spectrometric characteristics of the target components to design a multidimensional chromatography-mass spectrometry data filtering method, it is expected that the specificity and accuracy of filtering LC-MS raw data will be improved.

[0008] Overall, aldehyde and ketone compounds are important chemical components in tobacco, and non-targeted identification and identification of these compounds is of great significance. However, currently, efficient methods for achieving this goal are lacking. The identification and identification of aldehyde and ketone compounds in complex sample systems has the following difficulties: 1. Aldehyde and ketone compounds have small molecular weights and low ionization efficiency, making it difficult to directly achieve high-sensitivity detection using LC-MS technology; 2. The content of aldehydes and ketones has a wide distribution range, so the LC-MS method for analytical detection requires a wide dynamic range; 3. The raw LC-MS data contains a large amount of interfering components, which overwhelm the chromatographic-mass spectrometry signals of the target components, especially those with low abundance, making it very difficult to distinguish and identify the aldehyde and ketone components.

[0009] For this reason, the present invention is proposed. Summary of the Invention

[0010] The present invention chemically labels potential aldehyde and ketone compounds through derivatization, optimizes and establishes an LC-HRMS method to achieve highly sensitive detection of aldehyde and ketone derivatization products, and provides multidimensional data filtering technology based on statistical properties, chromatographic behavior, mass spectrometry characteristics, and isotope labeling information, enabling efficient differentiation and identification of aldehyde and ketone components in complex LC-MS data sets through automated processing.

[0011] The technical means of the present invention are as follows:

[0012] A non-targeted mass spectrometric method for identifying carbonyl-based components in tobacco is Step (1) of soaking a tobacco sample in water, adding a certain amount of acetonitrile, and shaking the sample to extract the tobacco sample; (2) collecting a certain amount of tobacco sample extract and adding a certain amount of 2,4-dinitrophenylhydrazine (DNPH) thereto for derivatization to prepare a first tobacco sample (CCs-DNPH); Step (3) of preparing a second isotope-labeled tobacco sample (CCs-DNPH-d3) by repeating step (2) but using DNPH-d3 instead of DNPH; (4) collecting the first tobacco sample and the second tobacco sample in a 1:1 ratio, thoroughly mixing them, and then filtering them through a filtration membrane to obtain a third tobacco sample; Step (5) of preparing a blank sample according to steps (1) and (2) in the same manner as steps (1) and (2), except that the tobacco sample is soaked in water; (6) performing UPLC-Orbitrap-HRMS analysis on each of the first tobacco sample, the second tobacco sample, the third tobacco sample, and the blank sample to obtain LC-MS raw data; (7) performing peak detection, peak alignment and peak grouping processes on the LC-MS raw data to obtain some original mass spectrometry characteristics of the target component and the interfering component present together, including m / z values, retention times and peak intensities; (8) performing multi-stage filtering on the raw mass spectrometry characteristic data obtained in step (7); Structural annotation for the final retained chromatographic peaks (Structural annotation) and (9) obtaining carbonyl-based components of tobacco by performing the identification.

[0013] Preferably, the conditions for UPLC-Orbitrap-HRMS analysis are: UPLC-HRMS was completed using an instrument platform consisting of a Dionex U3000 UHPLC system and a Q-Exactive mass spectrometer connected in series. The UPLC conditions were as follows: Syncronis C18 chromatography column (2.1 mm × 100 mm, 1.7 μm), column temperature: 40 °C, sample injection volume: 1 μL, mobile phase A: 0.1% formic acid aqueous solution, mobile phase B: acetonitrile, gradient elution program: 0-1 min (min) 95% B, 1-3 min 95%-60% B, 3-10 min 60%-10% B, 10-18 min 10% B, 18-19 min 10%-95% B, 19-20 min 95% B, flow rate: 0.2 mL / min. The mass analysis conditions were: spray voltage 3.7 kV, sheath gas flow rate 35 L / min, auxiliary gas flow rate 10 L / min, DL transfer tube temperature 350°C, data collection was performed in Full MS-DDA positive ion mode and negative ion mode, the mass-to-charge ratio scan range of the primary mass analysis was set to 100-1200 m / z, the resolution was 70,000, the mass-to-charge ratio scan resolution of the secondary mass analysis was 35,000, and the high-energy collision-induced dissociation voltage was 30 eV.

[0014] Preferably, the mass spectrometry characteristic data is subjected to multi-stage filtering, including mass spectrometry characteristic data filtering based on statistical characteristics, mass spectrometry characteristic data filtering based on mass defects, mass spectrometry characteristic data filtering based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and mass spectrometry characteristic filtering based on secondary multi-ion mass analysis, to finally obtain retained chromatographic peaks.

[0015] Preferably, in the step of filtering mass spectrometric characteristic data based on statistical characteristics, The coefficient of variation CV and fold change FC were calculated using the formula: CV = Int SQC / Int MQC × 100%, where Int SQC represents the standard deviation of the peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MQC represents the average peak intensity, FC=Int MQC / Int MBK where IntMQC represents the average peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MBK represents the average peak intensity of a particular chromatographic peak in a blank sample, If the CV≦30% and FC≧1.5, the corresponding chromatographic peak is retained.

[0016] Preferably, in the mass defect based mass spectrometry characteristic data filtering step, The formula is MD=|MZ-ceiling(MZ)|, where MZ represents the m / z value that is the exact mass of the precursor ion of a specific chromatographic peak detected in the first tobacco sample, and ceiling(MZ) represents the nominal mass that is the upper limit of the exact mass; If 0.02≦MD≦0.3 and m / z>209, the corresponding chromatographic peak is retained.

[0017] Preferably, in the mass spectrometry characteristic data filtering step based on the paired chromatographic peaks obtained by labeling through the derivatization reaction, for the paired chromatographic peaks detected in the third tobacco sample, |MZ-MZ d3 If |=3.0186, |IntP1-IntP2| / max(P1,P2)<30% and |RT1-RT2|<2, the corresponding chromatographic peaks are retained for the same time, where MZ and MZ d3 represent the m / z values, which are the accurate masses of the precursor ions of the non-isotopically labeled and isotope-d3-labeled derivatization products, respectively; IntP1 and IntP2 represent the peak areas of the first and second chromatographic peaks, respectively; and RT1 and RT2 represent the retention times of the first and second chromatographic peaks, respectively.

[0018] Preferably, in the mass analysis feature filtering method based on secondary multi-ion mass analysis, if fragment ions of m / z 76.018, m / z 120.008, m / z 122.024, m / z 135.019 and m / z 181.012 can be generated in the secondary mass analysis of the first tobacco sample, the corresponding chromatographic peaks are retained.

[0019] Preferably, methods for structural annotation or identification of the final retained chromatographic peaks include standard matching, database searching or degradation rule analysis.

[0020] Preferably, step (7) of performing peak detection, peak alignment, and peak grouping, step (8) of performing multistage filtering on the mass spectrometry characteristic data, and step (9) of performing structural annotation or identification on the finally retained chromatographic peaks are performed automatically by a software package. The present invention develops a data processing package (MSFiltering package) based on the R language, which performs peak detection, peak alignment, and peak grouping on LC-MS raw data to obtain mass spectrometry characteristics in which the target component and interfering components coexist, and realizes efficient discrimination and identification of aldehyde and ketone components in complex LC-MS data sets by automating multidimensional data filtering methods based on statistical characteristics, mass defects, paired chromatographic peaks, and multi-ion information, and structural annotation or identification on the finally retained chromatographic peaks.

[0021] [Effects of the invention] The present invention provides the following advantageous effects.

[0022] 1. The method of the present invention performs multi-stage filtering on mass spectrometry profile data to rapidly remove noise or interference components from the chromatographic-mass spectrometry information of raw data sets, and efficiently identify chromatographic peaks and mass spectrometry profiles that are definitely attributable to aldehyde and ketone chemical components from complex untargeted data sets. The extracted chromatographic peaks and mass spectrometry profiles attributable to aldehyde and ketone chemical components are annotated with their chemical structures by means of standard matching, database search, decomposition rule analysis, etc., to obtain composition information of aldehyde and ketone chemical components in cigarettes or food samples.

[0023] 2. The present invention develops a data processing package (MSFiltering package) based on the R language. This data processing package performs peak detection, peak alignment, and peak grouping processes on LC-MS raw data to obtain mass spectrometry characteristics of the coexistence of target components and interfering components. Multidimensional data filtering methods based on statistical characteristics, mass defects, isotope labeling, and secondary multi-ion mass spectrometry information are then used, and finally, automated processing such as structural identification of retained chromatographic peaks is used to efficiently distinguish and identify aldehyde and ketone components in complex LC-MS data sets. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a schematic diagram showing the number of characteristic peaks of a blank sample in a comparative example, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on statistical characteristics, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on mass defects, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and the number of characteristic peaks after filtering of mass spectrometry characteristic data based on secondary multi-ion mass spectrometry. [Figure 2]1 is a schematic diagram showing the number of characteristic peaks of a mixed sample in a comparative example, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on statistical characteristics, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on mass defects, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and the number of characteristic peaks after filtering of mass spectrometry characteristic data based on secondary multi-ion mass spectrometry. [Figure 3] 1 is a schematic diagram showing the reproducibility of a standard mixture sample in a comparative example, the reproducibility after filtering of mass spectrometry characteristic data based on statistical characteristics, the reproducibility after filtering of mass spectrometry characteristic data based on mass defects, the reproducibility after filtering of mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and the reproducibility after filtering of mass spectrometry characteristic data based on secondary multi-ion mass spectrometry. [Figure 4] 1 is a schematic diagram showing the number of characteristic peaks of a tobacco sample in an example, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on statistical characteristics, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on mass defects, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and the number of characteristic peaks after filtering of mass spectrometry characteristic data based on secondary multi-ion mass spectrometry. DETAILED DESCRIPTION OF THE INVENTION

[0025] The present invention will be described in more detail below with reference to specific embodiments, but it should not be understood that the scope of the present invention is limited to the following examples. Various substitutions and modifications based on the knowledge and common practices of those skilled in the art should be included in the scope of protection of the present invention as long as they do not deviate from the spirit of the present invention.

[0026] Example: A non-targeted mass spectrometric method for the identification of carbonyl-based components in tobacco Step (1) of soaking a tobacco raw material sample in water, adding a certain amount of acetonitrile, and shaking and extracting the sample to obtain a tobacco sample extract; (2) collecting a certain amount of sample extract and adding a certain amount of 2,4-dinitrophenylhydrazine (DNPH) thereto for derivatization to prepare a first tobacco sample (CCs-DNPH); Step (3) of preparing a second isotope-labeled tobacco sample (CCs-DNPH-d3) by repeating step (2) but using DNPH-d3 instead of DNPH; (4) collecting the first tobacco sample and the second tobacco sample in a 1:1 ratio, thoroughly mixing them, and then filtering them through a filtration membrane to obtain a third tobacco sample; Step (5) of preparing a blank sample according to step (1) and step (2) without soaking the tobacco raw material sample in water; (6) performing UPLC-Orbitrap-HRMS analysis on each of the first tobacco sample, the second tobacco sample, the third tobacco sample, the blank sample, and the mixed sample to obtain LC-MS raw data; (7) performing peak detection, peak alignment and peak grouping processes on the LC-MS raw data to obtain mass spectrometric characteristics of both the target component and the interfering component, including m / z values, retention times and peak intensities; (8) performing multiple filtering on the mass spectrometry signature data; and (9) performing structural identification of the finally retained chromatographic peaks to obtain the carbonyl components of tobacco.

[0027] The specific steps are as follows:

[0028] In step 1, the tobacco material was cut into pieces of approximately 0.5 cm x 0.5 cm, and 1.0 g of the tobacco material was placed in a 100 mL stoppered Erlenmeyer flask with an accuracy of 0.1 mg. 5 mL of water was added, and the sample was completely immersed in water. Then, 30 mL of acetonitrile was added accurately, and the sample was shaken and extracted for 30 minutes at a rotation speed of 150 r / min in a shaker to obtain a sample extract. In step 2, 1.0 mL of tobacco sample extract was accurately transferred to a 10 mL volumetric flask, 4 mL of DNPH was added, and the volume was adjusted to the required volume with acetonitrile. The mixture was then shaken uniformly and left at room temperature for 30 minutes for derivatization to prepare the first tobacco sample (CCs-DNPH). Repeat step 2, substituting DNPH-d3 for DNPH in step 3, to prepare a second isotope-labeled tobacco sample (CCs-DNPH-d3). In step 4, 1 mL of the first tobacco sample (CCs-DNPH) and the second tobacco sample (CCs-DNPH-d3) were taken in a 1:1 ratio and mixed thoroughly, and then filtered through a 0.22 μm organic phase filtration membrane to obtain a third tobacco sample; In step 5, prepare a blank sample according to the above steps, except that the tobacco material sample is soaked in water; In step 6, UPLC-Orbitrap-HRMS analysis is performed on each of the first tobacco sample, the second tobacco sample, the third tobacco sample, and the blank sample to obtain LC-MS raw data; Regarding the conditions for UPLC-Orbitrap-HRMS analysis, UPLC-HRMS was completed using an instrument platform consisting of a Dionex U3000 UHPLC system and a Q-Exactive mass spectrometer connected in series. The UPLC conditions were as follows: Syncronis C18 chromatography column (2.1 mm × 100 mm, 1.7 μm), column temperature: 40 °C, sample injection volume: 1 μL, mobile phase A: 0.1% formic acid aqueous solution, mobile phase B: acetonitrile, gradient elution program: 0-1 min 95% B, 1-3 min 95%-60% B, 3-10 min 60%-10% B, 10-18 min 10% B, 18-19 min 10%-95% B, 19-20 min 95% B, flow rate: 0.2 mL / min. The mass spectrometry conditions were: spray voltage 3.7 kV, sheath gas flow rate 35 L / min, auxiliary gas flow rate 10 L / min, DL transfer tube temperature 350 °C, data collection was performed in full MS-DDA positive ion mode and negative ion mode, the mass-to-charge ratio scan range of the primary mass spectrometry was set to 100-1200 m / z, the resolution was 70000, the mass-to-charge ratio scan resolution of the secondary mass spectrometry was 35000, and the high-energy collision-induced dissociation voltage was 30 eV. In step 7, the LC-MS raw data is subjected to peak detection, peak alignment and peak grouping processes to obtain mass spectrometric characteristics of the co-existence of the target component and the interfering component, including m / z values, retention times and peak intensities; In step 8, the mass spectrometry characteristic data of step 7 is subjected to multi-stage filtering, including mass spectrometry characteristic data filtering based on statistical characteristics, mass spectrometry characteristic data filtering based on mass defects, mass spectrometry characteristic data filtering based on paired chromatographic peaks obtained by labeling through derivatization reactions, and mass spectrometry characteristic filtering based on secondary multi-ion mass spectrometry, to finally obtain retained chromatographic peaks; In the step of filtering mass spectrometry characteristic data based on statistical characteristics, The coefficient of variation CV and fold change FC were calculated using the formula: CV = Int SQC / Int MQC × 100%, where Int SQC represents the standard deviation of the peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MQC represents the average peak intensity, FC=Int MQC / Int MBK where Int MQC represents the average peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MBK represents the average peak intensity of a particular chromatographic peak in a blank sample, If the CV≦30% and FC≧1.5, the corresponding chromatographic peak is retained; In the step of filtering mass spectrometry characteristic data based on mass defects, The formula is MD=|MZ-ceiling(MZ)|, where MZ represents the m / z value that is the exact mass of the precursor ion of a specific chromatographic peak detected in the first tobacco sample, and ceiling(MZ) represents the nominal mass that is the upper limit of the exact mass; If 0.02≦MD≦0.3 and m / z>209, the corresponding chromatographic peak is retained.

[0029] In the mass spectrometry characteristic data filtering step based on the paired chromatographic peaks obtained by labeling through the derivatization reaction, |MZ-MZ d3 If |=3.01605, |IntP1-IntP2| / max(P1,P2)<30% and |RT1-RT2|<2, the corresponding chromatographic peaks are retained for the same time, where MZ and MZ d3 represent the m / z values, which are the accurate masses of the precursor ions of the non-isotopically labeled and isotope-d3-labeled derivatization products, respectively; IntP1 and IntP2 represent the peak areas of the first and second chromatographic peaks, respectively; and RT1 and RT2 represent the retention times of the first and second chromatographic peaks, respectively.

[0030] In the mass analysis feature filtering method based on secondary multi-ion mass spectrometry, if the secondary mass analysis of the first tobacco sample can generate fragment ions of m / z 76.018, m / z 120.008, m / z 122.024, m / z 135.019 and m / z 181.012, the corresponding chromatographic peaks are retained.

[0031] In step 9, methods for structural annotation or identification of the final retained chromatographic peaks include standard matching, database searching, or degradation rule analysis.

[0032] Steps 7, 8 and 9 above are performed automatically by a software package.

[0033] The present invention develops a data processing package (MSFiltering package) based on the R language. This data processing package performs peak detection, peak alignment, and peak grouping processes on LC-MS raw data to obtain mass spectrometric characteristics of the coexistence of target components and interfering components. It then uses multidimensional data filtering methods based on statistical characteristics, mass defects, isotope labeling, and secondary multi-ion mass spectrometric information, and finally automates structural annotation or identification of the retained chromatographic peaks, thereby achieving efficient discrimination and identification of aldehyde and ketone components in complex LC-MS data sets.

[0034] Figure 4 is a schematic diagram showing the number of characteristic peaks of a tobacco sample in an example, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on statistical characteristics, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on mass defects, the number of characteristic peaks after filtering of mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through a derivatization reaction, and the number of characteristic peaks after filtering of mass spectrometry characteristic data based on secondary multi-ion mass spectrometry.

[0035] As can be seen from Figure 4, the original number of characteristic peaks of the tobacco sample was 12,601. After mass spectrometry characteristic data filtering (SCF) based on statistical characteristics, the number of retained chromatographic peaks was reduced to 3,452 (a 72.6% reduction). After mass spectrometry characteristic data filtering (MDF) based on mass defects, the number of retained chromatographic peaks was further reduced to 975 (a 92.3% reduction). After mass spectrometry characteristic data filtering (PPF) based on paired chromatographic peaks obtained by labeling via derivatization reactions, the number of retained chromatographic peaks was further reduced to 169 (a 98.6% reduction). After mass spectrometry characteristic data filtering (DFIF) based on secondary multi-ion mass spectrometry, the number of retained chromatographic peaks was further reduced to 93 (a 99.3% reduction). By evaluating these EIC masses, molecular formula predictions, and primary and secondary mass spectrometry masses, 70 chromatographic peaks were ultimately retained. Structural identification was performed for these retained chromatographic peaks by comparing retention times, MS, and MS / MS data based on standard comparisons and literature or database searches. For unknowns without standards or known mass spectrometry information, mass spectrometry decomposition rules were analyzed using Frontier 7.0 (Thermo Fisher Scientific) to annotate their chemical structures. The identification results are shown in Table 1, including information such as molecular formulas, secondary mass spectrometry analysis, and annotation results. Structural annotations were obtained for 40 chromatographic peaks, and the corresponding chemical structures of another 20 chromatographic peaks require further estimation. The annotated chromatographic peaks include hydrazones (DNPH-derivatized products) of aldehydes or ketones, such as formaldehyde, acetaldehyde, 2,3-butanedione, methylglyoxal, hydroxyacetone, 2-furaldehyde, 5-hydroxymethylfurfural, benzaldehyde, and salicylaldehyde.

[0036] The results show that there is a large amount of redundant information in raw LC-MS data, and the method of the present invention can efficiently remove interfering information through multi-stage mass spectrometry filtering and rapidly identify potential aldehyde and ketone chemical components from untargeted datasets.

[0037] In Comparative Example 1, the same multi-stage filtering was performed on the mass spectrometry characteristic data of the blank sample prepared in the Example. The results are shown in Figure 1. As can be seen from Figure 1, the original number of characteristic peaks of the blank sample was 7,372. After filtering the mass spectrometry characteristic data based on statistical characteristics, the number of retained chromatographic peaks was reduced to 641. After filtering the mass spectrometry characteristic data based on mass defects, the number of retained chromatographic peaks was reduced to 51. After filtering the mass spectrometry characteristic data based on paired chromatographic peaks obtained by labeling through derivatization reaction, the number of retained chromatographic peaks was reduced to 0.

[0038] In Comparative Example 2, the mass spectrometry characteristic data of the standard mixture sample is subjected to the same multi-stage filtering. The preparation steps of the standard mixture sample are the same as those in the Example. 24 aldehyde standards and 24 ketone standards were added to the standard mixture sample. The 24 aldehydes were formaldehyde, acetaldehyde, acrolein, glyoxal, n-propionaldehyde, crotonaldehyde, malondialdehyde, n-butyraldehyde, valeraldehyde, 2-furaldehyde, glutaraldehyde, hexanal, benzaldehyde, 5-methylfurfural, n-heptylaldehyde, phenylacetaldehyde, salicylic aldehyde, 1-octyl aldehyde, trans-cinnamaldehyde, 2,5-dimethylbenzaldehyde, p-methoxybenzaldehyde, 2,4-nonadienal, 2,4-decadienal, and decamethylbenzaldehyde. The 24 ketones were acetone, cyclopentanone, 2,3-butanedione, 3-methyl-2-cyclopentenone, cyclohexanone, 2-methyltetrahydrofuran-3-one, 3-hepten-2-one, 4-heptanone, acetophenone, 2,3-heptanedione, isophorone, α-ionone, hydroxyacetone, 2-pentanone, acetoin, 2,3-pentanedione, methyl isobutyl ketone, 2,3-hexanedione, 2-heptanone, acetoxy-2-propanone, 6-methyl-3,5-heptadiene-2-one, 6-methyl-6-hepten-2-one, 4-methylacetophenone, and 5-nonanone, each at a concentration of 0.1 mg / mL. The results of multistage filtering of the mass spectrometry characteristic data of the standard mixture sample are shown in Figure 2.As can be seen from Figure 2, the original number of characteristic peaks for the standard mixture sample was 8,021. After filtering the mass spectrometry characteristic data based on statistical characteristics, the number of retained chromatographic peaks was reduced to 2,440. After filtering the mass spectrometry characteristic data based on mass defects, the number of retained chromatographic peaks was reduced to 377. After filtering the mass spectrometry characteristic data based on paired chromatographic peaks obtained by derivatization, the number of retained chromatographic peaks was reduced to 109. After filtering the mass spectrometry characteristic data based on paired chromatographic peaks obtained by derivatization, the number of retained chromatographic peaks was reduced to 46. Observing the reconstructed TIC after filtering revealed that many mass spectrometry interference signals were removed, and target chromatographic peaks not present in the original TIC were prominently revealed. This dramatic reduction in the amount of mass spectrometry interference significantly narrowed the range of identification for aldehyde-ketone compounds. Figure 3 shows the recall obtained by filtering the mass spectrometry characteristic data for the standard mixture sample using different filtering methods. As can be seen from Figure 3, the recall rates of the raw dataset, statistical feature filtering, and the combination of statistical feature filtering and mass defect filtering were all 100%, while the recall rates of the combination of statistical feature filtering, mass defect filtering, paired chromatographic peak filtering, and multi-ion filtering were greater than 91%, indicating that all added standards could be effectively identified, proving the effectiveness of the multi-stage filtering method.

[0039] The embodiments are merely preferred embodiments of the present invention and are not intended to limit the present invention, and those skilled in the art may make various modifications and changes to the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0040] Table 1. Identification results of aldehyde and ketone components using UHPLC-Q-Orbitrap-MS / MS TIFF0007793645000001.tif204170TIFF0007793645000002.tif195170TIFF0007793645000003.tif196170TIFF0007793645000004.tif196170TIFF0007793645000005.tif196170TIFF0007793645000006.tif196170TIFF0007793645000007.tif196170TIFF0007793645000008.tif196170TIFF0007793645000009.tif196170TIFF0007793645000010.tif196170TIFF0007793645000011.tif204170

Claims

1. Step (1) of soaking a tobacco sample in water, adding a certain amount of acetonitrile, and shaking the sample to extract the tobacco sample; (2) collecting a certain amount of tobacco sample extract and adding a certain amount of 2,4-dinitrophenylhydrazine (DNPH) thereto for derivatization to prepare a first tobacco sample (CCs-DNPH); DNPH-d instead of DNPH 3 Step (2) was repeated using the second isotope-labeled tobacco sample (CCs-DNPH-d 3 (3) preparing a (4) collecting the first tobacco sample and the second tobacco sample in a 1:1 ratio, thoroughly mixing them, and then filtering them through a filtration membrane to obtain a third tobacco sample; Step (5) of preparing a blank sample similar to steps (1) and (2), except that the tobacco sample is soaked in water; (6) performing UPLC-Orbitrap-HRMS analysis on each of the first tobacco sample, the second tobacco sample, the third tobacco sample and the blank sample to obtain LC-MS raw data; (7) performing peak detection, peak alignment and peak grouping processes on the LC-MS raw data to obtain some original mass spectrometry characteristics of both the target component and the interfering component, including m / z values, retention times and peak intensities; Step (8) of performing multi-stage filtering on the original mass spectrometry characteristic data obtained in step (7), including mass spectrometry characteristic data filtering based on statistical characteristics, mass spectrometry characteristic data filtering based on mass defects, mass spectrometry characteristic data filtering based on paired chromatographic peaks obtained by labeling through derivatization reaction, and mass spectrometry characteristic filtering based on secondary multi-ion mass spectrometry, to finally obtain retained chromatographic peaks; and (9) performing structural annotation or identification on the final retained chromatographic peaks to obtain the tobacco carbonyl-based components.

2. Regarding the conditions for UPLC-Orbitrap-HRMS analysis, UPLC-HRMS was completed using an instrument platform consisting of a Dionex U3000 UHPLC system and a Q-Exactive mass spectrometer connected in series. The UPLC conditions were as follows: Syncronis C18 chromatography column (2.1 mm × 100 mm, 1.7 μm), column temperature: 40°C, sample injection volume: 1 μL, mobile phase A: 0.1% formic acid aqueous solution, mobile phase B: acetonitrile, gradient elution program: 0-1 min 95% B, 1-3 min 95%-60% B, 3-10 min 60%-10% B, 10-18 min 10% B, 18-19 min 10%-95% B, 19-20 min 95% B, flow rate: 0.2 mL / min. The method according to claim 1, wherein the mass spectrometry conditions are a spray voltage of 3.7 kV, a sheath gas flow rate of 35 L / min, an auxiliary gas flow rate of 10 L / min, a DL transfer pipe temperature of 350°C, data collection is performed in a positive ion mode and a negative ion mode of Full MS-DDA, the mass-to-charge ratio scan range of the primary mass spectrometry is set to 100 to 1200 m / z, the resolution is 70,000, the mass-to-charge ratio scan resolution of the secondary mass spectrometry is 35,000, and the high-energy collision-induced dissociation voltage is 30 eV.

3. In the step of filtering mass spectrometry characteristic data based on statistical characteristics, The coefficient of variation CV and fold change FC were calculated using the formula: CV = Int SQC / Int MQC × 100%, where Int SQC represents the standard deviation of the peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MQC represents the average peak intensity, FC=Int MQC / Int MBK where Int MQC represents the average peak intensity of a particular chromatographic peak in the first tobacco sample, and Int MBK represents the average peak intensity of a particular chromatographic peak in a blank sample, 2. The method of claim 1, wherein if CV≦30% and FC≧1.5, the corresponding chromatographic peak is retained.

4. In the step of filtering mass spectrometry characteristic data based on mass defects, The calculation formula is MD=|MZ-ceiling(MZ)|, where MZ represents the m / z value that is the exact mass of the precursor ion of a specific chromatographic peak detected in the first tobacco sample, and ceiling(MZ) represents the nominal mass that is the upper limit of the exact mass; 2. The method of claim 1, wherein if 0.02≦MD≦0.3 and m / z>209, the corresponding chromatographic peak is retained.

5. In the mass spectrometry characteristic data filtering step based on the paired chromatographic peaks obtained by labeling through the derivatization reaction, for the paired chromatographic peaks detected in the third tobacco sample, |MZ-MZ d3 |=3.0186, |IntP1-IntP2| / max(P1, P2)<30% and |RT1-RT2|<2, then the corresponding chromatographic peaks are retained for the same time, where MZ and MZ d3 10. The method of claim 1, wherein m / z values ​​represent the accurate masses of precursor ions of the non-isotopically labeled and isotopically d3-labeled derivatized products, respectively; IntP1 and IntP2 represent the peak areas of the first and second chromatographic peaks, respectively; and RT1 and RT2 represent the retention times of the first and second chromatographic peaks, respectively.

6. 2. The method of claim 1, wherein in the mass analysis feature filtering method based on secondary multi-ion mass spectrometry, if fragment ions at m / z 76.018, m / z 120.008, m / z 122.024, m / z 135.019 and m / z 181.012 can be generated in the secondary mass analysis of the first tobacco sample, the corresponding chromatographic peaks are retained.

7. 2. The method of claim 1, wherein the method for structural annotation or identification of the final retained chromatographic peaks comprises standard matching, database searching, or degradation rule analysis.

8. 2. The method of claim 1, wherein the steps of peak detection, peak alignment and peak grouping (7), multi-stage filtering of the mass spectrometry characteristic data (8), and structural annotation or identification of the final retained chromatographic peaks (9) are performed automatically by a software package.

Citation Information

Patent Citations

  • Method for measuring main carbonyl compounds in smoke-free tobacco by means of UPLC-IE method

    CN104950064A

  • Method for pre-processing tobacco or tobacco product and detection method thereof

    CN106018635A

  • Treatment method for smoke-free tobacco product and determination method for small-molecular aldehydes in smoke-free tobacco product

    CN107966518A

  • Non-targeted screening and quantitative detection methods for multiple pesticide residues in tobacco

    CN110646535A

  • UPLC-HRMS Profile mode non-target metabolism profile data automatic analysis method

    CN110806456A