A method and system for analyzing gas chromatography mass spectrometry data

By acquiring three-dimensional data using gas chromatography-mass spectrometry and separating overlapping peaks using a self-attention network, combined with chromatographic temperature data, high-precision automated analysis of organic pollutants in complex mixed samples was achieved. This solved the problem of low qualitative accuracy in traditional methods and provided accurate detection reports.

CN122430481APending Publication Date: 2026-07-21河南省许昌生态环境监测中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
河南省许昌生态环境监测中心
Filing Date
2026-05-25
Publication Date
2026-07-21

Smart Images

  • Figure CN122430481A_ABST
    Figure CN122430481A_ABST
Patent Text Reader

Abstract

The application discloses a gas chromatography-mass spectrometry data analysis method and system, relates to the gas chromatography-mass spectrometry combined technical field, and the method comprises the following steps: controlling a mass spectrometer to obtain three-dimensional chromatography-mass spectrometry data, performing overlapping peak separation by fusing local peak shape and global fragment correlation through a self-attention network, and obtaining a pure component mass spectrum. A molecular graph structure is constructed based on the graph, combined with measured temperature data, a candidate organic list is generated by using a gated recurrent unit; the confidence is determined by comparing the theoretical peak time and the measured time, and the highest confidence is selected as the target. After locking, it is automatically switched to a targeted collection mode for analysis and report generation. The application solves the technical problem that it is difficult to automatically separate single component information from mixed data and accurately lock target substances in the detection and analysis of complex mixed samples, achieves full-automatic and high-precision closed-loop analysis from original detection data to target substance confirmation and quantification, and reduces the technical effect of manual intervention and misjudgment risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of gas chromatography-mass spectrometry (GC-MS) technology, and particularly to a method and system for GC-MS data analysis. Background Technology

[0002] In ecological and environmental monitoring, accurate analysis of organic pollutants in environmental media is crucial for regional environmental risk assessment and pollution source tracing and control. Gas chromatography-mass spectrometry (GC-MS) is the core method for monitoring organic pollutants. Existing technologies mostly employ traditional data processing methods, relying on manual analysis or simple algorithms, which are effective for samples with single components and simple matrices. However, environmental samples have complex matrices, and the coexistence of multiple organic pollutants easily leads to overlapping chromatographic peaks and strong matrix interference. Traditional methods struggle to separate overlapping peaks and accurately obtain pure component mass spectrometry information. Furthermore, they do not incorporate key information such as chromatographic temperature and molecular structure, resulting in low qualitative accuracy and incomplete detection data, failing to meet the control requirements for accurate qualitative and quantitative analysis and efficient analysis of complex organic compounds in ecological and environmental monitoring. Summary of the Invention

[0003] This application provides a gas chromatography-mass spectrometry data analysis method and system, which solves the technical problem of difficulty in automatically separating information of a single component from mixed data and accurately locating the target substance in the detection and analysis of complex mixed samples.

[0004] The first aspect of this application provides a gas chromatography-mass spectrometry (GC-MS) data analysis method, comprising: controlling a mass spectrometer to detect a target in a full-scan mode to obtain three-dimensional GC-MS data; inputting the three-dimensional GC-MS data into an attention network to capture first feature data based on local chromatographic peak morphology and second feature data based on global mass spectrometry fragment correlation, performing overlapping peak separation under cross-constraint to obtain a pure component mass spectrum, wherein the first feature data provides time domain boundaries and the second feature data provides frequency domain correlation; constructing an organic molecular map structure based on the pure component mass spectrum, and simultaneously reading chromatographic temperature data based on measured peak elution times, inputting it into a gated loop unit to generate a candidate organic list; traversing the candidate organic list, predicting the theoretical peak elution time of each candidate organic, comparing it with the measured peak elution time to determine the confidence level, and selecting the candidate with the highest confidence level as the target organic; after locking the target organic, performing targeted parameter extraction and mode adaptive switching, sending the data to the scanning control module of the mass spectrometer for targeted acquisition and analysis, and generating a detection report.

[0005] A second aspect of this application provides a gas chromatography-mass spectrometry data analysis system, the system comprising: a three-dimensional chromatography-mass spectrometry data acquisition module, used to control the mass spectrometer to detect a target in a full-scan mode to obtain three-dimensional chromatography-mass spectrometry data; and a pure component mass spectrum acquisition module, used to input the three-dimensional chromatography-mass spectrometry data into a self-attention network, capture first feature data based on local chromatographic peak morphology and second feature data based on global mass spectrometry fragment correlation, perform overlapping peak separation under cross-constraints, and obtain a pure component mass spectrum, wherein the first feature data provides time-domain boundaries and the second feature data provides frequency-domain correlation; and candidate The organic compound list generation module is used to construct the molecular structure of organic compounds based on the mass spectrum of pure components, and simultaneously read chromatographic temperature data based on the measured peak elution time, inputting it into the gated loop unit to generate a candidate organic compound list; the target organic compound determination module is used to traverse the candidate organic compound list, predict the theoretical peak elution time of each candidate organic compound, compare it with the measured peak elution time to determine the confidence level, and select the candidate with the highest confidence level as the target organic compound; the detection report generation module is used to perform targeted parameter extraction and mode adaptive switching after locking the target organic compound, and send it to the scanning control module of the mass spectrometer for targeted acquisition and analysis to generate a detection report.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0007] This application acquires three-dimensional gas chromatography-mass spectrometry data of environmental media samples, obtains organic component data through intelligent splitting of overlapping peaks and multi-dimensional feature fusion analysis, calculates qualitative and quantitative information of organic matter, and corrects and adjusts the analysis results by combining chromatographic temperature sequence and molecular structure characteristics. This enables accurate detection of complex coexisting organic pollutants in the environment, making the detection results of organic pollutants in the ecological environment more accurate and reliable. It achieves a fully automated, high-precision closed-loop analysis from raw detection data to the confirmation and quantification of target substances, reducing the risk of human intervention and misjudgment. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a schematic flowchart of a gas chromatography-mass spectrometry data analysis method provided in an embodiment of this application.

[0010] Figure 2 This is a schematic diagram of the structure of a gas chromatography-mass spectrometry data analysis system provided in an embodiment of this application.

[0011] Figure labeling: 1. Three-dimensional chromatography-mass spectrometry data acquisition module; 2. Pure component mass spectrum acquisition module; 3. Candidate organic compound list generation module; 4. Target organic compound identification module; 5. Detection report generation module. Detailed Implementation

[0012] This application provides a gas chromatography-mass spectrometry data analysis method and system, which solves the technical problem of difficulty in automatically separating information of a single component from mixed data and accurately locating the target substance in the detection and analysis of complex mixed samples.

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0014] It should be noted that the terms "first," "second," etc., in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices.

[0015] Example 1, as Figure 1 As shown, a gas chromatography-mass spectrometry data analysis method is provided, wherein the method includes:

[0016] By controlling the mass spectrometer to detect the target in full-scan mode, three-dimensional chromatographic mass spectrometry data are obtained.

[0017] In this embodiment, the mass spectrometer is a precision analytical instrument used to determine the chemical composition of substances. In this scheme, the detection target is various organic substances in environmental media such as soil, water, and atmosphere.

[0018] In a specific embodiment, a soil sample is used as an exemplary environmental sample to be measured. This method can be widely applied to the detection of organic matters in various environmental samples such as water bodies, atmospheric particulate matters, sediment, etc. Only the pretreatment steps can be adjusted correspondingly according to the sample type, and the core processes of subsequent full-scan detection and three-dimensional data construction remain the same. All specific operation parameters in this embodiment are only examples and do not constitute a limitation on the protection scope of the present invention.

[0019] First, take the soil sample as an example to carry out the pretreatment operation: air-dry the collected soil sample naturally and then grind and sieve it. For example, a 80-mesh standard sieve can be used, and the sieve pore size can be flexibly adjusted according to the actual type of the soil. Use the Soxhlet extraction method to extract the organic components in the soil with an organic solvent. The organic solvent can be, for example, a mixed solution of n-hexane and acetone with a volume ratio of about 1:1, or other commonly used solvents in the industry such as dichloromethane and n-hexane can also be selected; the extraction temperature is set at 60 degrees Celsius and the extraction time is 12 hours, which can be appropriately increased or decreased according to the nature of the sample in actual operation. After the extraction is completed, carry out dehydration treatment with anhydrous sodium sulfate, concentrate by rotary evaporation, and then complete the purification through a C18 solid-phase extraction column or other conventional purification columns to remove solid impurities and interfering matrices, and obtain a sample solution to be measured that can be directly injected.

[0020] Then, transport the sample solution to be measured to the gas chromatograph, set the chromatographic column temperature programming and the carrier gas flow rate. The injection volume can be set at, for example, 1 μL, and the split injection mode can be adopted, and the split ratio can be set at 10:1, both of which can be adjusted according to the actual detection requirements. The chromatographic column can adopt a DB-5MS capillary chromatographic column or other similar stationary-phase chromatographic columns; the carrier gas can adopt high-purity helium with a purity not lower than 99.999%, and the carrier gas flow rate can be set at about 1.0 mL / min, which can be adjusted according to the actual separation effect. The temperature programming example is to keep the initial temperature at 40 degrees Celsius for about 2 minutes, and then increase the temperature to 280 degrees Celsius at a rate of about 5 degrees Celsius per minute and keep it for about 10 minutes. This program can be flexibly adjusted according to the types of target organic matters. Based on the differences in the adsorption and desorption rates of different organic components in the chromatographic column, component separation is achieved, and the separated components enter the ion source of the mass spectrometer in the order of peak emergence.

[0021] Open the full-scan detection mode through the control terminal supporting the mass spectrometer, and set the mass range and scan time interval of the mass spectrometer. The ion source can adopt, for example, an electron impact ionization source. Exemplarily, the ionization energy is 70 eV, the ion source temperature can be set at about 230 degrees Celsius, and the transfer line temperature can be set at about 280 degrees Celsius; the scan mass range is 35 - 450 amu, and the scan time interval is 0.2 seconds; the detector can adopt an electron multiplier. The mass spectrometer detector continuously collects ion signals and synchronously records the chromatographic retention time, ion mass-to-charge ratio, and ion signal intensity corresponding to each ion signal.

[0022] Finally, the collected retention time, mass-to-charge ratio, and signal intensity data were systematically integrated according to the time, mass-to-charge ratio, and signal intensity dimensions, respectively, and a complete three-dimensional chromatographic-mass spectrometry (GC-MS) dataset was constructed with time as the x-axis, mass-to-charge ratio as the y-axis, and intensity as the numerical axis. This data retains the original detection information of all organic components in various environmental samples, providing a comprehensive data foundation for subsequent organic component analysis and target pollutant identification.

[0023] By performing standardized pretreatment, stepwise chromatographic separation, and full-dimensional mass spectrometry signal acquisition on environmental samples, the complete acquisition of original three-dimensional detection data of organic components in environmental samples is achieved, providing a comprehensive and reliable raw data foundation for subsequent organic component analysis and target pollutant identification.

[0024] The three-dimensional chromatographic mass spectrometry data is input into an attention network to capture first feature data based on local chromatographic peak morphology and second feature data based on global mass spectrometry fragment correlation. Overlapping peak separation under cross-constraint is performed to obtain pure component mass spectra. The first feature data provides time domain boundaries, and the second feature data provides frequency domain correlation.

[0025] In this embodiment, the self-attention network is a neural network that captures the local and global correlation features of data, and is used to extract the correlation features of chromatographic peak morphology and fragment ion in three-dimensional chromatographic mass spectrometry data.

[0026] Optionally, the three-dimensional chromatographic mass spectrometry data, which consists of time, mass-to-charge ratio, and intensity, is reconstructed into sequence data and then input into a self-attention network. The first feature data and the second feature data are obtained by using the correlation between local chromatographic peak morphology and global mass spectrometry fragments as the attention target. The specific implementation process of this step will be described in detail below.

[0027] Next, based on the second feature data, the mass-to-charge ratio channels are first clustered to obtain the first clustering result. Then, cross-time window correlations are eliminated through time window constraints to obtain the second clustering result. Finally, the fragment ion intensities of the same organic compound are weighted and averaged to obtain the pure component mass spectrum. The specific implementation process of this step is also described in detail below.

[0028] The molecular structure of organic compounds is constructed based on the mass spectra of pure components. At the same time, chromatographic temperature data based on the measured peak elution time is read and input into the gated cycle unit to generate a list of candidate organic compounds.

[0029] In one embodiment of this application, an organic molecular map structure is constructed based on the pure component mass spectrum. After encoding the chromatographic temperature data, which includes the initial temperature, heating rate, final temperature, and holding time, into a temperature sequence, the sequence is input into a gated loop unit along with the organic molecular map structure to generate a candidate organic list. The specific implementation process of this step will be described in detail below.

[0030] The list of candidate organic compounds is traversed, the theoretical elution time of each candidate organic compound is predicted, and the confidence level is determined by comparing it with the measured elution time. The candidate with the highest confidence level is selected as the target organic compound.

[0031] Specifically, the candidate organic compound list is traversed to determine the first test group, which is obtained by splicing the first candidate organic compound with the temperature sequence. The theoretical peak time is predicted based on the first test group, and the measured peak time is compared with the theoretical peak time to generate the first confidence level. Then, the Nth test group based on the candidate organic compound list is determined in the same way and the Nth confidence level is generated. The first confidence level to the Nth confidence level are compared, and the candidate organic compound corresponding to the highest confidence level is selected as the target organic compound. The specific implementation process of this step will be described in detail below.

[0032] After identifying the target organic compound, the system performs targeted parameter extraction and adaptive mode switching, and sends the data to the mass spectrometer's scanning control module for targeted acquisition and analysis, generating a detection report.

[0033] Specifically, first, after locking onto the target organic compound, the optimal characteristic ion information and collision energy parameters, including the mass ratio of the parent ion, daughter ion, and characteristic fragments, are retrieved and sent to the mass spectrometer scanning control module. The scanning mode is then switched from full scan to multi-reaction monitoring mode to perform targeted acquisition and obtain targeted detection data. The specific implementation process of this step will be described in detail below.

[0034] Next, the targeted detection data is identified, and a concentration-response standard curve is established using the peak area of ​​the characteristic ion pairs of the target organic compound as the response value. The measured peak area is substituted into the standard curve to calculate the organic compound concentration. Qualitative confirmation is completed based on the deviation of the ion abundance ratio from the standard, and a detection report containing the organic compound name, retention time, organic compound concentration and ion comparison results is generated. The specific implementation process of this step is also described in detail below.

[0035] Furthermore, the method provided in this application embodiment includes:

[0036] Three-dimensional chromatographic mass spectrometry data are reconstructed into sequence data, wherein the three-dimensional chromatographic mass spectrometry data consists of time × mass-to-charge ratio × intensity, and each time step in the sequence data corresponds to a full mass spectrum vector of one scan cycle; the sequence data is input into an attention network, with local chromatographic peak morphology as the first attention target and global mass spectrometry fragment correlation as the second attention target, to capture first feature data and second feature data, wherein the local chromatographic peak morphology includes peak scale and symmetry, and the global mass spectrometry fragment correlation refers to the synchronous change of fragment ions of the same organic compound; based on the first feature data and the second feature data, the mass spectrum of the separated pure components is determined.

[0037] Specifically, environmental sample matrices are extremely complex. Soil, water, and atmospheric environmental samples generally contain a large amount of interfering substances such as humic substances, lipids, and pigments. During gas chromatography separation, co-elution often occurs, meaning that multiple organic compounds elute at similar or identical retention times, causing the mass spectrometry fragment signals of different components to overlap and form aliased mass spectra. Traditional single-peak extraction methods cannot accurately separate the pure mass spectrometry information of each component, thus affecting the accuracy of subsequent qualitative identification.

[0038] To address the aforementioned issues, the three-dimensional chromatographic mass spectrometry (GC-MS) data is first reconstructed into sequence data. GC-MS data consists of three dimensions: time, mass-to-charge ratio (MTR), and intensity. Following the chronological retention time, the MTR and intensity data for each scan cycle are expanded into a one-dimensional full-mass spectrum vector. All full-mass spectrum vectors are then arranged chronologically to form a complete sequence. Each scan cycle corresponds to a time step, and the dimension of the full-mass spectrum vector corresponds to the scan MTR range of the mass spectrometer. For example, when the scan MTR range is 35 to 450 amu, the dimension of each full-mass spectrum vector is approximately 416 dimensions; the length of the sequence data is consistent with the total number of scans. These dimensional values ​​are merely examples and will vary with the actual scan parameters.

[0039] Next, a lightweight self-attention network is constructed. This network consists of an input layer, a position encoding layer, a single-layer multi-head self-attention layer, a feedforward network layer, and a dual-branch output layer. The input layer receives sequence data with a dimension of sequence length × full mass spectrum vector dimension. The position encoding layer adds sinusoidal position encoding to the sequence data. The sine and cosine position encoding can adopt the Transformer standard sine and cosine position encoding formula to preserve the temporal order information. The single-layer multi-head self-attention layer adopts a 2-head attention structure to adapt to the extraction needs of two core features: local chromatographic peak morphology and global mass spectrometry fragment association. The hidden layer dimension can be set to 64. Two independent attention query branches are set: the first branch takes local chromatographic peak morphology as the attention target, and the attention window size is set to 5 time steps, which can be adjusted according to the actual data; the second branch takes global mass spectrometry fragment association as the attention target and adopts a global attention window. The feedforward network layer contains two linear layers and a ReLU activation function. The first linear layer maps the feature dimension to a higher dimension, such as 128 dimensions, and the second linear layer maps the feature dimension back to a lower dimension, such as 64 dimensions. The dual-branch output layer consists of two independent linear layers, connected in parallel from the output of the feedforward network layer, outputting first and second feature data respectively. The output of the first branch (local attention) passes through the first linear layer to obtain the first feature data, which characterizes the temporal boundary information of the chromatographic peak, such as start time, end time, and peak width. The output of the second branch (global attention) passes through the second linear layer to obtain the second feature data, which characterizes the synchronous changes and correlations of fragment ions, such as the similarity of intensity change trends in channels with different mass-to-charge ratios. The first feature data provides temporal boundary information, and the second feature data provides frequency domain correlation information.

[0040] The aforementioned lightweight self-attention network was pre-trained offline. The training dataset can be generated using publicly available standard mass spectrometry libraries, such as standard mass spectra of common environmental organic pollutants selected from the NIST mass spectrometry library. Two to five standard mass spectra were randomly selected and overlaid to simulate chromatographic mass spectrometry data with varying degrees of overlap, and the corresponding single-component features were labeled as training labels. In the training labels, the time-domain boundary labels are the start and end time indices of each single-component chromatographic peak, and the frequency-domain association labels are the mass-to-charge ratios of the characteristic ions corresponding to each single component. During training, batch gradient descent can be used, with a batch size of, for example, 32. The optimizer can be Adam, with an initial learning rate of, for example, 0.0001. The loss function is the mean squared error loss function. The loss values ​​between the first feature data and its corresponding time-domain boundary label, and between the second feature data and its corresponding frequency-domain association label are calculated separately. The two loss values ​​are then weighted and summed to obtain the total loss value. For example, with a total training run of 50 rounds, the learning rate can be decayed to 0.5 times its original value every 10 rounds. An independent validation set is used to monitor the model's training performance in real time. Training is terminated early when the total loss on the validation set fails to decrease for five consecutive rounds to prevent overfitting. The network weight parameters are saved after training. It should be noted that the above training process is an offline pre-training step. In actual detection applications, the saved weight parameters are directly loaded for forward computation, eliminating the need for repeated training.

[0041] In practical detection applications, the reconstructed sequence data is input into a pre-trained lightweight self-attention network, directly loading pre-trained weights. The position encoding layer adds positional information to the sequence data before inputting it into a single-layer multi-head self-attention layer. The first attention branch calculates attention weights within a local time window, capturing local morphological features such as peak size and symmetry for each chromatographic peak, generating the first feature data. The second attention branch calculates attention weights across the entire timeframe, capturing the synchronous changes in fragment ions of the same organic compound at different time steps, generating the second feature data. The dual-branch output layer outputs the feature data generated by the two branches separately. The first feature data is used to determine the temporal boundaries of chromatographic peaks for different components, and the second feature data is used to correlate all fragment ion signals of the same component.

[0042] Finally, fragment ions are clustered based on the second feature data, and cross-time window correlations are eliminated by time window constraints. Finally, the fragment intensity within the window is weighted and averaged to obtain the pure component mass spectrum. The specific implementation process of this step will be described in detail below.

[0043] By converting three-dimensional chromatographic mass spectrometry data into sequence data for an adaptive network, a dual-objective lightweight self-attention network is used to extract local chromatographic peak morphology and global mass spectrometry fragment correlation features, effectively solving the problem of mass spectrum aliasing caused by co-eluting of environmental samples. This achieves efficient and accurate extraction of overlapping peak features, providing reliable feature support for subsequent organic component separation.

[0044] Furthermore, the method provided in this application embodiment includes:

[0045] Based on the second feature data, the mass-to-charge ratio channels are clustered to determine the first clustering result, wherein fragment ions with high attention weights are classified as the same potential organic compounds that change synchronously; using the second feature data as a time window constraint, cross-time window associations are removed from the first clustering result to obtain the second clustering result, wherein the time window constraint is used to limit the elution start and end range of each organic compound on the chromatography; according to the second clustering result, the fragment ion intensities belonging to the same organic compound are weighted and averaged within each time window to obtain the pure component mass spectrum of the time window.

[0046] In this embodiment, the mass-to-charge ratio channel is an independent detection channel in the mass spectrometer that continuously acquires the intensity of ion signals corresponding to a single mass-to-charge ratio.

[0047] Optionally, based on the second feature data, the K-means clustering algorithm is used to cluster all mass-to-charge ratio channels. The second feature data includes the global attention weight and complete time-series variation features of each mass-to-charge ratio channel. During clustering, the feature vector corresponding to each mass-to-charge ratio channel is used as input. The maximum number of clusters is set to 10, which can be adjusted between 5 and 20 according to the expected amount of organic matter in the actual detected sample. The maximum number of iterations is 100, and the convergence threshold is 0.001. During the clustering process, mass-to-charge ratio channels with an attention weight higher than 0.5 and a Pearson correlation coefficient greater than 0.8 for the time-series variation trend are grouped into the same cluster. Each cluster corresponds to a group of synchronously changing fragment ions, representing a potential organic component.

[0048] Using the second feature data as a time window constraint, the first clustering results undergo cross-time window correlation elimination processing. Specifically, the start and end times of signal occurrence for all mass-to-charge ratio channels within each cluster are extracted from the second feature data. The start time of signal occurrence for a mass-to-charge ratio channel is defined as the first time step in which the channel's signal intensity exceeds three times the background noise, and the end time is defined as the last time step in which the signal intensity falls back to below three times the background noise. The minimum value of all start times is taken as the elution start time of the organic chromatographic compound corresponding to that cluster, and the maximum value of all end times is taken as the elution end time of the organic chromatographic compound. These are combined to obtain a complete elution time window. All mass-to-charge ratio channels in each cluster are traversed, and mass-to-charge ratio channels whose signal occurrence time exceeds the range of this time window are eliminated. This removes erroneous correlations between organic fragment ions eluted at different times, resulting in a second clustering result containing only fragment ions of the same organic characteristic.

[0049] Finally, based on the second clustering results, within the time window corresponding to each organic compound, a weighted average of the intensities of all fragment ions belonging to the same organic compound is calculated to obtain the pure component mass spectrum corresponding to each time window. During the weighted average calculation, the attention weights corresponding to each mass-to-charge ratio channel are used as weighting coefficients to sum the fragment ion intensities at each time step within the time window, and then the sum is divided by the total number of time steps within the time window to obtain the average intensity. This ultimately generates the pure component mass spectrum of the organic compound within the corresponding time window, which contains all characteristic fragment ions of the organic compound and their relative intensity information.

[0050] Organic components are initially divided by clustering based on the synchronous change characteristics of fragment ions. Then, cross-component miscorrelation is eliminated by time window constraint. Finally, high-purity mass spectra of pure components are generated by attention weighted averaging, which realizes efficient and accurate separation of overlapping chromatographic peaks and provides reliable basic data for subsequent qualitative and quantitative analysis of organic pollutants.

[0051] Furthermore, the method provided in this application embodiment includes:

[0052] Based on the mass spectrum of the pure components, an organic molecular map structure is constructed; chromatographic temperature data based on the measured peak elution time is read and encoded into a temperature sequence, wherein the chromatographic temperature data includes the initial temperature, heating rate, final temperature, and holding time; the organic molecular map structure and the temperature sequence are input into a gated loop unit to generate a candidate organic list.

[0053] Specifically, the pure component mass spectra are first initially matched with the NIST standard mass spectrometry library to obtain the top 20 candidate molecular formulas with the highest matching degree. Based on each candidate molecular formula, a molecular structure generation algorithm is used to generate all possible isomer structures. This can be implemented using commonly used industry tools such as RDKit or OpenBabel. Each isomer structure is represented as a SMILES string. The SMILES string is converted into a molecular graph structure. The molecular graph structure uses atoms as nodes and chemical bonds as edges. Each atom node is assigned a 4-dimensional node feature, including atomic number, formal charge, hybridization type, and number of connected atoms. Each chemical bond edge is assigned a 2-dimensional edge feature, including bond type and bond order. A single-layer graph convolutional network is used to aggregate the features of the molecular graph structure. The hidden layer dimension of the graph convolutional network is set to 64. During the aggregation process, each node integrates the feature information of its direct neighbor nodes. In this embodiment, the graph convolutional layer adopts a summation aggregation method, that is, the new feature vector of each node is equal to the sum of its own feature vector and the feature vectors of all neighboring nodes. After linear transformation and ReLU activation function, the result is output. The final output is a molecular embedding vector with a dimension of 64, which encodes key physicochemical properties of organic compounds such as molecular weight, functional group type, topological polar surface area, and lipid-water partition coefficient.

[0054] Next, chromatographic temperature data based on the measured peak elution time is read. This data includes the initial temperature, heating rate, final temperature, and holding time of each stage of the chromatographic program. The actual column temperature for each time step is calculated using a time step consistent with the mass spectrometry scan cycle. For example, if the initial temperature is 40 degrees Celsius and held for 2 minutes, then the temperature is increased to 280 degrees Celsius at a rate of 5 degrees Celsius per minute and held for 10 minutes, with a mass spectrometry scan interval of 0.2 seconds, the temperature for each time step is calculated sequentially, ultimately forming a one-dimensional temperature sequence with the same length as the sequence data.

[0055] Finally, the molecular embedding vectors corresponding to the organic molecular diagram structures are dimensionally aligned with the temperature sequences. The 64-dimensional molecular embedding vectors are broadcast to the same length as the temperature sequences, and then concatenated along the feature dimensions to form a fused feature sequence, which is then input into a single-layer gated recurrent unit network (GRU). The hidden layer dimension of this GRU is set to 64. The training dataset uses the molecular structures and corresponding standard retention times of organic compounds from the NIST standard mass spectrometry library. The training labels are the standard retention times of each organic compound. The mean squared error loss function is used during training, and the Adam optimizer is used. The initial learning rate is set to 0.0001, and the total number of training epochs is set to 30. During inference, the GRU first outputs the predicted retention time of each candidate organic compound, calculates the absolute deviation between the predicted retention time and the measured peak time, and converts the deviation into a matching probability between 0 and 1. The smaller the deviation, the higher the matching probability. Candidate organic compounds with a matching probability greater than 0.5 are sorted from high to low probability to generate a candidate organic compound list.

[0056] By constructing a molecular graph structure that encodes physicochemical properties and combining it with chromatographic temperature sequence features, and using a gated cyclic unit to fuse the two types of information for retention time prediction and matching degree calculation, the accurate screening of candidate organic compounds was achieved, effectively improving the accuracy and efficiency of subsequent qualitative identification.

[0057] Furthermore, the method provided in this application embodiment includes:

[0058] The organic molecular graph structure uses atoms as nodes and chemical bonds as edges, and aggregates the features of neighboring atoms through graph convolution layers. The organic molecular graph structure contains molecular embedding vectors, which encode key physicochemical properties.

[0059] Specifically, the aforementioned steps have completed the construction of the organic molecular graph structure. This structure uses atoms as nodes and chemical bonds as edges, aggregating features of neighboring atoms through graph convolutional layers. The example above uses a single-layer graph convolutional network with a hidden layer dimension of 64. Each node integrates feature information from its direct neighboring nodes, ultimately generating a molecular embedding vector. This molecular embedding vector encodes key physicochemical properties of the organic compound, such as functional group type, topological polar surface area, and lipid-water partition coefficient. These properties are directly related to the retention behavior of organic compounds in the chromatographic column, providing richer feature support for predicting the retention time of subsequent gated cycling units and further improving the accuracy of candidate organic compound screening.

[0060] Furthermore, the method provided in this application embodiment includes:

[0061] The candidate organic compound list is traversed to determine the first test group, wherein the first test group is obtained by splicing the first candidate organic compound with the temperature sequence; for the first test group, the theoretical peak time is predicted; and a first confidence level is generated by comparing the measured peak time with the theoretical peak time.

[0062] Specifically, a lightweight theoretical peak time predictor is constructed. This predictor employs a structure combining a single-layer gated recurrent unit with a linear output layer. The hidden layer dimension of the gated recurrent unit is set to 64, and the output dimension of the linear output layer is 1. The input to the predictor is a concatenated fusion feature sequence, which is obtained by broadcasting the 64-dimensional molecular embedding vector of the organic compound to a length equal to that of the temperature sequence and then concatenating it with the temperature sequence in the feature dimension. The output of the predictor is the theoretical peak time of the organic compound at the corresponding chromatographic temperature program.

[0063] A lightweight theoretical peak time predictor was pre-trained offline. The training dataset consisted of molecular structure data and corresponding standard retention time data for 50,000 common environmental organic pollutants from the NIST standard mass spectrometry library. All standard retention times in the training dataset corresponded to nonpolar DB-5 capillary columns. During training, the molecular embedding vector of each organic compound was concatenated with the temperature sequence encoded by the corresponding standard chromatographic temperature program as input, and the training label was the standard retention time of that organic compound. Batch gradient descent was used for training, with a batch size of 32. The Adam optimizer was used, with an initial learning rate of 0.0001 and a mean squared error loss function. The total number of training epochs was 30, with the learning rate decreasing to 0.5 times every 5 epochs. An independent validation set was also included. Training was terminated early if the validation set loss did not decrease for three consecutive epochs. After training, the predictor's weight parameters were saved.

[0064] In actual detection, the first candidate organic compound is selected from the candidate organic compound list, for example, the first one after sorting by matching degree. The molecular embedding vector of this candidate organic compound is concatenated with the temperature sequence of the current measured sample to generate the first test group. The first test group is input into the trained predictor, which outputs the theoretical peak time of the candidate organic compound under the current chromatographic conditions. Then, the first confidence level is calculated: first, the absolute deviation (in minutes) between the theoretical peak time and the measured peak time is calculated, and the absolute deviation is converted into a confidence value between 0 and 1 using an exponential decay function, for example, confidence level = exp(-0.1 × absolute deviation). The smaller the absolute deviation, the higher the confidence level, thus obtaining the first confidence level.

[0065] Finally, the candidate organic compound list is traversed to generate the first to Nth confidence levels for each candidate organic compound, and the candidate organic compound with the highest confidence level is selected as the target organic compound. The specific implementation process of this step will be described in detail below.

[0066] By constructing and training a lightweight peak elution time predictor, the theoretical peak elution time is accurately predicted by combining molecular structure characteristics and chromatographic temperature program. Then, the confidence level is obtained by time deviation conversion, which realizes the quantitative evaluation of candidate organic compounds and provides an objective basis for the accurate identification of subsequent target organic compounds.

[0067] Furthermore, the method provided in this application embodiment includes:

[0068] The Nth test group is determined based on the list of candidate organic compounds, and the Nth confidence level is generated, where N is the total number of candidate organic compounds. By comparing the first confidence level up to the Nth confidence level, the candidate organic compound corresponding to the highest confidence level is selected as the target organic compound.

[0069] In one embodiment, the candidate organic compound list contains N candidate organic compounds (N is the total number of candidate organic compounds), corresponding to the generation of the first test group to the Nth test group, and the first confidence level to the Nth confidence level are obtained in sequence. The method of test group construction, theoretical peak time prediction and confidence level calculation is consistent with the aforementioned process of generating the first confidence level.

[0070] Next, a maximum value filtering method based on traversal comparison is adopted. The values ​​of the first confidence level to the Nth confidence level are read sequentially, and each confidence level value is compared with the others to identify the highest confidence level. Then, a one-to-one correspondence between confidence levels and candidate organic compounds is established to match the candidate organic compounds corresponding to the highest confidence level.

[0071] Furthermore, if multiple candidate organic compounds correspond to the same maximum confidence level, the candidate organic compound with the higher mass spectrometry matching degree is selected first. The mass spectrometry matching degree is the similarity between the measured mass spectrum of the target pure component in the sample and the standard mass spectrum of the corresponding candidate organic compound in the NIST standard mass spectrometry library. It is calculated by comparing the mass-to-charge ratio and abundance of the measured mass spectrometry fragment peaks with the corresponding parameters of the standard mass spectrum. The value ranges from 0 to 1. The larger the value, the higher the matching degree. Finally, the selected candidate organic compound is determined as the target organic compound, and the qualitative identification of the unknown organic compound is completed.

[0072] By comparing and screening the confidence levels of all candidate organic compounds in batches, the target organic compound with the highest matching degree is accurately identified, the judgment logic when multiple candidates have the highest confidence level is improved, and the accuracy and efficiency of identifying unknown organic pollutants in gas chromatography-mass spectrometry data are enhanced.

[0073] Furthermore, the method provided in this application embodiment includes:

[0074] After locking onto the target organic compound, the optimal characteristic ion information and collision energy parameters are retrieved by indexing. The optimal characteristic ion information includes the mass ratio of the parent ion, daughter ion, and characteristic fragment. The optimal characteristic ion information and collision energy parameters are sent to the scanning control module of the mass spectrometer to perform targeted acquisition under mode switching to obtain targeted detection data. The mode switching is from full scan mode to multiple reaction monitoring mode.

[0075] Optionally, an environmental organic pollutant parameter database can be pre-established. This database stores the optimal characteristic ion information and collision energy parameters for all target organic compounds. The optimal characteristic ion information includes the parent ion mass-to-charge ratio, daughter ion mass-to-charge ratio, and characteristic fragment mass ratio. The collision energy parameter is the optimal collision voltage value for the corresponding characteristic ion pair. Each organic compound uses its CAS number as a unique index key. After identifying the target organic compound, its CAS number is extracted as a search index. A precise matching search is performed in the environmental organic pollutant parameter database to retrieve all the optimal characteristic ion information and collision energy parameters corresponding to the target organic compound. The aforementioned parent ion mass-to-charge ratio, daughter ion mass-to-charge ratio, characteristic fragment mass ratio, and optimal collision voltage value can be obtained using mass spectrometry parameter optimization methods in this field (e.g., parent ion scanning and daughter ion scanning via standard injection), or they can be obtained from publicly available mass spectrometry libraries or literature.

[0076] Next, the optimal characteristic ion information and collision energy parameters obtained are sent to the mass spectrometer's scanning control module via the standard communication protocol. After receiving the parameters, the scanning control module automatically parses the parameter content and generates a scanning task list for multi-reaction monitoring mode. The scanning task list includes the parent ion mass-to-charge ratio, daughter ion mass-to-charge ratio, collision energy, and residence time for each characteristic ion pair. The residence time is uniformly set to 50 milliseconds.

[0077] The scanning control module performs adaptive mode switching based on the measured peak time of the target organic compound. 30 seconds before the expected peak time of the target organic compound, the mass spectrometer is switched from full scan mode to multi-reaction monitoring mode. The peak end time of the target organic compound is set to the measured peak time plus 1 minute. After the peak ends, it automatically switches back to full scan mode and performs continuous scanning and acquisition according to the generated scanning task list to fully acquire the characteristic ion pair signals of the target organic compound and obtain targeted detection data.

[0078] By establishing a parameter database in advance to quickly retrieve the optimal detection parameters, and combining it with adaptive mode switching based on peak elution time, high-sensitivity targeted acquisition of target organic compounds is achieved without affecting the full-scan qualitative analysis, significantly improving the detection accuracy of low-concentration organic pollutants.

[0079] Furthermore, the method provided in this application embodiment includes:

[0080] The targeted detection data is identified, and a concentration-response standard curve is established using the peak area of ​​the characteristic ion pair of the target organic compound as the response value. The measured peak area is obtained and substituted into the standard curve to calculate the organic compound concentration. Qualitative confirmation is performed based on the deviation of the ion abundance ratio from the standard, and a detection report is generated. The detection report includes the organic compound name, retention time, organic compound concentration, and ion comparison results.

[0081] In one embodiment, a series of standard working solutions for the target organic compound are prepared. For example, hexane is used as the solvent. Five concentration gradient standard solutions with concentrations of 0.1 μg / L, 0.5 μg / L, 1 μg / L, 5 μg / L, and 10 μg / L are prepared sequentially, with each concentration gradient standard solution prepared in triplicate. All standard working solutions are injected and analyzed under the same chromatographic and mass spectrometric conditions as the analyte sample, with an injection volume of 1 μL. Multiple reaction monitoring (MRM) data for each concentration gradient standard solution are collected sequentially. The peak areas of the characteristic ion pairs corresponding to the target organic compound are extracted, and the average peak area of ​​the three parallel samples at each concentration gradient is calculated.

[0082] Using the concentration of the standard working solution as the x-axis and the average peak area of ​​the corresponding concentration gradient as the y-axis, a linear regression fitting using the least squares method is performed to establish an external standard method concentration-response standard curve. The linear equation and correlation coefficient of the standard curve are obtained, with a requirement that the correlation coefficient be no less than 0.995. If the correlation coefficient does not meet the requirement, the standard solution is prepared again and the fitting is performed. If the internal standard method is used for quantitative analysis, the same concentration of deuterated internal standard is added to all standard working solutions and the sample to be tested. The peak areas of the characteristic ion pairs of the target organic compound and the internal standard are extracted, and the ratio of the peak area of ​​the target compound to the peak area of ​​the internal standard is calculated. Using the concentration of the standard working solution as the x-axis and the ratio of the average peak area of ​​the corresponding concentration gradient as the y-axis, a linear regression fitting using the least squares method is performed to establish an internal standard method concentration-corrected peak area standard curve.

[0083] The measured peak areas of characteristic ion pairs of the target organic compound in the targeted detection data are read. If the external standard method is used, the measured peak areas are substituted into the linear equation of the established concentration-response standard curve. If the internal standard method is used, the ratio of the peak areas of the measured target compound to the internal standard is substituted into the linear equation of the concentration-correction peak area standard curve to calculate the original concentration of the target organic compound in the sample. If the sample was diluted during pretreatment, the original concentration is multiplied by the dilution factor to obtain the final actual concentration. The ion abundance ratio of the quantitative ion peak area to the qualitative ion peak area of ​​the target organic compound in the targeted detection data is calculated. The measured ion abundance ratio is compared with the ion abundance ratio of the same concentration standard. If the relative deviation between the two does not exceed 20%, the qualitative confirmation is deemed qualified. The concentration gradient settings, injection volume, correlation coefficient requirements, and ion abundance ratio deviation thresholds in the above embodiments are illustrative examples and can be adjusted according to the response characteristics of the target organic compound and the detection standard requirements in actual applications.

[0084] Finally, the names of the target organic compounds, the measured retention times, and the calculated organic compound concentration and ion abundance ratio are integrated to generate a standardized test report.

[0085] Accurate concentration-response standard curves were established using both external and internal standard methods. Combined with clear rules for qualitative confirmation based on ion abundance ratios, accurate quantification and final confirmation of the target organic compounds were achieved, ensuring the reliability and traceability of the detection results.

[0086] In summary, the gas chromatography-mass spectrometry data analysis method provided in this application has the following technical effects:

[0087] This application reconstructs three-dimensional chromatographic mass spectrometry data into sequence data, extracts features using a lightweight self-attention network, associates fragment ions and separates pure components, and combines a pre-trained model with parameter adjustment to accurately separate co-eluting components. This makes the detection results of unknown organic compounds in environmental samples more accurate and reliable, achieving a fully automated, high-precision closed-loop analysis from raw detection data to the confirmation and quantification of target substances, reducing the risk of human intervention and misjudgment.

[0088] Example 2, as Figure 2 As shown, based on the same inventive concept as in Embodiment 1 above, this application provides a gas chromatography-mass spectrometry data analysis system, the system comprising:

[0089] The three-dimensional chromatography-mass spectrometry data acquisition module 1 is used to control the mass spectrometer to detect the target in full scan mode and obtain three-dimensional chromatography-mass spectrometry data.

[0090] Pure component mass spectrum acquisition module 2 is used to input the three-dimensional chromatographic mass spectrometry data into an attention network, capture first feature data based on local chromatographic peak morphology and second feature data based on global mass spectrometry fragment correlation, perform overlapping peak separation under cross constraints, and obtain pure component mass spectrum, wherein the first feature data provides time domain boundary and the second feature data provides frequency domain correlation.

[0091] Candidate organic compound list generation module 3 is used to construct the molecular map structure of organic compounds based on the mass spectrum of pure components, and at the same time read the chromatographic temperature data based on the measured peak elution time, and input it into the gated loop unit to generate a candidate organic compound list.

[0092] The target organic compound determination module 4 is used to traverse the candidate organic compound list, predict the theoretical peak time of each candidate organic compound, compare it with the measured peak time to determine the confidence level, and select the candidate with the highest confidence level as the target organic compound.

[0093] The detection report generation module 5 is used to lock onto the target organic matter, perform targeted parameter extraction and mode adaptive switching, and send the data to the scanning control module of the mass spectrometer for targeted acquisition and analysis to generate a detection report.

[0094] Furthermore, the pure component mass spectrum acquisition module 2 is used to perform the following steps:

[0095] Three-dimensional chromatographic mass spectrometry data are reconstructed into sequence data, wherein the three-dimensional chromatographic mass spectrometry data consists of time × mass-to-charge ratio × intensity, and each time step in the sequence data corresponds to a full mass spectrum vector of one scan cycle; the sequence data is input into an attention network, with local chromatographic peak morphology as the first attention target and global mass spectrometry fragment correlation as the second attention target, to capture first feature data and second feature data, wherein the local chromatographic peak morphology includes peak scale and symmetry, and the global mass spectrometry fragment correlation refers to the synchronous change of fragment ions of the same organic compound; based on the first feature data and the second feature data, the mass spectrum of the separated pure components is determined.

[0096] Furthermore, the pure component mass spectrum acquisition module 2 is used to perform the following steps:

[0097] Based on the second feature data, the mass-to-charge ratio channels are clustered to determine the first clustering result, wherein fragment ions with high attention weights are classified as the same potential organic compounds that change synchronously; using the second feature data as a time window constraint, cross-time window associations are removed from the first clustering result to obtain the second clustering result, wherein the time window constraint is used to limit the elution start and end range of each organic compound on the chromatography; according to the second clustering result, the fragment ion intensities belonging to the same organic compound are weighted and averaged within each time window to obtain the pure component mass spectrum of the time window.

[0098] Furthermore, the candidate organic compound list generation module 3 is used to perform the following steps:

[0099] Based on the mass spectrum of the pure components, an organic molecular map structure is constructed; chromatographic temperature data based on the measured peak elution time is read and encoded into a temperature sequence, wherein the chromatographic temperature data includes the initial temperature, heating rate, final temperature, and holding time; the organic molecular map structure and the temperature sequence are input into a gated loop unit to generate a candidate organic list.

[0100] Furthermore, the candidate organic compound list generation module 3 is used to perform the following steps:

[0101] The organic molecular graph structure uses atoms as nodes and chemical bonds as edges, and aggregates the features of neighboring atoms through graph convolution layers. The organic molecular graph structure contains molecular embedding vectors, which encode key physicochemical properties.

[0102] Furthermore, the target organic matter determination module 4 is used to perform the following steps:

[0103] The candidate organic compound list is traversed to determine the first test group, wherein the first test group is obtained by splicing the first candidate organic compound with the temperature sequence; for the first test group, the theoretical peak time is predicted; and a first confidence level is generated by comparing the measured peak time with the theoretical peak time.

[0104] Furthermore, the target organic matter determination module 4 is used to perform the following steps:

[0105] The Nth test group is determined based on the list of candidate organic compounds, and the Nth confidence level is generated, where N is the total number of candidate organic compounds. By comparing the first confidence level up to the Nth confidence level, the candidate organic compound corresponding to the highest confidence level is selected as the target organic compound.

[0106] Furthermore, the test report generation module 5 is used to perform the following steps:

[0107] After locking onto the target organic compound, the optimal characteristic ion information and collision energy parameters are retrieved by indexing. The optimal characteristic ion information includes the mass ratio of the parent ion, daughter ion, and characteristic fragment. The optimal characteristic ion information and collision energy parameters are sent to the scanning control module of the mass spectrometer to perform targeted acquisition under mode switching to obtain targeted detection data. The mode switching is from full scan mode to multiple reaction monitoring mode.

[0108] Furthermore, the test report generation module 5 is used to perform the following steps:

[0109] The targeted detection data is identified, and a concentration-response standard curve is established using the peak area of ​​the characteristic ion pair of the target organic compound as the response value. The measured peak area is obtained and substituted into the standard curve to calculate the organic compound concentration. Qualitative confirmation is performed based on the deviation of the ion abundance ratio from the standard, and a detection report is generated. The detection report includes the organic compound name, retention time, organic compound concentration, and ion comparison results.

[0110] The gas chromatography-mass spectrometry data analysis system provided in this embodiment of the invention can execute the gas chromatography-mass spectrometry data analysis method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method.

[0111] Although this application makes various references to certain modules in the system according to the embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy distinction between each other and are not used to limit the scope of protection of this invention.

[0112] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application. In some cases, the actions or steps described in this application can be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

Claims

1. A method of gas chromatography mass spectrometry data analysis, characterized by, The method includes: Three-dimensional chromatographic mass spectrometry data were obtained by controlling the mass spectrometer to detect the target in full scan mode. The three-dimensional chromatographic mass spectrometry data is input into an attention network to capture first feature data based on local chromatographic peak morphology and second feature data based on global mass spectrometry fragment correlation. Overlapping peak separation under cross-constraint is performed to obtain pure component mass spectra. The first feature data provides time domain boundaries and the second feature data provides frequency domain correlation. The molecular structure of organic compounds is constructed based on the mass spectra of pure components. At the same time, chromatographic temperature data based on the measured peak elution time is read and input into the gated cycle unit to generate a list of candidate organic compounds. The list of candidate organic compounds is traversed, the theoretical elution time of each candidate organic compound is predicted, and the confidence level is determined by comparing it with the measured elution time. The candidate with the highest confidence level is selected as the target organic compound. After identifying the target organic compound, the system performs targeted parameter extraction and adaptive mode switching, and sends the data to the mass spectrometer's scanning control module for targeted acquisition and analysis, generating a detection report.

2. A method of analysis of gas chromatography mass spectral data as claimed in claim 1 wherein, The mass spectra of the pure components were obtained, including: The three-dimensional chromatographic mass spectrometry data is reconstructed into sequence data, wherein the three-dimensional chromatographic mass spectrometry data consists of time × mass-to-charge ratio × intensity, and each time step in the sequence data corresponds to a full mass spectrum vector of one scan cycle. The sequence data is input into an attention network, with local chromatographic peak morphology as the first attention target and global mass spectrometry fragment association as the second attention target, to capture first feature data and second feature data. The local chromatographic peak morphology includes peak size and symmetry, and the global mass spectrometry fragment association refers to the synchronous change of fragment ions of the same organic compound. Based on the first feature data and the second feature data, the mass spectrum of the separated pure components is determined.

3. The gas chromatography-mass spectrometry data analysis method as described in claim 2, characterized in that, Based on the first feature data and the second feature data, determine the mass spectrum of the separated pure components, including: Based on the second feature data, the mass-to-charge ratio channels are clustered to determine the first clustering result, in which fragment ions with high attention weights are classified as the same potential organic matter that changes synchronously. Using the second feature data as a time window constraint, cross-time window associations are removed from the first clustering result to obtain the second clustering result, wherein the time window constraint is used to limit the elution start and end range of each organic compound on the chromatogram; Based on the second clustering results, the intensity of fragment ions belonging to the same organic compound is weighted and averaged within each time window to obtain the pure component mass spectrum of the time window.

4. The gas chromatography-mass spectrometry data analysis method as described in claim 1, characterized in that, Generate a list of candidate organic compounds, including: Based on the mass spectrum of the pure components, construct the molecular structure of the organic compounds. Read the chromatographic temperature data based on the measured peak elution time and encode it into a temperature sequence, wherein the chromatographic temperature data includes the initial temperature, heating rate, final temperature, and holding time; The molecular diagram of the organic compound and the temperature sequence are input into a gated loop unit to generate a list of candidate organic compounds.

5. The gas chromatography-mass spectrometry data analysis method as described in claim 4, characterized in that, The organic molecular graph structure uses atoms as nodes and chemical bonds as edges, and aggregates the features of neighboring atoms through graph convolution layers. The organic molecular graph structure contains molecular embedding vectors, which encode key physicochemical properties.

6. The gas chromatography-mass spectrometry data analysis method as described in claim 4, characterized in that, The theoretical elution times of each candidate organic compound are predicted, and the confidence levels are determined by comparing them with the measured elution times, including: Traverse the list of candidate organic compounds to determine the first test group, wherein the first test group is obtained by splicing the first candidate organic compound with the temperature sequence; For the first test group, predict the theoretical peak time; A first confidence level is generated by comparing the measured peak time with the theoretical peak time.

7. The gas chromatography-mass spectrometry data analysis method as described in claim 6, characterized in that, Determine the Nth test group based on the candidate organic compound list and generate the Nth confidence level, where N is the total number of candidate organic compounds; By comparing the first confidence level up to the Nth confidence level, the candidate organic compound corresponding to the highest confidence level is selected as the target organic compound.

8. The gas chromatography-mass spectrometry data analysis method as described in claim 1, characterized in that, Targeted collection includes: After locking onto the target organic compound, the optimal characteristic ion information and collision energy parameters are retrieved by indexing. The optimal characteristic ion information includes the mass ratio of the parent ion, daughter ion, and characteristic fragment. The optimal characteristic ion information and collision energy parameters are sent to the scanning control module of the mass spectrometer to perform targeted acquisition under mode switching and obtain targeted detection data. The mode switching is from full scan mode to multi-reaction monitoring mode.

9. The gas chromatography-mass spectrometry data analysis method as described in claim 8, characterized in that, Generate a test report, including: The target detection data is identified, and a concentration-response standard curve is established using the peak area of ​​the characteristic ion pairs of the target organic compound as the response value. The measured peak area is obtained and substituted into the standard curve to calculate the organic matter concentration. Qualitative confirmation is performed based on the deviation of the ion abundance ratio from the standard, and a test report is generated. The test report includes the name of the organic matter, retention time, organic matter concentration and ion comparison results.

10. A gas chromatography-mass spectrometry data analysis system, characterized in that, The system for implementing the gas chromatography-mass spectrometry data analysis method according to any one of claims 1-9, the system comprising: The three-dimensional chromatography-mass spectrometry data acquisition module is used to control the mass spectrometer to detect the target in full scan mode and obtain three-dimensional chromatography-mass spectrometry data. The pure component mass spectrum acquisition module is used to input the three-dimensional chromatographic mass spectrometry data into the attention network, capture the first feature data based on the local chromatographic peak morphology and the second feature data based on the global mass spectrometry fragment correlation, perform overlapping peak separation under cross constraints, and obtain the pure component mass spectrum. The first feature data provides the time domain boundary and the second feature data provides the frequency domain correlation. The candidate organic compound list generation module is used to construct the molecular map structure of organic compounds based on the mass spectrum of pure components, and at the same time read the chromatographic temperature data based on the measured peak elution time, and input it into the gated loop unit to generate the candidate organic compound list; The target organic compound determination module is used to traverse the candidate organic compound list, predict the theoretical peak time of each candidate organic compound, compare it with the measured peak time to determine the confidence level, and select the candidate with the highest confidence level as the target organic compound. The detection report generation module is used to extract targeted parameters and adaptively switch modes after identifying the target organic matter, and then send the data to the scanning control module of the mass spectrometer for targeted acquisition and analysis to generate a detection report.