An intelligent batch processing method, system, device and medium for detecting data
By adopting an automatic alignment method for GCMS data based on RT tolerance, the problems of low efficiency and poor accuracy in GCMS multi-sample data processing are solved, and efficient and accurate multi-sample data analysis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA TOBACCO HUNAN IND CORP
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-26
AI Technical Summary
GCMS detection suffers from low efficiency, poor accuracy, and limited functionality in processing multi-sample data, and is also plagued by retention time drift, data dispersion, and inter-sample variability.
An automatic alignment method for GCMS data based on RT tolerance is adopted. By extracting target data, establishing a tolerance matching algorithm to obtain data groups, and performing data integration and intelligent labeling, the automatic alignment and integration of multiple sample data is achieved.
It significantly improves the efficiency and quality of GCMS data processing, solves the core pain points of traditional processing methods, and achieves efficient and accurate multi-sample data analysis.
Smart Images

Figure CN122087474A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to an intelligent batch processing method, system, device and medium for detection data. Background Technology
[0002] Gas chromatography-mass spectrometry (GC-MS) is an important qualitative and quantitative tool in chemical analysis. In practical applications, after analyzing each sample, GC-MS generates an analytical data file containing the following key information: retention time (RT), hit name, quality of match, molecular weight, CAS number, and peak area. However, due to the following technical issues, direct comparison and analysis of data from multiple samples is difficult: (1) Retention time drift problem: Due to factors such as column aging, temperature fluctuation, and flow rate change, the retention time of the same substance in different samples varies by ±0.01-0.03 minutes (the main phenomenon), and some even longer. (2) Data Discreteness Problem: Each sample generates an independent file, with the same data format but discrete content; (3) Differences between samples: Different samples may detect different amounts of substances, and some substances may not be detected in some samples.
[0003] However, currently, operators manually open each sample file, search for and compare substances with the same retention time, manually copy and paste data into a summary table, and manually calculate and organize statistical information. Therefore, this method suffers from problems such as extremely low efficiency, high error rate, and limited functionality.
[0004] Therefore, how to solve the problems of low efficiency, poor accuracy and limited functionality in GCMS multi-sample data processing in the existing technology is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] To address the aforementioned issues, this invention provides an intelligent batch processing method, system, device, and medium for detection data. By employing an automatic GCMS data alignment method based on RT tolerance, it can significantly improve the efficiency and quality of GCMS data processing, thus resolving the core pain points of traditional processing methods.
[0006] The first objective of this invention is to provide an intelligent batch processing method for detection data; The technical solution provided by this invention is as follows: A smart batch processing method for detection data includes the following steps: Extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample; Data sets are obtained based on the target data and the tolerance matching algorithm; Data is integrated based on the data set to obtain information about the target substance; The target substance information is intelligently labeled, and the intelligently labeled target substance information is output.
[0007] Preferably, the step of extracting target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample specifically includes: By identifying the valid parsed data regions in all GCMS parsing files, the retention time (RT), substance name, quality of match, molecular weight, CAS number, and peak area (Area) are extracted. The target data includes the retention time RT, the substance name, the matching quality, the molecular weight, the CAS number, and the peak area Area.
[0008] Preferably, the step of obtaining the data set based on the target data and the tolerance matching algorithm specifically includes: Create an empty data set G={}; Iterate through all substances Mi (i=1,2,...,N, where N is the total number of substances) in all samples, and process each substance Mi according to the judgment conditions. The specific judgment process is as follows: The retention time RTi of substance Mi; In the current data set G, check if there exists a data set Gj (j=1,2,...,M, where M is the number of data sets in the current data set G) that satisfies the condition |RTj_ref-RTi|≤tolerance threshold T: If a Gj that meets the conditions exists, then add the substance Mi to the data set Gj and update the reference retention time RTj_ref of the data set Gj. If no Gj meets the conditions, create a new data group G_new, set the initial reference retention time RT_new_ref=RTi, and add the data group G_new to the data group G; The iterative judgment process continues until all substances have been processed to obtain a complete data set G.
[0009] Preferably, after obtaining the data set based on the target data and the tolerance matching algorithm, the method further includes: The median retention time RT_median is calculated based on the data set.
[0010] Preferably, the tolerance threshold T is specifically 0.01 minutes, 0.02 minutes, or 0.03 minutes.
[0011] Preferably, the step of integrating the data based on the data set to obtain target substance information specifically includes: Select the median RT value of all retention times within the data set; The material information with the highest quality within the data set is selected as the target material information.
[0012] Preferably, the step of intelligently tagging the target substance information and outputting the intelligently tagged target substance information specifically includes: Peak area marking, mass marking, and dual marking are performed based on the target substance information; The target visual data is generated by combining peak area marking, mass marking, and double-marked target material information and output.
[0013] The second objective of this invention is to provide an intelligent batch processing system for detection data; The technical solution provided by this invention is as follows: An intelligent batch processing system for detection data includes: an extraction module, an acquisition module, a data integration module, and a tag output module; The extraction module is used to extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample. The acquisition module is used to acquire a data set based on the target data and the tolerance matching algorithm; The data integration module is used to integrate data according to the data set in order to obtain target substance information; The tagging output module is used to perform intelligent tagging based on the target substance information and output the intelligently tagged target substance information.
[0014] The third objective of this invention is to provide an electronic device; The technical solution provided by this invention is as follows: An electronic device, comprising: At least one processor; and A memory communicatively connected to the at least one processor, the memory storing a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method steps of any one of the intelligent batch processing methods for detection data.
[0015] A fourth objective of this invention is to provide a computer-readable storage medium; The technical solution provided by this invention is as follows: A computer-readable storage medium for storing a computer program for causing a computer to perform the steps of any one of the intelligent batch processing methods for detecting data.
[0016] Compared with existing technologies, the present invention provides an intelligent batch processing method for detection data, comprising: extracting target data based on the analytical data of GC-MS samples; obtaining data groups based on the target data and a tolerance matching algorithm; integrating the data groups to obtain target substance information; intelligently labeling the target substance information; and outputting the intelligently labeled target substance information. This method, through an automatic GCMS data alignment method based on RT tolerance, can significantly improve the efficiency and quality of GCMS data processing, and solves the core pain points of traditional processing methods.
[0017] The present invention also provides an intelligent batch processing system for detection data. Since this system and the intelligent batch processing method for detection data solve the same technical problem and belong to the same technical concept, they should have the same beneficial effects, and will not be described in detail here. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 A flowchart of an intelligent batch processing method for detection data provided in one embodiment; Figure 2 A schematic diagram of a tolerance matching algorithm provided in one embodiment; Figure 3 A schematic diagram of the structure of an intelligent batch processing system for detection data provided in one embodiment; Figure 4 This is a schematic diagram of the structure of an electronic device provided in one embodiment. Detailed Implementation
[0020] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] like Figure 1 As shown, an embodiment of the present invention provides an intelligent batch processing method for detection data, comprising the following steps: S1. Extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample; S2. Obtain a data set based on the target data and the tolerance matching algorithm; S3. Integrate the data based on the data set to obtain information about the target substance; S4. Perform intelligent tagging based on the target substance information, and output the intelligently tagged target substance information.
[0022] In steps S1 to S4, based on the raw analytical data obtained from the gas chromatography-mass spectrometry (GC-MS) analysis of the samples, key data information related to target detection is accurately identified and extracted. The extracted target data is then processed and compared using a tolerance matching algorithm to filter and obtain data sets that meet preset conditions. These data sets are systematically organized and integrated, and further refined and confirmed through data integration technology to accurately obtain the required detailed information about the target substances. Based on the obtained target substance information, an intelligent labeling method is used to classify and label it, and finally, the intelligently labeled target substance information is visualized and output. This method, through an automatic GCMS data alignment method based on RT tolerance, can significantly improve the efficiency and quality of GCMS data processing, solving the core pain points of traditional processing methods. Furthermore, this method enables a complete process from data reading to result output, possessing a standardized, automated, and intelligent processing mode.
[0023] Preferably, the step of extracting target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample specifically includes: By identifying the valid parsed data regions in all GCMS parsing files, the retention time (RT), substance name, quality of match, molecular weight, CAS number, and peak area (Area) are extracted. The target data includes the retention time RT, the substance name, the matching quality, the molecular weight, the CAS number, and the peak area Area.
[0024] In practical applications, the system automatically scans all GCMS parsing files in a specified folder (which can be in .xls format), then identifies the valid data area in each file (skipping non-data rows such as headers) and extracts the following key information for each sample: retention time (RT), substance name, quality of match, molecular weight, CAS number, and peak area (Area). Simultaneously, a unique identifier (sample name) is assigned to each sample.
[0025] Preferably, the step of obtaining the data set based on the target data and the tolerance matching algorithm specifically includes: Create an empty data set G={}; Iterate through all substances Mi (i=1,2,...,N, where N is the total number of substances) in all samples, and process each substance Mi according to the judgment conditions. The specific judgment process is as follows: The retention time RTi of substance Mi; In the current data set G, check if there exists a data set Gj (j=1,2,...,M, where M is the number of data sets in the current data set G) that satisfies the condition |RTj_ref-RTi|≤tolerance threshold T: If a Gj that meets the conditions exists, then add the substance Mi to the data set Gj and update the reference retention time RTj_ref of the data set Gj. If no Gj meets the conditions, create a new data group G_new, set the initial reference retention time RT_new_ref=RTi, and add the data group G_new to the data group G; The iterative judgment process continues until all substances have been processed to obtain a complete data set G.
[0026] In practical applications, a tolerance threshold T for the retention time RT is set. This tolerance threshold T is preferably T = 0.01, 0.02, or 0.03 minutes. Substances from different samples whose retention time differences do not exceed the tolerance threshold T are matched into the same data set. Figure 2 (a) shows the RT1 distribution, where samples A, B, and C are all green, indicating a successful match; Figure 2 (b) shows the RT2 distribution, where samples A and B are both green, indicating a successful match, while sample C is red, indicating that it exceeds the lower limit of the tolerance threshold T.
[0027] Specifically, an empty data set G is established, all substances Mi (i=1,2,...,N, where N is the total number of substances) in all samples are traversed, and each substance Mi is processed according to the judgment conditions. The specific judgment process is as follows: Step 1: Initialization Phase Create an empty data set G={}, at which point set G contains no data sets; Step 2: Iterative Processing Phase For the first substance M1 in the first sample, since the set G is empty and no data group can be found, the first data group G1 is created and the first substance M1 is added to the set G. The set G is then represented as: G={G1}. Step 3: Continuous Iteration For the second substance M2 (whether it is the same sample or a different sample): Now that set G contains G1 (and is no longer empty), search for a matching data group in G={G1}: If it matches, add it to G1; if it doesn't match, create G2 and add it to G. Step 4: Algorithm complete Once all substances from all samples have been processed, set G contains all data grouped according to tolerance matching rules.
[0028] The core logic in this embodiment is that the "empty" state of set G exists only at the beginning of the algorithm, and G dynamically grows as substances are processed one by one. This is a standard "initialization → iterative filling" algorithm pattern. The core technical challenge of GCMS data analysis is solved by an automatic GCMS data alignment method based on RT tolerance, achieving intelligent matching of multi-sample data.
[0029] Preferably, after obtaining the data set based on the target data and the tolerance matching algorithm, the method further includes: The median retention time RT_median is calculated based on the data set.
[0030] In practical applications, the median is the value located at the midpoint of a data set arranged in ascending (or descending) order. It is a typical positional mean and is unaffected by extreme values. The median is primarily used for ordinal data, but can also be used for numerical data; however, it cannot be used for categorical data.
[0031] If the sequence is odd, the median is equal to the nth... The number of elements; if the sequence is even, the median is equal to the first element. and The median is the average of the number of data points. For a given set of data, the median is unique.
[0032] The following examples illustrate the specific steps for calculating the median RT: For a dataset Gj, suppose it contains the following RT values (in min): Sample A: 12.345; Sample B: 12.348; Sample C: 12.350; Sample D: 12.352; The calculation process is as follows: 1. Based on the RT values of the above samples, sort them as follows: 12.345, 12.348, 12.350, 12.352; 2. Calculate the median: (12.348 + 12.350) / 2 = 12.349 3. Keep 4 decimal places: RT_median_k=12.3490min.
[0033] Preferably, the step of integrating the data based on the data set to obtain target substance information specifically includes: Select the median RT value of all retention times within the data set; The material information with the highest quality within the data set is selected as the target material information.
[0034] In practical applications, for each data set Gk, it's important to note that this data set Gk differs from the previous data set Gj. Data set Gj is based on retention time, while this data set Gk involves another variable, namely the matching quality. Therefore, it is denoted by Gk. Retention time: Take the median of all RT values in the group, RT_median_k, and retain 4 decimal places; Material information: Select the material information with the highest quality within the group as representative; Substance name = argmax(Quality); Molecular weight = the molecular weight corresponding to the highest quality; CAS number = CAS number corresponding to the highest quality; Sample data: Record the peak area (Area) of each sample in the group; undetected areas are marked as "ND".
[0035] Preferably, the step of intelligently tagging the target substance information and outputting the intelligently tagged target substance information specifically includes: Peak area marking, mass marking, and dual marking are performed based on the target substance information; The target visual data is generated by combining peak area marking, mass marking, and double-marked target material information and output.
[0036] In practical applications, each data set Gk is labeled using methods such as peak area labeling, quality labeling, and double labeling, specifically as follows: (1) Peak area marking: Find the sample with the largest peak area S_max_area within the group, and mark the peak area value of S_max_area in red in the summary table.
[0037] (2) Quality marking: Find the sample with the highest matching quality within the group, S_max_quality, and mark the cell of S_max_quality with a light yellow background in the summary table. Also, record the actual retention time RT_actual of S_max_quality.
[0038] (3) Double marking: If a sample is both the sample with the largest peak area (S_max_area) and the sample with the highest matching quality (S_max_quality) in the group, it will be highlighted with red text and a light yellow background.
[0039] The labeled data is then used to generate a master summary table containing integrated information for all data groups, a quality traceability column recording the actual RT value of the highest-quality sample in each data group, and a processing log recording the processing procedure, statistical information, and error information. The results are then automatically saved by timestamp. This embodiment, by establishing a two-dimensional intelligent labeling system that considers both peak area and quality as key indicators, provides intuitive and visual analysis results.
[0040] The present invention will be further described in detail with reference to the following embodiments: Example 1: Batch processing of 33 GCMS samples as an example 1. Data preparation stage This batch processing experiment involved 33 parsing files from an Agilent gas chromatography-mass spectrometry (GCMS) system. All files were stored and transmitted in .xls spreadsheet format. Each file contained a relatively large amount of data, ranging from approximately 90 to 200 lines, covering chromatographic peak information of different substances. The data content mainly included key parameters such as retention time (RT), peak area, and qualitative information of substances.
[0041] 2. Parameter Configuration and Settings Regarding the matching tolerance, a high-precision mode was selected, with a time window set at ±0.01 minutes to improve the accuracy of material matching; In terms of output settings, the system will generate a summary table of data with undetected (ND) markers for easy subsequent analysis; At the same time, a two-dimensional labeling function has been enabled, which can highlight key data from different perspectives.
[0042] 3. Detailed processing procedure (1) First, run the processing system and specify the directory path containing 33 GCMS data files; (2) The system automatically scans the folder, reads all .xls format files completely, and accurately extracts the valid chromatographic data from them; (3) Using a high-precision tolerance matching algorithm, substances with similar retention times are classified and integrated to form 650 valid data groups; (4) Perform the following standardization operations for each data set: Calculate the median retention time of the substance in all samples, and round the result to four decimal places. Filter and record the qualitative information of the highest quality substances in this dataset; The sample number with the largest peak area is highlighted in red. The sample data with the highest quality score is highlighted against a light yellow background; Simultaneously record the actual retention time value corresponding to the highest quality sample; (5) The system automatically generates a summary table, which includes: 650 complete data records, corresponding to all identified substance groups; 35 columns of detailed data, including 5 columns of basic information, 33 columns of sample data, a tolerance column, and a dedicated column for quality RT; (6) Synchronously generate detailed processing log documents, which include: Document processing quantity statistics table; Summary of data group integration; Record any exceptions or errors that occur during the processing; A complete processing time statistics report.
[0043] 4. Processing Results and Performance Analysis Total processing time: 43.7 seconds, demonstrating high processing efficiency; The number of successfully integrated data sets was 650, demonstrating a good data integration effect. The system completed 1,300 automatic markings (including 650 peak area maximum markings and 650 highest quality markings); After manual random sampling verification, 100 key data points were checked, and the results showed that all data were accurate, with an accuracy rate of 100%.
[0044] Example 2: Comparative experiments with different tolerances: Experimental conditions: 33 samples in the same group, using: tolerance T=0.01min (high precision mode); Experiments were conducted with tolerance T = 0.02 min (standard mode) and tolerance T = 0.03 min (relaxed mode); the experimental results are shown in Table 1 below: Table 1 Experimental Results
[0045] Furthermore, through the sample analysis results batch processing summary table in Table 2 below, it can be concluded that the data summarized by this method not only greatly shortens the time compared to conventional manual summarization, but also has high data reliability. It can also intuitively show which sample has the largest peak area and the sample with the highest index matching degree and its retention time.
[0046] Manual summarization often requires opening each sample's analytical data file one by one, finding each analytical time point, and copying the corresponding data of each sample to the summary table. Although it can eventually achieve the goal, it is very inefficient when summarizing more than 10 samples, especially for routine samples which typically have more than 30, or even more than 200, data points with retention times. Moreover, it is prone to errors such as data not being pasted into the correct location, making subsequent analysis often futile.
[0047] Table 2 Summary of Batch Processing Results of Sample Analysis Obtained Using This Method (Excerpt)
[0048] In the table above, underlined data is marked in red font in actual use; slanted data is marked with a light yellow background; data that is both underlined and slanted is highlighted using both red font and light yellow background.
[0049] like Figure 3 As shown in the figure, an intelligent batch processing system for detection data according to an embodiment of the present invention includes: an extraction module, an acquisition module, a data integration module, and a tag output module; The extraction module is used to extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample. The acquisition module is used to acquire a data set based on the target data and the tolerance matching algorithm; The data integration module is used to integrate data according to the data set in order to obtain target substance information; The tagging output module is used to perform intelligent tagging based on the target substance information and output the intelligently tagged target substance information.
[0050] In practical applications, the intelligent batch processing system for detection data includes an extraction module, an acquisition module, a data integration module, and a labeling output module. The acquisition module is connected to both the extraction and data integration modules; the labeling output module is connected to the data integration module. The extraction module extracts target data from the GC-MS sample analysis data and transmits the target data to the acquisition module. The acquisition module acquires data sets based on the target data and a tolerance matching algorithm and transmits the data sets to the data integration module. The data integration module integrates the data sets to obtain target substance information and transmits this information to the labeling output module. The labeling output module intelligently labels the target substance information and outputs the intelligently labeled target substance information. This system, through its extraction, acquisition, data integration, and labeling output modules, and based on an automatic GCMS data alignment method with RT tolerance, significantly improves the efficiency and quality of GCMS data processing, addressing the core pain points of traditional processing methods.
[0051] Furthermore, embodiments of this application also disclose an electronic device, Figure 4 This is a structural diagram of an electronic device according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0052] Figure 4 This is a schematic diagram of an electronic device provided in an embodiment of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the intelligent batch processing method for detection data disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0053] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create an intelligent batch processing channel for detection data between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0054] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222 and data 223, etc., and the storage method can be temporary storage or permanent storage.
[0055] The operating system 221 manages and controls the various hardware devices on the electronic device 20 and the computer program 222 to enable the processor 21 to perform operations and processing on the data 223 in the memory 22. The operating system 221 can be Windows Server, Netware, Unix, Linux, etc. The computer program 222, in addition to including a computer program capable of performing the intelligent batch processing method for detection data executed by the electronic device 20 as disclosed in any of the foregoing embodiments, may further include computer programs capable of performing other specific tasks. The data 223 may include data received by the intelligent batch processing device for detection data from external devices, as well as data collected by its own input / output interface 25.
[0056] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0057] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned intelligent batch processing method for detection data. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0058] It should be understood that the use of terms such as "method," "apparatus," "unit," and / or "module" in this application is merely to distinguish one method of different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0059] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements. An element defined by the phrase "comprising an..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.
[0060] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0061] If a flowchart is used in this application, it is used to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.
[0062] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An intelligent batch processing method for detection data, characterized in that, Includes the following steps: Extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample; Data sets are obtained based on the target data and the tolerance matching algorithm; Data is integrated based on the data set to obtain information about the target substance; The target substance information is intelligently labeled, and the intelligently labeled target substance information is output.
2. The intelligent batch processing method for detection data according to claim 1, characterized in that, The extraction of target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample specifically includes: By identifying the valid parsed data regions in all GCMS parsing files, the retention time (RT), substance name, quality of match, molecular weight, CAS number, and peak area (Area) are extracted. The target data includes the retention time RT, the substance name, the matching quality, the molecular weight, the CAS number, and the peak area Area.
3. The intelligent batch processing method for detection data according to claim 2, characterized in that, The step of obtaining the data set based on the target data and the tolerance matching algorithm specifically includes: Create an empty data set G={}; Iterate through all substances Mi (i=1,2,...,N, where N is the total number of substances) in all samples, and process each substance Mi according to the judgment conditions. The specific judgment process is as follows: The retention time RTi of substance Mi; In the current data set G, check if there exists a data set Gj (j=1,2,...,M, where M is the number of data sets in the current data set G) that satisfies the condition |RTj_ref-RTi|≤tolerance threshold T: If a Gj that meets the conditions exists, then add the substance Mi to the data set Gj and update the reference retention time RTj_ref of the data set Gj. If no Gj meets the conditions, create a new data group G_new, set the initial reference retention time RT_new_ref=RTi, and add the data group G_new to the data group G; The iterative judgment process continues until all substances have been processed to obtain a complete data set G.
4. The intelligent batch processing method for detection data according to claim 1, characterized in that, After obtaining the data set based on the target data and the tolerance matching algorithm, the process further includes: The median retention time RT_median is calculated based on the data set.
5. The intelligent batch processing method for detection data according to claim 3, characterized in that, The tolerance threshold T is specifically 0.01 minutes, 0.02 minutes, or 0.03 minutes.
6. The intelligent batch processing method for detection data according to claim 1, characterized in that, The step of integrating the data based on the data set to obtain target substance information specifically includes: Select the median RT value of all retention times within the data set; The material information with the highest quality within the data set is selected as the target material information.
7. The intelligent batch processing method for detection data according to claim 1, characterized in that, The step of intelligently tagging the target substance information and outputting the intelligently tagged target substance information specifically includes: Peak area marking, mass marking, and dual marking are performed based on the target substance information; The target visual data is generated by combining peak area marking, mass marking, and double-marked target material information and output.
8. An intelligent batch processing system for detection data, characterized in that, include: The module includes an extraction module, an acquisition module, a data integration module, and a tag output module. The extraction module is used to extract target data based on the analytical data of the gas chromatography-mass spectrometry (GC-MS) sample. The acquisition module is used to acquire a data set based on the target data and the tolerance matching algorithm; The data integration module is used to integrate data according to the data set in order to obtain target substance information; The tagging output module is used to perform intelligent tagging based on the target substance information and output the intelligently tagged target substance information.
9. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor, the memory storing a computer program executable by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium is used to store a computer program that causes a computer to perform the method described in any one of claims 1-7.