Method for determining organic pollutants in alkylated greasy dirt water based on comprehensive two-dimensional gas chromatography-mass spectrometry

By applying all-two-dimensional gas chromatography mass spectrometry in wastewater combined with internal standard method, multiple linear regression model optimization, data smoothing treatment and support vector machine secondary verification, the uncertainty and accuracy of qualitative quantitative analysis of organic pollutants in wastewater are solved, especially in the quantitative analysis of low-content components, higher accuracy of analysis results is achieved, providing important technical support for sewage treatment and environmental monitoring.

CN120177656AInactive Publication Date: 2025-06-20NINGXIA RUIKE XINYUAN CHEM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364258.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of uncertainty and accuracy in the qualitative quantitative analysis of organic pollutants in wastewater, especially the inaccurate quantitative results caused by the differences in response factors of different compounds, and the difficulty in quantitative analysis of low-content components in complex wastewater samples.

Method used

Qualitative and quantitative analysis of wastewater samples was performed using the methods of full two-dimensional gas chromatography mass spectrometry combined with internal standard method, multiple linear regression model optimization, data smoothing processing and secondary verification of support vector machines. The accuracy and reliability of the analysis results are ensured by comparing with the standard compound database, calculating response factors, correcting peak area, optimizing mass fractions, smoothing low-content component data, and secondary verification of preliminary qualitative results.

Benefits of technology

It improves the accuracy and reliability of qualitative quantitative analysis of complex components in sewage, especially in the quantitative analysis of low-content components, which enhances the accuracy of the analysis results and provides important technical support for sewage treatment and environmental monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120177656A_ABST
    Figure CN120177656A_ABST
Patent Text Reader

Abstract

The invention provides a method for determining organic pollutants in alkylated greasy dirt water based on comprehensive two-dimensional gas chromatography-mass spectrometry. The method comprises the following steps: acquiring a sewage sample, and carrying out preliminary analysis to obtain two-dimensional retention time data and mass spectrogram data of each component; comparing the two-dimensional retention time data and the mass spectrogram data with a standard compound database to obtain a preliminary qualitative result; calculating the peak area of each component in the preliminary qualitative result, and introducing a standard substance through an internal standard method to obtain a response factor corresponding to each component; correcting the peak area of each component according to the response factor corresponding to each component to obtain corrected peak area data of each component; performing multiple linear regression model optimization on the mass fractions of the components to obtain optimized mass fractions of the components; performing secondary verification on the preliminary qualitative result through a support vector machine to obtain a qualitative result after secondary verification; and outputting the qualitative result after the secondary verification and the mass fraction of each component after the smoothing processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection technology, and particularly to a method for determining organic pollutants in alkylated oil sewage based on comprehensive two-dimensional gas chromatography-mass spectrometry. Background Art

[0002] In the business scenario of fine chemical sewage treatment, the qualitative and quantitative analysis of organic pollutants is a key step. The current method relies on chromatography-mass spectrometry coupling technology. First, qualitative analysis is carried out according to the two-dimensional retention time and mass spectrometry diagram of the sample, and each chromatographic peak is compared with the standard spectrum diagram to identify each hydrocarbon compound in the sewage. However, in sewage samples of different batches, there may be significant differences in the peak areas of the same compound, which brings a certain degree of uncertainty to the qualitative analysis. Further quantitative analysis is based on calculating the mass fraction of each component based on the peak area. Specifically, by calculating the ratio of the peak area of a single component to the total peak area of all components, the mass fraction of this component in the sewage is determined. The problem is that this calculation method assumes that the response factors of all components are the same, while in fact, the response factors of different compounds may vary significantly, which leads to doubts about the accuracy of the quantitative results. In addition, whether high precision requirements are feasible in actual operation, and whether the errors introduced in the peak area measurement and calculation process will affect the accuracy of the final result are all technical issues worthy of in-depth exploration. Especially considering the complexity and diversity of the sewage components, how to ensure the accuracy and reliability of qualitative and quantitative analysis has become a core problem to be solved urgently. Summary of the Invention

[0003] The present invention provides a method for determining organic pollutants in alkylated oil sewage based on comprehensive two-dimensional gas chromatography-mass spectrometry, mainly including:

[0004] Obtain a sewage sample, and preliminarily analyze to obtain two-dimensional retention time data and mass spectrometry diagram data of each component;

[0005] Compare the two-dimensional retention time data and mass spectrometry diagram data with a standard compound database to obtain a preliminary qualitative result;

[0006] Calculate the peak area of each component in the preliminary qualitative result, introduce a standard product by the internal standard method, and obtain the response factor corresponding to each component;

[0007] According to the response factor corresponding to each component, correct the peak area of each component to obtain the corrected peak area data of each component;

[0008] According to the corrected peak area data of each component, calculate the mass fraction of each component, where the mass fraction of a single component = (the corrected peak area of the single component / the total of the corrected peak area data of each component) × 100%;

[0009] Optimize the mass fractions of the respective components using a multiple linear regression model to obtain the optimized mass fractions of the respective components;

[0010] If the mass fraction in the optimized mass fractions of the respective components is less than the threshold, perform data smoothing on the mass fraction less than the threshold to obtain the mass fractions of the respective components after smoothing;

[0011] Perform secondary verification on the preliminary qualitative result through a support vector machine to obtain the qualitative result after secondary verification;

[0012] Output the qualitative result after secondary verification and the mass fractions of the respective components after smoothing.

[0013] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0014] The present invention discloses a method for qualitative and quantitative analysis of components in sewage. Aiming at the problems of complex components and large content differences in sewage samples, which lead to difficulties in qualitative and quantitative analysis, firstly, the two-dimensional retention time data and mass spectrometry data of the sewage sample are compared with the standard compound database to obtain a preliminary qualitative result, and the internal standard method is introduced to calculate the response factors of each component to correct the peak area, and then the preliminary mass fractions of each component are calculated. Then, a multiple linear regression model is used to optimize the mass fraction, and data smoothing is performed on the mass fraction below the threshold to improve the quantitative accuracy of low-content components. Finally, a support vector machine is used to perform secondary verification on the preliminary qualitative result to ensure the reliability of the qualitative result. The present invention integrates various technical means such as data comparison, internal standard correction, model optimization, data smoothing and secondary verification, realizes accurate and reliable qualitative and quantitative analysis of complex components in sewage samples, especially improves the quantitative accuracy of low-content components, and provides important technical support for sewage treatment and environmental monitoring. Description of the Drawings

[0015] Figure 1 It is a flowchart of a method for determining organic pollutants in alkylated oil sewage based on comprehensive two-dimensional gas chromatography-mass spectrometry of the present invention. Detailed Embodiments

[0016] The technical solutions in the embodiments of the present invention will be clearly and detailedly described below with reference to the drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0017] Such as Figure 1 , a method for determining organic pollutants in alkylated oil sewage based on comprehensive two-dimensional gas chromatography-mass spectrometry in this embodiment may specifically include:

[0018] Step S101: Obtain a sewage sample and preliminarily analyze it to obtain two-dimensional retention time data and mass spectrometry data for each component.

[0019] Obtain two-dimensional chromatographic data of the sewage sample. According to a pre-established chromatographic peak identification algorithm, determine the initial time of each component in the chromatogram. Based on the initial time of the component and a preset time difference, determine the termination time of each component, obtain the component types and corresponding primary mass spectrometry data. According to the component types, match the corresponding secondary mass spectrometry data from a pre-established mass spectrometry database, calculate the similarity between the primary mass spectrometry data and the secondary mass spectrometry data. If the similarity is greater than a preset threshold, determine the component name according to the corresponding secondary mass spectrometry data. According to the component name, obtain the standard chromatogram of the component from the database, compare the actual chromatogram data of the component with the standard chromatogram, and determine whether there is distortion in the actual chromatogram data. If there is distortion, perform correction processing. According to the corrected chromatogram data, integrally calculate the peak area of each component, and combine it with a pre-established calibration curve to determine the concentration of each component. According to the concentration of each component and a pre-established concentration database, determine whether the concentration of each component in the sewage sample exceeds the standard. If it exceeds the standard, generate an alarm signal. According to the alarm signal and a pre-established sewage treatment plan library, obtain the corresponding sewage treatment plan, and execute the sewage treatment plan through the control system.

[0020] Step S102: Compare the two-dimensional retention time data and mass spectrometry data with a standard compound database to obtain a preliminary qualitative result.

[0021] After obtaining the two-dimensional retention time data and mass spectrometry data of the sample to be tested, preprocess the two-dimensional retention time data of the sample to be tested to remove background noise signals. According to the preprocessed two-dimensional retention time data, extract characteristic peak information to obtain the retention time and peak area of the characteristic peaks. Through the extracted characteristic peak information, combined with the fragment ion information in the mass spectrometry data, calculate the molecular mass of the compounds in each characteristic peak. According to the molecular mass information of each standard substance in the standard compound database, screen out the standard substances with the same molecular mass as each characteristic peak to obtain a preliminary screened set of standard substances. According to the retention time information of each standard substance in the standard compound database, remove the standard substances in the preliminary screened set of standard substances whose retention times are inconsistent with the retention times of the characteristic peaks to obtain a set of candidate standard substances. Obtain the mass spectrometry data of each standard substance in the set of candidate standard substances, and determine whether the similarity coefficient is greater than a preset threshold by calculating the similarity coefficient between the mass spectrometry data of the candidate standard substances and the mass spectrometry data of the sample to be tested. If it is greater, determine the candidate standard substance as the target substance to obtain a preliminary qualitative result. According to the isotope peak information of each target substance in the obtained preliminary qualitative result, use the isotope internal standard method to obtain the quantitative result of the target substance.

[0022] Specifically, assume that after a sample to be tested is analyzed by a gas chromatography-mass spectrometry (GC-MS) instrument, a set of two-dimensional retention time data and corresponding mass spectrometry data are obtained. First, baseline correction is performed on the two-dimensional retention time data. For example, the adaptive iteratively reweighted penalized least squares algorithm is used, with the number of iterations set to 10 and the penalty factor set to 1000 to remove baseline drift in the data. Then, wavelet transform is used for denoising. The db4 wavelet basis function is selected, the decomposition level is 5 layers, and the threshold is set to 3 times the standard deviation to remove high-frequency noise signals. After the preprocessing is completed, the continuous wavelet transform peak detection algorithm is adopted, with the wavelet scale range set to 1 to 100 and the signal-to-noise ratio threshold set to 5 to detect and extract all characteristic peaks. Next, in combination with the mass spectrometry data corresponding to the characteristic peaks, the molecular mass of the compound in the characteristic peak is calculated through the mass-to-charge ratio of the main fragment ions shown in the mass spectrometry. Standard substances with this molecular mass are retrieved in the standard compound database, and multiple standard substances are preliminarily screened. The retention times are further compared. Assume that the retention time of benzaldehyde in the standard compound database is 130 minutes, the retention time of acetophenone is 156 minutes, and the retention time of the characteristic peak is 134 minutes. Therefore, acetophenone is excluded, and several standard substances including benzaldehyde are used as the candidate standard substance set. The mass spectra of these candidate standard substances are obtained. Taking benzaldehyde as an example, the mass-to-charge ratio and relative abundance of the main ions in its standard mass spectrum are 77 (60%), 105 (100%), and 106 (8%). The similarity coefficient between it and the mass spectrum of the sample to be tested is calculated using the cosine similarity algorithm, and the similarity coefficient is 95. The threshold of the similarity coefficient is set to 9. Since 95 is greater than 9, benzaldehyde is determined as the target substance, and a preliminary qualitative result is obtained. Assume that the preliminary qualitative result contains two target substances, benzaldehyde and toluene. Isotope internal standard method is used for quantification. Benzaldehyde-d5 and toluene-d8 are selected as internal standards, and the concentration of the internal standards is 10 μg / mL. The peak area of the isotope peak 107 of benzaldehyde is 50000 obtained from the mass spectrometry data, and the peak area of the isotope peak 112 of its isotope internal standard benzaldehyde-d5 is 20000. Similarly, the peak area of the isotope peak 92 of toluene is 60000, and the peak area of the isotope peak 100 of its isotope internal standard toluene-d8 is 25000. According to the calculation formula of the isotope internal standard method, the concentrations of benzaldehyde and toluene are 25 μg / mL and 24 μg / mL respectively, thus obtaining the quantitative results of the target substances.

[0023] Step S103: Calculate the peak areas of the components in the preliminary qualitative result, introduce a standard product through the internal standard method, and obtain the response factors corresponding to the components.

[0024] According to the original chromatogram data file, use the automatic integration algorithm to detect the chromatographic peaks therein and record the start time of the chromatographic peaks. According to the start time of the chromatographic peaks, perform integration operations to obtain the peak areas obtained by integrating each chromatographic peak, and generate a list of peak areas. According to the known concentration of the reference substance and its corresponding peak area, establish the relationship between the concentration of the reference substance and the peak area, and obtain the calibration curve equation of the reference substance y = ax + b, where y represents the peak area, x represents the concentration, and a and b represent the slope and intercept of the calibration curve equation respectively. According to the retention time of the chromatographic peaks of the substance to be measured, match the retention time of the corresponding components in the reference substance to determine the components of the substance to be measured, and generate a list of components. According to the calibration curve of the reference substance and the peak area of the substance to be measured, calculate the calibrated peak areas of each component, and generate a list of calibrated peak areas. According to the calibrated peak areas of each component and the calibration curve of the reference substance, calculate the concentrations of each component, and generate a list of component concentrations. According to the list of component concentrations and the concentration of the internal standard substance, calculate the relative response factors of each component through the normalization algorithm, and generate a list of response factors.

[0025] Step S104: Correct the peak areas of the components according to the response factors corresponding to the components, and obtain the peak area data after correction for each component.

[0026] Read the data table of the response factors corresponding to each component and the data table of the original peak areas of each component, and determine the set of components to be corrected. According to the determined set of components and the chromatographic conditions, obtain the original peak areas of each component from the data table of the original peak areas, and obtain the set of peak areas to be corrected. If there is a missing component original peak area data in the set of peak areas to be corrected, perform data integrity judgment, and use the cubic spline interpolation algorithm to predict the missing values to obtain the predicted peak areas of each component. According to the response factors and the original peak areas or predicted peak areas of each component, perform peak area correction calculations to obtain the preliminary corrected peak areas corresponding to each component. For the preliminary corrected peak areas of each component, combined with the original peak widths of each component, use the standard deviation normalization algorithm to obtain the standard corrected peak areas of each component. According to the standard corrected peak areas, combined with the data interface type, transmit the standard corrected peak areas of each component to the data receiving end through the data interface, and obtain the data receiving status feedback by the receiving end. If the data receiving status feedback by the receiving end is successful reception, then associate and store the standard corrected peak areas of each component and the corresponding component names to determine the final correction result.

[0027] Step S105: Calculate the mass fractions of the components according to the peak area data after correction for each component, where the mass fraction of a single component = (the peak area after correction of the single component / the sum of the peak area data after correction of each component) × 100%.

[0028] That is: after qualitative analysis of the chromatographic peaks, calculate the mass fractions of the monomeric compounds and component compounds in the sample according to the following formula:

[0029]

[0030] where: ωi—the mass fraction of a certain monomer or a certain component in the sample, % (m / m);

[0031] Ai—the peak area value of a certain monomer or a certain component in the sample;

[0032] ∑Ai—the sum of the peak area values of all components in the sample.

[0033] Obtain the chromatographic peak area data of each component to obtain a chromatographic peak area data set. According to the chromatographic peak area data set, perform correction processing on the chromatographic peak area data of each component to obtain the corrected peak area data of each component. Determine whether the corrected peak area data of each component is greater than a preset threshold. If it is greater than the preset threshold, determine that the corrected peak area data of each component is valid data. Calculate the sum of the corrected peak area data of all components. By accumulating the corrected peak area data of each component, determine the sum of the corrected peak area data of all components. Calculate the mass fraction of a single component. Divide the corrected peak area of a single component by the sum of the corrected peak area data of all components to obtain the calculation result of the mass fraction of a single component. According to the calculation result of the mass fraction of a single component, perform percentage conversion. Multiply the calculation result of the mass fraction of a single component by one hundred percent to determine the percentage value of the mass fraction of a single component. Collect the component concentration data during the chromatographic separation process. Combine the obtained percentage values of the mass fractions of each component to construct a set of corresponding relationships between the mass fractions of each component and the component concentrations. Through the gradient boosting decision tree of the machine learning algorithm, train the set of corresponding relationships between the mass fractions of each component and the component concentrations to obtain a corresponding relationship model between the mass fractions of each component and the component concentrations. Collect the chromatographic data of the sample to be tested. Input the corresponding relationship model between the mass fractions of each component and the component concentrations, and use the support vector machine algorithm or the K-nearest neighbor algorithm in combination with the corresponding relationship model between the mass fractions of each component and the component concentrations to obtain the predicted values of the mass fractions of each component in the sample to be tested.

[0034] Step S106, optimize the mass fractions of the components by using a multiple linear regression model to obtain the optimized mass fractions of the components.

[0035] Collect the initial data of each component and input it as the training data of the multiple linear regression model. According to the input of the training data, perform data fitting to obtain the initial multiple linear regression model. Set the objective function, which incorporates the component proportion information and aims to minimize the difference between the predicted value and the target data as the optimization objective. Adjust the model parameters through an iterative algorithm to obtain the optimized model parameters. Take the mass fractions of the components of the sample to be measured as the input of the optimized model to obtain the predicted values of the mass fractions of each component after regression analysis. Compare with the pre-established mass fraction threshold to determine whether the predicted value deviates from the threshold range. If the predicted value deviates from the threshold range, re-collect the data. Based on the re-collected data, construct a multiple linear regression model to determine the final mass fractions of each component.

[0036] Specifically, assume we have a mixture sample containing three components (A, B, C) and we need to determine the mass fractions of each component. First, we prepare a series of standard samples with known mass fractions. For example, in the first standard sample, the mass fractions of A, B, and C are 10%, 20%, and 30% respectively; in the second standard sample, the mass fractions of A, B, and C are 15%, 25%, and 35% respectively; in the third standard sample, the mass fractions of A, B, and C are 20%, 30%, and 40% respectively, and so on. We prepare 10 standard samples in this way. Use an analytical instrument (such as a spectrometer) to measure these standard samples to obtain the corresponding instrument response signals. For example, the spectral data of the first standard sample is (2, 5, 8), the spectral data of the second standard sample is (8, 2, 5), the spectral data of the third standard sample is (4, 9, 2), and so on. Take these spectral data as independent variables and the corresponding mass fractions as dependent variables to construct a multiple linear regression model: Y = β0 + β1X1 + β2X2 + β3X3, where Y is the mass fraction of the component, X1, X2, X3 are the spectral data, and β0, β1, β2, β3 are the model parameters to be estimated. Perform data fitting through the least squares method to obtain the initial model parameters. For example, β0 = 1, β1 = 5, β2 = 8, β3 = 2. Set the objective function, for example: Objective function = Σ(Predicted mass fraction - Actual mass fraction) 2 + λΣ(Sum of mass fractions of each component - 1) 2, where λ is the weight coefficient, for example, taking the value of 1. The first part of this objective function represents the difference between the predicted value and the actual value. The second part utilizes the prior information that the sum of the mass fractions of each component should be 1 (i.e., 100%). The model parameters are adjusted by iterative algorithms such as the gradient descent method to minimize the objective function. For example, after 100 iterations, the optimized model parameters are obtained: β0 = 05, β1 = 6, β2 = 9, β3 = 3. Now there is a sample to be measured, and its spectral data is (0, 5, 8). Inputting these data into the optimized model, the predicted mass fractions of each component are obtained: 18% for A, 28% for B, and 42% for C. We have preset the threshold ranges of the mass fractions of each component. For example, for A it is 15% - 25%, for B it is 20% - 35%, and for C it is 30% - 50%. Since the predicted values are all within the threshold ranges, we consider the prediction results to be reliable. If the predicted value exceeds the threshold range, for example, the predicted mass fraction of A is 5%, which is significantly lower than the set threshold range. At this time, the system will automatically mark this sample and trigger the data re - acquisition process, perform repeated measurements on this sample to obtain new spectral data. Then, using the new spectral data, reconstruct the multiple linear regression model and optimize the parameters to obtain new predicted values. If the new predicted values still exceed the threshold range, the system can further increase the number of repeated measurements or adopt other correction methods until reliable prediction results are obtained, and finally determine the mass fractions of each component.

[0037] Step S107, if there is a mass fraction less than the threshold among the optimized mass fractions of each component, perform data smoothing on the mass fraction less than the threshold to obtain the mass fractions of each component after smoothing.

[0038] Obtain the optimized initial mass fractions of each component, calculate the sum of the mass fractions of each component through a pre-set algorithm to obtain a total mass fraction. According to the ratio of the optimized initial mass fraction of each component to the total mass fraction, determine whether the mass fraction of each component is lower than a pre-set threshold. If there is a component mass fraction lower than the preset threshold, record the component information corresponding to the component mass fraction. For the component mass fraction that is lower than the preset threshold, use the moving average algorithm to calculate the component mass fraction that is lower than the threshold to obtain the smoothed value of the component mass fraction. According to the calculation result of the moving average algorithm, use the data normalization algorithm to process the mass fractions of each component to obtain a new sum of the mass fractions of each component and get a new total mass fraction. According to the new mass fractions of each component after being processed by the data normalization algorithm, classify the mass fractions of each component through the support vector machine algorithm to obtain the classification results of the mass fractions of each component. According to the classification results of the support vector machine algorithm, perform dimensionality reduction processing through the principal component analysis algorithm to obtain the mass fraction data of each component after dimensionality reduction processing. According to the mass fraction data of each component after dimensionality reduction processing by the principal component analysis algorithm, perform predictive analysis through the logistic regression algorithm to obtain the predictive results of the mass fractions of each component.

[0039] Step S108, perform a secondary verification on the preliminary qualitative result through a support vector machine to obtain the qualitative result after secondary verification.

[0040] According to the preliminary qualitative result, construct a data sample set. According to the data sample set, obtain the feature space. According to the feature space, perform model training. If there are outliers in the data sample set, perform data preprocessing to obtain the preprocessed data sample set. If there are no outliers in the data sample set, no processing is required. According to the preprocessed data sample set, determine the classification model to obtain the initial classification model. Through the initial classification model, construct a hyperplane to obtain the classification decision boundary. According to the classification decision boundary, perform a secondary verification on the preliminary qualitative result to obtain the qualitative result after secondary verification.

[0041] Step S109, output the qualitative result after secondary verification and the mass fractions of each component after smoothing processing.

[0042] Collect raw data, where the raw data contains data information of each component. Preprocess the raw data, and the preprocessing includes data format conversion to obtain the data to be verified that conforms to the data verification standard format. According to the data to be verified, perform data splitting to split it into multiple component data. Perform the first verification on the multiple component data respectively. If the first verification result conforms to the preset range, use the first verification result as the input data for the second verification. If the first verification result does not conform to the preset range, re-collect the component data corresponding to the first verification result, and perform the first verification on the re-collected data. According to the first verification result, perform the second verification on the multiple component data that conform to the preset range. If the second verification result conforms to the preset range, use the second verification result as the qualitative result. If the second verification result does not conform to the preset range, re-collect the component data corresponding to the second verification result, and perform the first verification on the re-collected data. According to the second verification result, use the weighted average algorithm for the multiple component data that conform to the preset range to obtain the smoothed data containing the mass fractions of each component. According to the smoothed data, use the cubic spline interpolation algorithm to obtain the mass fraction curves of each component, and perform data fitting on the mass fraction curves of each component to obtain the fitting functions of the mass fractions of each component. According to the fitting functions, use the support vector machine algorithm to obtain the classification models of the mass fractions of each component, and train the classification models of the mass fractions of each component to obtain the trained classification models of the mass fractions of each component. According to the trained classification models of the mass fractions of each component, use the random forest algorithm to classify the qualitative result to obtain the final classification result, and output the final classification result and the mass fractions of each component corresponding to the trained classification models of the mass fractions of each component.

[0043] The present invention discloses a method for qualitative and quantitative analysis of components in sewage. Aiming at the problems of complex components and large content differences in sewage samples, which lead to difficulties in qualitative and quantitative analysis, first, the two-dimensional retention time data and mass spectrometry data of sewage samples are compared with the standard compound database to obtain a preliminary qualitative result, and the internal standard method is introduced to calculate the response factors of each component to correct the peak area, and then the preliminary mass fractions of each component are calculated. Then, a multiple linear regression model is used to optimize the mass fractions, and data smoothing is performed on the mass fractions below the threshold to improve the quantitative accuracy of low-content components. Finally, the support vector machine is used to perform a secondary verification on the preliminary qualitative result to ensure the reliability of the qualitative result. The present invention integrates various technical means such as data comparison, internal standard correction, model optimization, data smoothing, and secondary verification, realizes accurate and reliable qualitative and quantitative analysis of complex components in sewage samples, especially improves the quantitative accuracy of low-content components, and provides important technical support for sewage treatment and environmental monitoring.

[0044] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the concept of the present application. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. A method for determining organic pollutants in alkylation oil wastewater based on comprehensive two-dimensional gas chromatography-mass spectrometry, characterized in that: The method comprises: Obtain sewage samples and conduct preliminary analysis to obtain two-dimensional retention time data and mass spectrum data of each component; Comparing the two-dimensional retention time data and mass spectrum data with a standard compound database to obtain preliminary qualitative results; Calculating the peak area of ​​each component in the preliminary qualitative results, introducing a standard by an internal standard method, and obtaining a response factor corresponding to each component; Correcting the peak area of ​​each component according to the response factor corresponding to each component to obtain corrected peak area data of each component; Calculate the mass fraction of each component according to the corrected peak area data of each component, wherein the mass fraction of a single component = (corrected peak area of ​​the single component / sum of corrected peak area data of each component) × 100%; Performing multiple linear regression model optimization on the mass fractions of the components to obtain optimized mass fractions of the components; If the mass fractions of the optimized components are less than the threshold value, data smoothing is performed on the mass fractions less than the threshold value to obtain the mass fractions of the components after smoothing; Performing secondary verification on the preliminary qualitative results by using a support vector machine to obtain a qualitative result after secondary verification; The qualitative results after the secondary verification and the mass fractions of the components after the smoothing process are output.

2. The method according to claim 1, characterized in that The sewage sample is obtained and the two-dimensional retention time data and mass spectrum data of each component are obtained by preliminary analysis, including: Obtain two-dimensional chromatographic data of the sewage sample, and determine the initial time of each component in the chromatogram according to a pre-established chromatographic peak recognition algorithm; According to the initial time of the component and the preset time difference, the end time of each component is determined, and the component type and the corresponding primary mass spectrum data are obtained; According to the component type, the corresponding secondary mass spectrum data is matched from the pre-established mass spectrum database, and the similarity between the primary mass spectrum data and the secondary mass spectrum data is calculated. If the similarity is greater than a preset threshold, the component name is determined according to the corresponding secondary mass spectrum data; According to the component name, the standard chromatogram of the component is obtained from the database, and the actual spectrum data of the component is compared with the standard chromatogram to determine whether the actual spectrum data is distorted. If distortion is present, correction processing is performed; According to the corrected spectrum data, the peak area of ​​each component is obtained by integration calculation, and the concentration of each component is determined by combining with the pre-established calibration curve; According to the concentration of each component and the pre-established concentration database, it is determined whether the concentration of each component in the sewage sample exceeds the standard. If it exceeds the standard, an alarm signal is generated; According to the alarm signal and the pre-established sewage treatment solution library, the corresponding sewage treatment solution is obtained and executed through the control system.

3. The method according to claim 1, characterized in that The two-dimensional retention time data and mass spectrum data are compared with a standard compound database to obtain preliminary qualitative results, including: After obtaining the two-dimensional retention time data and mass spectrum data of the sample to be tested, the two-dimensional retention time data of the sample to be tested is preprocessed to remove background noise signals; Extract characteristic peak information according to the pre-processed two-dimensional retention time data to obtain the retention time and peak area of ​​the characteristic peak; The molecular mass of the compound in each characteristic peak is calculated by combining the extracted characteristic peak information with the fragment ion information in the mass spectrum data; According to the molecular mass information of each standard substance in the standard compound database, the standard substances with the same molecular mass as each characteristic peak are screened out to obtain a preliminary screened standard substance set; According to the retention time information of each standard substance in the standard compound database, the standard substances whose retention time is inconsistent with the retention time of the characteristic peak in the preliminary screened standard substance set are removed to obtain a candidate standard substance set; Obtaining the mass spectrum data of each standard substance in the candidate standard substance set, and calculating the similarity coefficient between the mass spectrum data of the candidate standard substance and the mass spectrum data of the sample to be tested, determining whether the similarity coefficient is greater than a preset threshold, and if so, determining that the candidate standard substance is the target substance, and obtaining a preliminary qualitative result; According to the isotope peak information of each target substance in the preliminary qualitative results, the isotope internal standard method was used to obtain the quantitative results of the target substance.

4. The method according to claim 1, characterized in that: The peak area of ​​each component in the preliminary qualitative result is calculated, and a standard substance is introduced by an internal standard method to obtain a response factor corresponding to each component, including: According to the original chromatogram data file, the chromatographic peak is detected using the automatic integration algorithm, and the starting time of the chromatographic peak is recorded; According to the starting time of the chromatographic peak, an integral operation is performed to obtain the peak area obtained by integrating each chromatographic peak, and a peak area list is generated; According to the known concentration of the standard substance and its corresponding peak area, the relationship between the concentration of the standard substance and the peak area is constructed to obtain the standard substance calibration curve equation y=ax+b, where y represents the peak area, x represents the concentration, and a and b represent the slope and intercept of the calibration curve equation respectively; According to the retention time of the chromatographic peak of the substance to be tested, the retention time of the chromatographic peak of the corresponding component in the standard substance is matched to determine the components of the substance to be tested and generate a component list; According to the calibration curve of the standard substance and the peak area of ​​the substance to be tested, the calibrated peak area of ​​each component is calculated and a calibrated peak area list is generated; According to the peak area of ​​each component after calibration and the calibration curve of the standard substance, the concentration of each component is calculated and a component concentration list is generated; According to the component concentration list and the internal standard substance concentration, the relative response factor of each component is calculated through the normalization algorithm to generate a response factor list.

5. The method according to claim 1, characterized in that The step of correcting the peak area of ​​each component according to the response factor corresponding to each component to obtain the corrected peak area data of each component comprises: Read the response factor data table corresponding to each component and the original peak area data table of each component to determine the component set to be calibrated; According to the determined component set and chromatographic conditions, the original peak area of ​​each component is obtained from the original peak area data table to obtain the peak area set to be corrected; If there are missing original peak area data of components in the peak area set to be corrected, data integrity judgment is performed, and the missing values ​​are predicted using the cubic spline interpolation algorithm to obtain the predicted peak area of ​​each component; According to the response factor and the original peak area or predicted peak area of ​​each component, a peak area correction calculation is performed to obtain the preliminary corrected peak area corresponding to each component; Based on the preliminary corrected peak area of ​​each component and the original peak width of each component, the standard deviation normalization algorithm was used to obtain the standard corrected peak area of ​​each component; According to the standard calibrated peak area and the data interface type, the standard calibrated peak area of ​​each component is transmitted to the data receiving end through the data interface to obtain the data receiving status fed back by the receiving end; If the data receiving status fed back by the receiving end is that the receiving is successful, the standard calibration peak area of ​​each component and the corresponding component name are stored in association to determine the final calibration result.

6. The method according to claim 1, characterized in that The mass fraction of each component is calculated according to the corrected peak area data of each component, wherein the mass fraction of a single component = (corrected peak area of ​​the single component / sum of corrected peak area data of each component) × 100%, including: Acquire the chromatographic peak area data of each component to obtain a chromatographic peak area data set; Correcting the chromatographic peak area data of each component according to the chromatographic peak area data set to obtain the corrected peak area data of each component, and determining whether the corrected peak area data of each component is greater than a preset threshold value, and if so, determining that the corrected peak area data of each component is valid data; Calculate the sum of the corrected peak area data of all components, and determine the sum of the corrected peak area data of all components by accumulating the corrected peak area data of each component; Calculate the mass fraction of a single component by dividing the corrected peak area of ​​a single component by the sum of the corrected peak area data of all components to obtain the mass fraction calculation result of a single component; According to the calculation result of the mass fraction of a single component, percentage conversion is performed, and the calculation result of the mass fraction of a single component is multiplied by 100% to determine the percentage value of the mass fraction of the single component; Collect component concentration data during chromatographic separation, combine the obtained mass fraction percentage values ​​of each component, construct a set of corresponding relationships between the mass fraction of each component and the component concentration, train the set of corresponding relationships between the mass fraction of each component and the component concentration through the machine learning algorithm gradient boosting decision tree, and obtain a model of corresponding relationships between the mass fraction of each component and the component concentration; The chromatographic data of the sample to be tested is collected, and the corresponding relationship model between the mass fraction of each component and the component concentration is input. The support vector machine algorithm or the K nearest neighbor algorithm is combined with the corresponding relationship model between the mass fraction of each component and the component concentration to obtain the predicted value of the mass fraction of each component of the sample to be tested.

7. The method according to claim 1, characterized in that The method of performing multiple linear regression model optimization on the mass fractions of the components to obtain the optimized mass fractions of the components comprises: Collect initial data of each component and use it as training data input for the multiple linear regression model; According to the training data input, data fitting is performed to obtain the initial multiple linear regression model; Set the objective function, which integrates the component proportion information and takes minimizing the difference between the predicted value and the target data as the optimization goal; Adjust the model parameters through iterative algorithms to obtain optimized model parameters; The component mass fractions of the sample to be tested are used as the input of the optimized model to obtain the predicted values ​​of the mass fractions of each component after regression analysis; A pre-established quality score threshold is used to determine whether the predicted value deviates from the threshold range. If the predicted value deviates from the threshold range, the data is collected again; Based on the re-collected data, a multiple linear regression model was constructed to determine the final mass fraction of each component.

8. The method according to claim 1, characterized in that If the mass fractions of the optimized components are less than the threshold value, data smoothing is performed on the mass fractions less than the threshold value to obtain the mass fractions of the components after smoothing, including: Obtaining the optimized initial mass fraction of each component, calculating the sum of the mass fractions of each component by a preset algorithm, and obtaining a total mass fraction; According to the ratio of the optimized initial mass fraction of each component to the total mass fraction, it is judged whether the mass fraction of each component is lower than a preset threshold value. If there is a component mass fraction lower than the preset threshold value, the component information corresponding to the component mass fraction is recorded; For the component mass fractions that are lower than the preset threshold, a moving average algorithm is used to calculate the component mass fractions that are lower than the threshold to obtain a smoothed value of the component mass fraction; According to the calculation results of the moving average algorithm, the mass fraction of each component is processed by using the data normalization algorithm to obtain the sum of the new mass fractions of each component and a new total mass fraction; According to the new mass scores of each component after the data normalization algorithm is processed, the mass scores of each component are classified by the support vector machine algorithm to obtain the classification results of the mass scores of each component; According to the classification results of the support vector machine algorithm, the principal component analysis algorithm is used to perform dimensionality reduction processing to obtain the mass fraction data of each component after dimensionality reduction processing; According to the mass fraction data of each component after dimensionality reduction processing by principal component analysis algorithm, prediction analysis is carried out through logistic regression algorithm to obtain the prediction results of the mass fraction of each component.

9. The method according to claim 1, characterized in that: The preliminary qualitative results are verified twice by a support vector machine to obtain the qualitative results after the second verification, including: Based on the preliminary qualitative results, a data sample set was constructed; According to the data sample set, the feature space is obtained; Perform model training based on feature space; If there are outliers in the data sample set, data preprocessing is performed to obtain a preprocessed data sample set. If there are no outliers in the data sample set, no processing is required. According to the preprocessed data sample set, a classification model is determined to obtain an initial classification model; Through the initial classification model, a hyperplane is constructed to obtain the classification decision boundary; According to the classification decision boundary, the preliminary qualitative results are verified for the second time to obtain the qualitative results after the second verification.

10. The method according to claim 1, characterized in that The output of the qualitative results after the secondary verification and the mass fractions of the components after the smoothing process includes: Collecting raw data, which includes data information of each component, preprocessing the raw data, which includes data format conversion, to obtain data to be verified that meets the data verification standard format; According to the data to be verified, the data is split into multiple component data, and the multiple component data are respectively verified for the first time. If the first verification result meets the preset range, the first verification result is used as the input data for the second verification. If the first verification result does not meet the preset range, the component data corresponding to the first verification result is re-collected, and the re-collected data is verified for the first time; According to the first verification result, a second verification is performed on multiple component data that meet the preset range. If the second verification result meets the preset range, the second verification result is used as a qualitative result. If the second verification result does not meet the preset range, the component data corresponding to the second verification result is re-collected, and the re-collected data is subjected to the first verification; According to the second verification results, a weighted average algorithm is used for multiple component data that meet the preset range to obtain smoothed data containing the mass fraction of each component; According to the smoothed data, the cubic spline interpolation algorithm is used to obtain the mass fraction curve of each component, and the mass fraction curve of each component is fitted with data to obtain the fitting function of the mass fraction of each component; According to the fitting function, a support vector machine algorithm is used to obtain a classification model of the mass fraction of each component, and the classification model of the mass fraction of each component is trained to obtain a trained classification model of the mass fraction of each component; According to the trained classification model of mass scores of each component, the random forest algorithm is used to classify the qualitative results to obtain the final classification result, and the final classification result and the mass scores of each component corresponding to the trained classification model of mass scores of each component are output.