A gas chromatography detection method for determining benzene series in water is optimized by optimizing split ratio
Patent Information
- Application Number
- CN202511769237.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-11-28
AI Technical Summary
[0003]有鉴于此,本发明提供一种优化分流比测定水中苯系物的气相色谱检测方法,能够解决现有技术中存在气相色谱测定水中苯系物时分流比参数难以兼顾甲醇溶剂峰与苯峰分离效果和目标物检测灵敏度的技术问题
[0020] This invention achieves intelligent prediction of the optimal split ratio parameter by constructing a split ratio optimization prediction model, solving the technical problem of balancing separation effect and detection sensitivity. The model employs multi-scale convolution to extract complex feature patterns of the split ratio-resolution-response value relationship. Through independent component analysis blind source separation technology, it effectively separates the aliased methanol interference signal and benzene series response signal in the feature space, overcoming the shortcomings of traditional feature extraction methods that misclassify interference components as valid signals or submerge valid signals in the interference background. The separated interference feature components are used to evaluate the degree of solvent peak masking and optimize resolution prediction, while the target feature components purely reflect the detector response capability and optimize sensitivity prediction. The outputs of the two parallel prediction branches are weighted and fused through multi-objective analysis to generate a comprehensive evaluation index, enabling the model to simultaneously optimize two mutually constraining objectives at the source signal level. The optimal split ratio parameter predicted by the model can retain the target analyte response value to the maximum extent while ensuring effective separation of methanol and benzene, improving the detection capability and quantitative accuracy of low-concentration benzene series compounds. In summary, this invention solves the technical problem mentioned in the background art where the split ratio parameter is difficult to balance separation effect and detection sensitivity.
Smart Images

Figure CN121385153B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of ecological environment monitoring technology, and specifically relates to a gas chromatography detection method for determining benzene series compounds in water with optimized split ratio. Background Technology
[0002] Gas chromatography is a commonly used analytical method for determining benzene series pollutants in water bodies, and it is widely used in environmental monitoring, water quality assessment, and other fields. Traditional methods control the injection volume by adjusting the split ratio parameter. While a lower split ratio parameter can achieve higher detection sensitivity, the methanol solvent peak easily interferes with the benzene peak, leading to insufficient separation and affecting quantitative accuracy. To improve separation, analysts often increase the split ratio parameter to reduce solvent interference; however, this significantly reduces the peak area response value of the target analyte, lowering the detection capability for low-concentration samples. Current technologies mainly rely on empirical exploration or single-factor experiments to test the separation and response intensity under different split ratio parameters, lacking a systematic optimization strategy and making it difficult to accurately grasp the balance between separation effect and detection sensitivity. Given the increasingly stringent requirements for trace pollutant monitoring, the lack of scientific predictive methods for split ratio parameter selection often prevents analytical methods from simultaneously meeting the dual requirements of effective separation and high-sensitivity detection. In other words, current technologies suffer from the technical problem of split ratio parameters failing to balance separation effect and detection sensitivity. Summary of the Invention
[0003] In view of this, the present invention provides a gas chromatography detection method for benzene series compounds in water with optimized split ratio, which can solve the technical problem in the prior art that the split ratio parameter is difficult to balance the separation effect of methanol solvent peak and benzene peak and the detection sensitivity of target analytes when determining benzene series compounds in water by gas chromatography.
[0004] This invention is implemented as follows: It provides a gas chromatography detection method for optimizing the split ratio of benzene compounds in water, including a sample pretreatment stage where pH is adjusted and sodium chloride is added; a column selection stage where columns meeting the resolution parameter requirements are chosen; a three-dimensional dataset of split ratio, resolution parameter, and peak area response value is established; the three-dimensional dataset is input into a split ratio optimization prediction model; the split ratio optimization prediction model extracts features through multi-scale convolution and then separates the fused feature vector into interference feature components and target feature components through an independent component analysis blind source separation unit; these are then input into two parallel fully connected prediction branches to predict the resolution parameter and response intensity parameter, respectively; a multi-objective weighted fusion layer generates a comprehensive evaluation index and outputs the optimal split ratio parameter; in the headspace equilibrium condition optimization stage, a Gaussian process regression Bayesian optimization algorithm is used to determine the optimal equilibrium temperature and optimal equilibrium time, and a calibration curve is established; in the actual sample determination stage, analysis is performed according to the optimal split ratio parameter and optimal equilibrium conditions; and in the matrix effect correction stage, a calibration curve is established using the standard addition method.
[0005] Specifically, the sample pretreatment stage involves adding ascorbic acid and hydrochloric acid solutions to the water sample to be tested to adjust the pH value to below 2, and then adding sodium chloride at a mass-volume ratio of 0.3 g / mL and dissolving it completely.
[0006] Specifically, the column screening stage involves using a DB-624 medium-polarity column and a DB-WAX strong-polarity column to separate and test the benzene series standard solutions, recording the resolution parameters of p-xylene and m-xylene, and selecting columns with a resolution parameter greater than 1.5 as analytical columns.
[0007] The step of establishing a three-dimensional dataset of split ratio-resolution parameter-peak area response value involves selecting a test point every 5 units within a split ratio range of 5:1 to 30:1, performing headspace gas chromatography analysis on a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L, and recording the resolution parameter between the methanol solvent peak and the benzene peak, as well as the peak area response values of the eight benzene series compounds.
[0008] Specifically, the structure of the split ratio optimization prediction model is as follows: the input layer receives a three-dimensional dataset, which is then processed by data standardization and enters the feature extraction layer. The feature extraction layer contains three parallel one-dimensional convolutional branches with kernel sizes of 3, 5, and 7, respectively. The output is concatenated into a fused feature vector through a feature fusion layer.
[0009] Specifically, the independent component analysis blind source separation unit treats the fused feature vector as a mixed representation of multiple independent source signals, finds the unmixing matrix by maximizing the non-Gaussianity criterion, optimizes the separation matrix parameters using a fast independent component analysis iterative algorithm, and separates the interference feature components and the target feature components from the fused feature vector.
[0010] Specifically, the calculation of the comprehensive evaluation index generated by the multi-objective weighted fusion layer is as follows: the comprehensive evaluation index is equal to the weighted sum of the quotient of the separation degree parameter divided by the standard separation degree threshold and the quotient of the response intensity parameter divided by the standard response intensity threshold. The separation degree weight is set to 0.6 and the response intensity weight is set to 0.4.
[0011] Specifically, the optimal split ratio parameter is used as the subsequent analysis condition when the predicted resolution parameter is greater than 1.2 and the peak area response value of benzene is greater than 40% of the initial peak area response threshold.
[0012] Specifically, in the headspace equilibrium condition optimization stage, the initial equilibrium temperature is set to 60℃ and the initial equilibrium time is set to 30min. A Gaussian process regression Bayesian optimization algorithm is used to optimize the parameters within the equilibrium temperature range of 50℃ to 70℃ and the equilibrium time range of 20min to 40min, with the objective function value being the maximization of the total peak area response value of the eight benzene series compounds.
[0013] Specifically, the application of the Gaussian process regression Bayesian optimization algorithm involves selecting the squared exponential covariance function to describe the correlation between different points in the input space, using the expectation improvement acquisition function to select the next experimental point, executing the experiment and updating the Gaussian process posterior distribution until the optimal equilibrium temperature and optimal equilibrium time are found.
[0014] The step of establishing the calibration curve involves creating a series of mixed standard working solutions with concentration gradients of 50 μg / L, 200 μg / L, 500 μg / L, 2000 μg / L, 4000 μg / L, and 5000 μg / L. Each concentration level is measured in parallel three times. Peak identification and deconvolution separation algorithms are used to process partially overlapping chromatographic peaks and establish the calibration curve.
[0015] Specifically, the peak identification and deconvolution separation algorithm involves preprocessing the original chromatographic signal with wavelet denoising, identifying the peak position by combining the zero crossover point of the second derivative with the sign change of the first derivative, establishing a multi-peak superposition model for the overlapping peak region, and optimizing the model parameters using the Levenberg-Marquardt nonlinear least squares algorithm.
[0016] Specifically, in the actual sample determination stage, the pretreated water sample to be tested is subjected to gas chromatography analysis according to the optimal split ratio parameter, optimal equilibrium temperature and optimal equilibrium time. When the resolution parameter between the methanol solvent peak and the benzene peak is less than 1.0, the optimal split ratio parameter is increased by 5 units and the determination is repeated.
[0017] Specifically, the matrix effect correction stage involves adding standard solutions at three concentration levels (low, medium, and high) to the water sample containing oil or high molecular weight organic matter. These three concentration levels represent 50%, 100%, and 150% of the estimated concentration of the sample, respectively. The actual concentration of benzene compounds in the water sample is then calculated using linear regression.
[0018] Specifically, the resolution parameter is calculated by dividing the difference in retention times of the two chromatographic peaks by the sum of the half-peak widths of the two chromatographic peaks and then multiplying by 2.
[0019] Specifically, the implementation of the independent component analysis blind source separation unit involves centering and whitening the fusion feature matrix as preprocessing steps, and iteratively optimizing the separation matrix parameters to maximize the negative entropy of each independent component after separation. The iteration termination condition is that the Frobenius norm of the change in the separation matrix parameters between two adjacent iterations is less than 0.0001 or the number of iterations reaches 200.
[0020] This invention achieves intelligent prediction of the optimal split ratio parameter by constructing a split ratio optimization prediction model, solving the technical problem of balancing separation effect and detection sensitivity. The model employs multi-scale convolution to extract complex feature patterns of the split ratio-resolution-response value relationship. Through independent component analysis blind source separation technology, it effectively separates the aliased methanol interference signal and benzene series response signal in the feature space, overcoming the shortcomings of traditional feature extraction methods that misclassify interference components as valid signals or submerge valid signals in the interference background. The separated interference feature components are used to evaluate the degree of solvent peak masking and optimize resolution prediction, while the target feature components purely reflect the detector response capability and optimize sensitivity prediction. The outputs of the two parallel prediction branches are weighted and fused through multi-objective analysis to generate a comprehensive evaluation index, enabling the model to simultaneously optimize two mutually constraining objectives at the source signal level. The optimal split ratio parameter predicted by the model can retain the target analyte response value to the maximum extent while ensuring effective separation of methanol and benzene, improving the detection capability and quantitative accuracy of low-concentration benzene series compounds. In summary, this invention solves the technical problem mentioned in the background art where the split ratio parameter is difficult to balance separation effect and detection sensitivity. Attached Figure Description
[0021] Figure 1 This is a flowchart of the method of the present invention.
[0022] Figure 2 This is a comparison chart of chromatographic separation effects under different split ratio parameters.
[0023] Figure 3 The image shows the separation effect of interference feature components and target feature components after blind source separation in independent component analysis.
[0024] Figure 4 This is a comparison chart showing the results of peak identification and deconvolution separation algorithms in processing overlapping peaks.
[0025] Figure 5 A schematic diagram of the structure of the prediction model for optimizing the split ratio. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below.
[0027] like Figure 1The diagram shows a flowchart of a gas chromatography detection method for benzene series compounds in water with optimized split ratio, provided by the present invention. This method includes the following steps:
[0028] S01. In the sample pretreatment stage, ascorbic acid and hydrochloric acid solution are added to the water sample to be tested to adjust the pH value to below 2. Then, sodium chloride is added at a mass-volume ratio of 0.3 g / mL and fully dissolved. The sodium chloride is used to enhance the salting-out effect.
[0029] S02. In the column screening stage, DB-624 medium polarity column and DB-WAX strong polarity column were used to separate and test the benzene series standard solutions, respectively. The resolution parameters of p-xylene and m-xylene were recorded, and the column with a resolution parameter greater than 1.5 was selected as the analytical column.
[0030] S03. Establish a split ratio response relationship dataset. Select a test point every 5 units within the split ratio range of 5:1 to 30:1, and perform headspace gas chromatography analysis on a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L. Record the resolution parameters of the methanol solvent peak and the benzene peak, as well as the peak area response values of the eight benzene series compounds, and establish a three-dimensional dataset of split ratio-resolution parameter-peak area response value.
[0031] S04. Input the three-dimensional dataset of split ratio-resolution parameter-peak area response value into the split ratio optimization prediction model for processing. The split ratio optimization prediction model outputs the optimal split ratio parameter and the predicted resolution parameter. When the predicted resolution parameter is greater than 1.2 and the peak area response value of benzene is greater than 40% of the initial peak area response threshold, the optimal split ratio parameter is used as the subsequent analysis condition.
[0032] S05, Headspace equilibrium condition optimization stage: The initial equilibrium temperature was set to 60℃ and the initial equilibrium time was set to 30min. The Gaussian process regression Bayesian optimization algorithm was used to optimize the parameters in the range of 50℃ to 70℃ and the range of 20min to 40min. The objective function value was to maximize the total peak area response value of the eight benzene series compounds. The optimal equilibrium temperature and optimal equilibrium time were recorded.
[0033] S06. Establish a series of mixed standard working solutions with concentration gradients of 50 μg / L, 200 μg / L, 500 μg / L, 2000 μg / L, 4000 μg / L, and 5000 μg / L. Each concentration level is measured in parallel three times. Peak identification and deconvolution separation algorithms are used to process partially overlapping chromatographic peaks, separate each single peak, calculate the peak area of each single peak, and establish a calibration curve.
[0034] S07. In the actual sample determination stage, the water sample to be tested, which has been pretreated in step S01, is placed in a headspace bottle and analyzed by gas chromatography according to the optimal split ratio parameter determined in step S04 and the optimal equilibrium temperature and optimal equilibrium time determined in step S05. When the resolution parameter between the methanol solvent peak and the benzene peak is less than 1.0, the optimal split ratio parameter is increased by 5 units and the determination is repeated.
[0035] S08. In the matrix effect correction stage, for water samples containing oil or high molecular weight organic matter, a calibration curve is established using the standard addition method. Standard solutions at three concentration levels (low, medium, and high) are added to the water sample, respectively, representing 50%, 100%, and 150% of the estimated concentration of the sample. The actual concentration of benzene compounds in the water sample is calculated using linear regression.
[0036] The specific structure of the split ratio optimization prediction model is as follows: the input layer receives a three-dimensional dataset of split ratio, separation degree parameter, and peak area response value, which, after data standardization, enters the feature extraction layer; the feature extraction layer contains three parallel one-dimensional convolutional branches with kernel sizes of 3, 5, and 7, respectively, each extracting feature patterns at different scales; the outputs of the three parallel one-dimensional convolutional branches are concatenated through a feature fusion layer, and the fused feature vector is input to the independent component analysis blind source separation unit; the independent component analysis blind source separation unit treats the fused feature vector as a mixed representation of multiple independent source signals, finds the unmixing matrix by maximizing the non-Gaussianity criterion, optimizes the separation matrix parameters using a fast independent component analysis iterative algorithm, and separates the methanol-related signals from the fused feature vector. The interference-related feature components and the target feature components related to the target object's response are separated. The separated interference feature components and target feature components are respectively input into two parallel fully connected prediction branches. The first fully connected prediction branch predicts the separation degree parameter, and the second fully connected prediction branch predicts the response intensity parameter. The outputs of the two parallel fully connected prediction branches generate a comprehensive evaluation index through a multi-objective weighted fusion layer. The comprehensive evaluation index is calculated as follows: the comprehensive evaluation index is equal to the weighted sum of the quotient of the separation degree parameter divided by the standard separation degree threshold and the quotient of the response intensity parameter divided by the standard response intensity threshold, where the separation degree weight is set to 0.6 and the response intensity weight is set to 0.4. The final output layer of the split ratio optimization prediction model outputs the optimal split ratio parameter according to the principle of maximizing the comprehensive evaluation index.
[0037] The steps for establishing the training dataset for the split ratio optimization prediction model specifically include: collecting split ratio parameters, corresponding methanol-benzene separation parameters, and peak area response values of eight benzene series compounds used by different laboratories on different chromatographic instruments when determining benzene series compounds. The peak area response value data covers a split ratio parameter range of 5:1 to 30:1; performing quality assessment on each group of experimental data, removing experimental data with a separation parameter less than 0.5 or abnormal fluctuations in peak area response values; and grouping the qualified experimental data according to the split ratio parameters, with each split ratio parameter level down to... The dataset contains at least 20 sets of valid experimental data. The separation parameter and the peak area response value are normalized by subtracting the minimum value of the parameter corresponding to the original value from the original value and then dividing by the range of the parameter corresponding to the original value. Data augmentation samples are constructed by adding a Gaussian-distributed random perturbation to the qualified experimental data, with the perturbation standard deviation set to 10% of the standard deviation of the qualified experimental data. Five sets of data augmentation samples are generated for each set of qualified experimental data. All experimental data are divided into training, validation, and test sets in an 8:1:1 ratio.
[0038] The specific steps of training the split ratio optimization prediction model include: initializing model parameters, with the parameters of the convolutional layers of the one-dimensional convolutional branch using the He Kaiming initialization method and the parameters of the fully connected layers of the fully connected prediction branch using the Xavier initialization method; setting the batch size to 32 and the initial learning rate to 0.001, and updating the parameters using an adaptive moment estimation optimizer; in the feature extraction layer training stage, comparing the outputs of the three parallel one-dimensional convolutional branches with manually labeled feature importance tags, and calculating the feature extraction error using the cross-entropy loss function; in the independent component analysis blind source separation unit training stage, performing centering and whitening preprocessing on the fused feature matrix composed of the fused feature vectors, and iteratively optimizing the separation matrix parameters to ensure that the separated independent components are optimized. The negative entropy of the discrete components is maximized, and the iteration termination condition is that the Frobenius norm of the change in the separation matrix parameters between two adjacent iterations is less than 0.0001 or the number of iterations reaches 200. During the training phase of the fully connected prediction branch, the mean squared error loss function is used to calculate the separation degree prediction error and the response intensity prediction error, respectively. The total loss function is the weighted sum of the separation degree prediction error and the response intensity prediction error. After each training cycle, the performance of the split ratio optimization prediction model is evaluated on the validation set. An early stopping mechanism is triggered when the validation set loss does not decrease for 10 consecutive training cycles. After training, the generalization performance of the split ratio optimization prediction model is evaluated on the test set, and the mean absolute error of the separation degree parameter prediction and the mean absolute percentage error of the response intensity parameter prediction are calculated.
[0039] The salting-out effect refers to the phenomenon that when inorganic salts are added to an aqueous solution, the solubility of nonpolar or weakly polar organic compounds in the aqueous phase decreases, and more of them are transferred to the gas phase. The mechanism is that inorganic salt ions form hydration with water molecules, reducing the number of free water molecules available to dissolve organic compounds. At the same time, the presence of salt ions increases the ionic strength of the solution, disrupting the weak interaction between organic molecules and water molecules.
[0040] The resolution parameter is used to quantitatively evaluate the degree of separation between two adjacent chromatographic peaks. The calculation method is to divide the difference in retention time of the two chromatographic peaks by the sum of the half-peak widths of the two chromatographic peaks and then multiply by 2. When the resolution parameter is greater than 1.5, it indicates that the two chromatographic peaks have achieved baseline separation.
[0041] The application steps of the Gaussian process regression Bayesian optimization algorithm are as follows: The initial values of the equilibrium temperature and equilibrium time are used as two-dimensional input variables, and the total peak area response value of the eight benzene compounds is used as the objective function value output. A squared exponential covariance function is selected to describe the correlation between different points in the input space. The length scale parameter and signal variance parameter of the squared exponential covariance function are estimated by maximizing the marginal likelihood function. The Gaussian process posterior distribution of the objective function value is calculated based on existing experimental data points. The posterior mean represents the predicted total peak area response value, and the posterior variance represents the prediction uncertainty. The expected improvement acquisition function is used to select the next experimental point. This expected improvement acquisition function gives greater weight to regions with high predicted values and high uncertainty, balancing the utilization of known optimal solutions and the exploration of unknown regions. The experiment at the selected parameter points is executed, and new experimental data is added to the experimental dataset. The Gaussian process posterior distribution is updated, and the above application steps are repeated until the optimal equilibrium temperature and optimal equilibrium time that maximize the objective function value are found.
[0042] The peak identification and deconvolution separation algorithm's processing flow is as follows: The original chromatographic signal undergoes wavelet denoising preprocessing to remove baseline drift and high-frequency noise; the first and second derivatives of the original chromatographic signal are calculated, and potential peak positions are identified by combining the zero-crossing points of the second derivative with the sign changes of the first derivative; the sharpness of each candidate peak is evaluated using curvature analysis, and false peaks with curvature less than a curvature threshold are eliminated; a multi-peak stacking model is established for the identified overlapping peak regions, assuming each single peak conforms to a Gaussian function or an exponentially modified Gaussian function, and the multi-peak stacking model parameters include peak height, peak position, peak width, and asymmetry factor; the Levenberg-Marquardt nonlinear least squares algorithm is used to optimize the multi-peak stacking model parameters to minimize the sum of squared residuals between the fitted curve and the measured curve; during the iteration process, physical constraints are applied to the multi-peak stacking model parameters to ensure that the peak height is positive and the peak width is within a reasonable range; after separating each single peak, numerical integration is performed on each single peak to calculate its peak area, using either the trapezoidal rule or Simpson's rule for numerical integration.
[0043] The implementation mechanism of the independent component analysis blind source separation unit in the split ratio optimization prediction model is as follows: The fused feature vector output by the feature fusion layer is reconstructed into the fused feature matrix, where each column of the fused feature matrix represents the value of a feature dimension across all samples; the fused feature matrix is centered, i.e., its mean is subtracted from each feature dimension; the centered fused feature matrix is whitened using principal component analysis, making the covariance matrix of the whitened fused feature matrix an identity matrix, and the whitening process eliminates the second-order correlation between features; the separation matrix parameters are initialized to an identity matrix, and each row vector of the separation matrix parameters is updated using the fixed-point iterative algorithm of the fast independent component analysis iterative algorithm; the whitening is calculated in each iteration. The linear combination of features and the current separation vector is used to apply a nonlinear function and its derivative to the linear combination signal. The nonlinear function is selected as a hyperbolic tangent function or a cubic function to maximize non-Gaussianity. The separation vector is updated by gradient ascent to maximize the negative entropy of the output signal. The negative entropy is defined as the difference between the differential entropy of the output signal and the entropy of the homoscedastic Gaussian distribution. The updated separation vector is subjected to Schmitt orthogonalization to ensure that the separation vectors are mutually orthogonal. The iteration is repeated until the separation vector converges or the maximum number of iterations is reached. The final separation matrix parameters are applied to the whitening feature matrix to obtain statistically independent feature components. The interference feature components typically exhibit time-domain waveform characteristics and frequency-domain energy distribution, while the target feature components exhibit amplitude variation patterns positively correlated with concentration. The technical effect of the independent component analysis blind source separation unit for the split ratio optimization prediction model is reflected in its ability to automatically identify and separate the contributions of different physical processes aliased in the feature space, distinguishing between methanol solvent interference signals and target benzene series response signals without manual annotation. Traditional feature extraction methods struggle to handle mixed multi-source signals, often misclassifying interfering components as valid signals or submerging valid signals in the background interference, leading to deviations from the optimal solution in split ratio optimization. The independent component analysis blind source separation unit, by maximizing the statistical independence criterion, fundamentally reveals the intrinsic structure of the data, decomposing complex mixed observation signals into a few independent source signals with clear physical meaning. The separated interfering feature components are used to assess the masking effect of the methanol peak on the benzene peak, guiding the first fully connected prediction branch to make more accurate judgments; the target feature components directly reflect the detector response capability of benzene series compounds under different split ratio parameters, providing pure information input for the second fully connected prediction branch. Source signal-level separation allows the split ratio optimization prediction model to simultaneously optimize both separation effectiveness and detection sensitivity, avoiding the dilemma of sacrificing one for the other in traditional methods.For the entire detection method, the optimal split ratio parameter predicted by the split ratio optimization prediction model can retain the peak area response value of the target analyte to the maximum extent while ensuring effective separation of methanol and benzene. This solves the core contradiction that increasing the split ratio parameter improves separation but reduces sensitivity, improves the detection capability and quantitative accuracy of low-concentration benzene series compounds, expands the applicable concentration range of the detection method, and meets the actual needs of monitoring trace pollutants in environmental water samples.
[0044] The standard addition method establishes a linear relationship between the amount of standard substance added and the measured response value by adding a standard substance of known concentration to the sample to be tested. The original concentration of the target substance in the sample to be tested is calculated by the intercept of the linear regression equation. The standard addition method can effectively eliminate the influence of matrix effect on the measurement results.
[0045] The Frobenius norm is the arithmetic square root of the sum of the squares of all elements of a matrix. It is used to measure the size of a matrix and is often used in iterative algorithms to determine whether the changes in a matrix have converged to an acceptable range of precision.
[0046] Negative entropy is an indicator in information theory that measures the degree to which a random variable is non-Gaussian. According to the central limit theorem, the sum of multiple independent random variables tends to a Gaussian distribution. Therefore, maximizing the negative entropy is equivalent to finding the most independent signal component. The larger the negative entropy value, the more the signal component deviates from the Gaussian distribution and the more likely it is to be an independent source signal.
[0047] The initial peak area response threshold is the peak area response value of benzene obtained by measuring a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L at a split ratio parameter of 5:1.
[0048] The estimated concentration of the sample is obtained by preliminarily measuring the water sample under the optimal split ratio parameter, the optimal equilibrium temperature, and the optimal equilibrium time, and then estimating it based on the calibration curve.
[0049] The specific implementation methods of the above steps are described in detail below.
[0050] The specific implementation of step S01 is as follows: First, the collected water sample to be tested is transferred to a sample processing container. 25 mg of ascorbic acid is added to the sample processing container as a reducing agent to eliminate oxidative interference from residual chlorine. Then, a 10% hydrochloric acid solution is added dropwise, and the pH value of the water sample is monitored in real time using a pH meter. When the pH value drops below 2, acid addition is stopped to inhibit microbial activity and stabilize the target compound. Then, proceed as follows... Weigh out the sample at a ratio of 0.3 g / mL to the volume of the water sample. The solid was added to the sample processing container and stirred continuously for 5 to 10 minutes using a magnetic stirrer until... Completely dissolved, the The addition of [a substance] utilizes the salting-out effect to reduce the solubility of benzene compounds in the aqueous phase, thereby improving their transfer efficiency to the headspace phase. Finally, the treated water sample is transferred to a headspace vial and immediately sealed for storage. The purpose of step S01 is to create optimal conditions for subsequent headspace extraction by adjusting the chemical environment parameters of the water sample.
[0051] The specific implementation of step S02 involves preparing a mixed standard solution of eight benzene series compounds with a mass concentration of 10 mg / L. The mixed standard solution is then separated and determined using a DB-624 medium-polarity column and a DB-WAX high-polarity column under the same chromatographic conditions. The chromatographic conditions are set as follows: column temperature 40℃, constant temperature for 10 min; carrier gas helium flow rate 1.0 mL / min; split ratio 10:1. The retention times and half-peak widths of the two chromatographic peaks (p-xylene and m-xylene) are recorded. The difference in retention times between the two chromatographic peaks is then divided according to the resolution parameter calculation formula. The resolution parameter is obtained by multiplying the sum of the half-peak widths of the two chromatographic peaks by 2. For the DB-624 medium polar column, the resolution parameter is usually less than 0.8, indicating that p-xylene and m-xylene cannot be effectively separated. However, the DB-WAX strong polar column, due to its stronger stationary phase polarity, can generate different interaction forces based on the positional differences of the substituents of isomers, thereby achieving a baseline separation effect with a resolution parameter greater than 1.5. Therefore, the DB-WAX strong polar column is selected as the analytical column for subsequent experiments. Step S02 solves the technical difficulty of isomer separation.
[0052] The specific implementation of step S03 involves preparing a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L. The mixed standard solution is added to a headspace vial at a volume of 10.0 mL and sealed. Six test points are sequentially set with split ratio parameters of 5:1, 10:1, 15:1, 20:1, 25:1, and 30:1. Headspace gas chromatography analysis of the mixed standard solution is performed under each split ratio parameter. The headspace conditions are set as follows: equilibrium temperature 60℃, equilibrium time 30 min, injection volume 1.0 mL; chromatographic conditions are set as follows: injection port temperature 200℃, detector temperature 250℃, column temperature program 40℃ held for 5 min, followed by... The temperature was increased to 90℃ at a rate of 5℃ / min and held for 1 min. Chromatographic signals were acquired using a flame ionization detector. Peak identification processing was performed on the chromatograms obtained under each split ratio parameter condition. The retention time difference and half-peak width of the methanol solvent peak and benzene peak were measured, and the resolution parameter was calculated. At the same time, the peak area response values of the chromatographic peaks of the eight benzene series compounds were calculated by integrating the peaks. The split ratio parameter, the resolution parameter, and the peak area response value were organized into a three-dimensional dataset of split ratio-resolution parameter-peak area response value according to a one-to-one correspondence. Step S03 established a quantitative data foundation for the relationship between the split ratio parameter and the separation effect and detection sensitivity through systematic experiments.
[0053] The specific implementation of step S04 involves loading the three-dimensional dataset of split ratio-separation parameter-peak area response value established in step S03 as input data into the split ratio optimization prediction model. The input layer of the split ratio optimization prediction model standardizes the three-dimensional dataset, normalizing the numerical range to between 0 and 1. The standardized data then enters the feature extraction layer. The three parallel one-dimensional convolutional branches of the feature extraction layer use convolutional kernels with kernel sizes of 3, 5, and 7, respectively, to perform convolution operations on the input data to extract local and global feature patterns. The feature fusion layer concatenates the feature vectors output by the three one-dimensional convolutional branches along the feature dimension to form a fused feature vector. The fused feature vector is input to the independent component analysis blind source separation unit. Based on the statistical independence assumption, the independent component analysis blind source separation unit decomposes the mixed feature signal into interference feature components and target feature components. The two separated feature components are input into the first fully connected prediction branch and the second fully connected prediction branch, respectively. The first fully connected prediction branch outputs the prediction separation parameter, and the second fully connected prediction branch outputs the response intensity parameter. The multi-objective weighted fusion layer calculates the comprehensive evaluation index by weighting the prediction separation parameter and the response intensity parameter with weight coefficients of 0.6 and 0.4, respectively. Finally, the output layer selects the split ratio parameter that maximizes the comprehensive evaluation index as the optimal split ratio parameter, and outputs the prediction separation parameter corresponding to the optimal split ratio parameter. When the prediction separation parameter is greater than 1.2 and the peak area response value of benzene is greater than 40% of the initial peak area response threshold, it is confirmed that the optimal split ratio parameter meets the separation requirements and sensitivity requirements. Step S04 realizes intelligent parameter optimization that maximizes the response intensity of the target substance while ensuring effective separation of methanol and benzene.
[0054] The specific implementation of step S05 involves setting the initial equilibrium temperature to 60℃ and the initial equilibrium time to 30 min as the starting search points for the Gaussian process regression Bayesian optimization algorithm. This algorithm first uses the squared exponential covariance function to describe the correlation structure between different experimental points in the parameter space. Based on the total peak area response data of the initial experimental points, it constructs the prior distribution of the Gaussian process model. It then updates the posterior distribution through Bayesian inference to predict the response values and uncertainties of untested parameter points. Finally, it uses the expected improvement acquisition function to calculate the expected improvement value for each candidate point in the parameter space. This expected improvement acquisition function assigns higher weights near the current optimal value and in regions of high uncertainty to balance the overall performance. In the process of departmental development and global exploration, the parameter point with the largest expected improvement value is selected as the next experimental point. Under the equilibrium temperature and equilibrium time conditions corresponding to the next experimental point, headspace gas chromatography experiments are performed, and the total peak area response values of eight benzene series compounds are measured. The experimental point and its corresponding total peak area response value are added to the experimental dataset, and the posterior distribution of the Gaussian process model is updated. The above optimization iteration process is repeated until the change in the total peak area response value is less than 2% after three consecutive iterations or the number of iterations reaches 15. The parameter combination that maximizes the total peak area response value is recorded as the optimal equilibrium temperature and optimal equilibrium time. In step S05, the global optimal solution of the headspace equilibrium conditions is efficiently found with a small number of experiments using the Bayesian optimization algorithm.
[0055] The specific implementation of step S06 is as follows: First, prepare a mixed standard working solution of eight benzene series compounds with a mass concentration of 100 mg / L. Transfer 1.00 mL of the mixed standard substance of eight benzene series compounds in methanol into a 10 mL volumetric flask, dilute with ultrapure water and bring to the mark, mix well, and then take 6 portions. 3g of each solution was placed in six headspace vials. Following the concentration gradient requirements, 10.00mL, 10.00mL, 10.00mL, 9.8mL, 9.6mL, and 9.5mL of ultrapure water and 5μL, 20μL, 50μL, 200μL, 400μL, and 500μL of the prepared mixed standard solution were added to each of the six headspace vials respectively. The vials were immediately sealed and gently shaken to prepare concentrations of 50μg / L, 200μg / L, 500μg / L, 2000μg / L, 4000μg / L, and 500μL respectively. A series of mixed standard working solutions of 0.00 μg / L were analyzed by headspace gas chromatography at each concentration level under the optimal split ratio parameters determined in step S04 and the optimal equilibration temperature and time determined in step S05. Each concentration level was measured in triplicate to evaluate the precision. The obtained chromatograms were processed using a peak identification and deconvolution separation algorithm. The peak identification and deconvolution separation algorithm first performs wavelet transform denoising on the original chromatographic signal to remove baseline drift and random noise, and then calculates the second derivative of the chromatographic signal. The zero-crossing point position is identified by numerical identification and the potential peak position is determined by combining the sign change of the first derivative. For partially overlapping chromatographic peak regions, a multi-peak superposition model is established, assuming that each single peak follows a Gaussian function distribution or an exponentially modified Gaussian function distribution. The parameters of the multi-peak superposition model include the peak height parameter, peak center position parameter, peak width parameter, and peak asymmetry factor parameter of each single peak. The Levenberg-Marquardt nonlinear least squares iterative algorithm is used to optimize and fit the parameters of the multi-peak superposition model, so that the sum of squared residuals between the superposition curve predicted by the model and the actual measured chromatographic curve reaches the minimum value. During the iteration process, a constraint condition greater than zero is applied to the peak height parameter to ensure the rationality of the physical meaning. After deconvolution separation, the single peak area of each separated single peak is calculated by numerical integration using the trapezoidal rule. The single peak area of the eight benzene series compounds is used to establish a linear regression relationship with the corresponding mass concentration to obtain the calibration curve. The correlation coefficient, intercept, and slope parameters of each calibration curve are calculated. Step S06 solves the problem of accurate quantification of partially overlapping chromatographic peaks and establishes a reliable concentration-response relationship.
[0056] The specific implementation of step S07 involves taking 10.0 mL and 3.0 g of the water sample to be tested, which has undergone pretreatment in step S01. Place the sample in a headspace vial, immediately seal it with clamps and shake well. Place the vial on the heating platform of the headspace sampler. Set the headspace conditions to the optimal equilibrium temperature and optimal equilibrium time determined in step S05. Simultaneously set the injection valve temperature to 100℃, the transfer line temperature to 100℃, and the injection volume to 1.0 mL to ensure complete transfer of the gas phase sample to the gas chromatography system. The gas chromatograph performs split injection according to the optimal split ratio parameters determined in step S04. The chromatographic column is the DB-WAX high-polarity column selected in step S02. The injection port temperature is set to 200℃, the detector temperature is set to 250℃, and the column temperature program is set to an initial temperature of 40℃, held for 5 min, then increased to 90℃ at a rate of 5℃ / min and held for 1 min. The carrier gas helium flow rate is set to 1.0 mL / min, and the hydrogen flow rate of the flame ionization detector is set to 30 mL / min. The air flow rate was set to 300 mL / min and the tail gas nitrogen flow rate was set to 25 mL / min. Chromatographic signals were collected and peak identification was performed on the chromatogram. The retention time and half-peak width of the methanol solvent peak and the benzene peak were measured, and the resolution parameter between the two peaks was calculated. When the resolution parameter was less than 1.0, it indicated that the methanol solvent peak interfered with the benzene peak. At this time, the optimal split ratio parameter was increased by 5 units to reset the split ratio, and the same water sample was re-analyzed by headspace gas chromatography. The adjustment was repeated until the resolution parameter was greater than or equal to 1.0. The peak areas of the eight benzene series compounds in the well-separated chromatogram were calculated by integrating the peaks. The mass concentration of each benzene series compound in the water sample was calculated by linear interpolation based on the calibration curve established in step S06. Step S07 achieved accurate determination of the actual water sample and ensured the reliability of the analysis results by dynamically adjusting the split ratio parameter.
[0057] The specific implementation of step S08 involves first performing a preliminary determination on the water sample to be tested. A preliminary chromatogram is obtained using the analytical procedure of step S07. The approximate concentration range of each benzene series compound in the water sample is estimated by comparing it with the calibration curve established in step S06, serving as the estimated concentration of the sample. For water samples with complex matrices containing oils or high-molecular-weight organic matter, the complex matrix components can alter the partition coefficient of benzene series compounds between the aqueous and gas phases, leading to systematic errors in the standard curve method. Therefore, a standard addition method is used to correct for matrix effects. Specifically, four equal volumes of the water sample to be tested are placed in four headspace vials. The first headspace vial is left unspecified as an unspiked sample. The second, third, and fourth headspace vials are filled with standard solutions at concentrations of 50%, 100%, and 150% of the estimated sample concentration, respectively. The volume of the standard solution added is controlled to be less than 5% of the total volume of the water sample to avoid significant changes in the matrix due to dilution effects. All four headspace vials are filled according to the method described in step S01. The pH value was adjusted to below 2, and then headspace gas chromatography was performed according to the analytical conditions in step S07 to obtain the peak area response values of the unspiked sample and the three spiked samples. A standard addition correction curve was established with the spiking amount as the abscissa and the peak area response value as the ordinate. The slope and intercept parameters of the regression line were obtained by linear regression fitting using the least squares method. The absolute value of the intercept of the regression line on the abscissa is the actual concentration of the target benzene series compounds in the water sample to be tested. The principle of the standard addition method is to make the standard substance and the target substance in the test sample undergo the same matrix environment and analytical process, thereby automatically compensating for the response deviation caused by the matrix effect. The matrix interference is eliminated by extrapolation to achieve accurate quantification. For the test water sample with particularly high oil content, the test water sample is diluted by 2 to 5 times with ultrapure water before the standard addition method to reduce the matrix interference intensity. Step S08 ensures the accuracy and reliability of the determination results of benzene series compounds in complex matrix water samples.
[0058] It should be noted that one of the key technical ideas of this invention is to use an independent component analysis blind source separation unit to decompose the mixed feature signals in chromatographic analysis at the source signal level. Traditional feature extraction methods based on principal component analysis or artificial feature engineering cannot effectively distinguish the aliasing phenomenon of methanol solvent interference and target analyte response in the feature space. This makes it difficult to balance the two mutually restrictive objectives of split ratio optimization decision-making, namely separation effect and detection sensitivity. However, independent component analysis reveals the inherent generation mechanism of observation data by maximizing the statistical independence of signal components. It decomposes the complex mixed observation signal into independent source signals with clear physical meaning in an unsupervised manner. This allows the interference correlation component to be used to accurately assess the degree of masking of the methanol peak to the benzene peak, while the target analyte correlation component purely reflects the detector response capability of benzene series compounds. This source signal level separation provides pure feature input for subsequent prediction branches to remove interference. It fundamentally solves the technical contradiction of increasing the split ratio to improve separation but at the same time reducing sensitivity. This allows the model to output the optimal split ratio parameter that maximizes the retention of the target analyte response while ensuring effective separation.
[0059] The second key technical approach involves introducing a Gaussian process regression Bayesian optimization algorithm to globally optimize headspace equilibrium conditions. Traditional grid search or single-factor optimization methods require a large number of experiments and are prone to getting trapped in local optima. Bayesian optimization treats parameter optimization as a sequential decision problem, constructing a probabilistic surrogate model of the objective function through Gaussian process regression. It continuously updates the posterior distribution using existing experimental data to predict the function value and its uncertainty in untested regions. It adopts an expectation-based improvement acquisition function to achieve a dynamic balance between exploring unknown high-uncertainty regions and utilizing known high-response regions. Each time, the most valuable experimental point is selected for testing, thereby approaching the global optimum with minimal experimental cost. This algorithm is particularly suitable for scenarios with high single-experiment costs, such as chromatographic condition optimization. It can find the equilibrium temperature and equilibrium time combination that makes eight different boiling point benzene compounds simultaneously reach the best response within 10 to 15 experiments, significantly improving the efficiency of method development and ensuring the global optimality of the optimization results.
[0060] The synergistic effect of the aforementioned key technologies lies in constructing an end-to-end intelligent optimization system from raw experimental data to optimal analytical conditions. The blind source separation unit of independent component analysis solves the core problem of split ratio optimization, namely, how to decouple interference factors from the target response at the feature level, providing accurate decision-making basis for the model. Meanwhile, the Gaussian process regression Bayesian optimization algorithm efficiently searches for the global optimal solution in the headspace condition parameter space. The combination of the two enables the entire detection method to achieve data-driven intelligent optimization in both the split ratio and headspace conditions, two key dimensions. Compared with traditional empirical trial-and-error methods or single parameter adjustment methods, the synergistic optimization strategy not only significantly reduces the number of experiments and time costs required for method development, but more importantly, it uses machine learning algorithms to mine the complex nonlinear relationships hidden in the experimental data, finding the optimal combination of parameters that is difficult to discover through human experience. Thus, while ensuring the effective separation of isomers such as p-xylene and m-xylene, it maximizes the detection sensitivity of low-boiling-point benzene and high-boiling-point styrene, expands the linear range and detection limit of the method, and improves the quantitative accuracy and robustness of trace benzene series compounds in complex matrix water samples.
[0061] It should be noted that this invention also solves the following technical problems: First, it addresses the problem of low efficiency in exploring headspace equilibrium parameters. Traditional methods use orthogonal or full-factor experiments to optimize equilibrium temperature and time, requiring numerous experiments and struggling to find the global optimum. This invention employs a Gaussian process regression Bayesian optimization algorithm. By constructing a probabilistic surrogate model of the objective function and utilizing the expectation to improve the acquisition function, it balances the utilization of known optimal solutions and the exploration of unknown regions. Within a limited number of experiments, it efficiently approximates the optimal parameter combination that maximizes the total peak area response of the eight benzene series compounds, significantly improving the method development efficiency. Second, it solves the problem of difficulty in accurately quantifying partially overlapping chromatographic peaks. Isomers such as p-xylene and m-xylene are prone to peak overlap under conventional chromatographic conditions, and traditional peak area integration methods have large errors. This invention uses a peak identification and deconvolution separation algorithm. Through wavelet denoising preprocessing and identifying potential peak positions using the zero-crossing point of the second derivative, it establishes a multi-peak superposition model and optimizes the model parameters using a nonlinear least squares algorithm. This achieves mathematical separation of overlapping peaks and accurate calculation of single-peak areas, improving the quantitative accuracy of benzene series isomers in complex samples.
[0062] Specifically, the principle of this invention is as follows: The fundamental reason why this invention can solve the above-mentioned technical problems lies in breaking through the limitations of traditional single optimization objectives and establishing a technical route for the collaborative optimization of separation degree and response intensity as dual objectives. The model input layer receives separation degree and peak area response data under different split ratio parameters. The multi-scale convolutional branch can capture the different time-scale features of the impact of split ratio changes on separation effect and sensitivity. The feature fusion layer integrates multi-scale information to form a comprehensive feature expression. The independent component analysis blind source separation unit, by maximizing the statistical independence criterion, utilizes the essential differences in statistical characteristics between methanol interference signals and benzene series response signals to decompose the mixed observation features into independent source signals with clear physical meaning. This source signal-level separation eliminates the influence of interference factors on response intensity evaluation and also eliminates the interference of target object signals on separation degree evaluation. Two parallel fully connected prediction branches perform predictions based on pure feature components, avoiding the phenomenon of loss of sensitivity due to improved separation degree in traditional methods. The multi-objective weighted fusion layer constructs a comprehensive evaluation index based on the relative importance of separation degree and response intensity, ensuring that the optimal split ratio parameter of the output maximizes the detection sensitivity while meeting the separation degree requirement. This collaborative optimization mechanism makes the selection of the split ratio parameter scientific and predictable.
[0063] The following provides a specific embodiment 1 of the present invention. The specific implementation methods of steps S01, S03, S06 and S07 in this embodiment 1 are the same as those described above, and will not be repeated in detail here. The specific implementation methods of other steps are described in detail below.
[0064] Separation parameters in step S02 The calculation formula is expressed as follows:
[0065] ;
[0066] In the formula, This is the resolution parameter, which is dimensionless. Retention time of the later elution peak, in minutes; Retention time of the first eluting peak, in minutes; The half-peak width of the first peak to emerge is expressed in min. This represents the half-width of the later peak, in minutes. A value greater than 1.5 indicates that the two chromatographic peaks have achieved baseline separation.
[0067] In the specific implementation of step S04, the comprehensive evaluation index of the diversion ratio optimization prediction model The calculation formula is expressed as follows:
[0068] ;
[0069] In the formula, This is a comprehensive evaluation indicator, dimensionless. The separation parameter is dimensionless and is used to predict the separation degree. The standard resolution threshold is 1.5, which is dimensionless. The response intensity parameter refers to the peak area response value of benzene, with units of mV·s; The standard response intensity threshold is set to the initial peak area response threshold. The unit is mV·s; 0.6 is the separation weight, dimensionless; 0.4 is the response intensity weight, dimensionless. When Greater than 1.2 and When the peak area response exceeds 40% of the initial peak area response threshold, the optimal split ratio parameter is used as the subsequent analysis condition. The initial peak area response threshold is the peak area response value of benzene obtained by measuring a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L at a split ratio parameter of 5:1, expressed in mV·s.
[0070] In the specific implementation of step S05, the Gaussian process regression Bayesian optimization algorithm is used for parameter optimization, and the squared exponential covariance function is used. The statement is as follows:
[0071] ;
[0072] In the formula, For input points and The covariance between them, expressed in mV·s squared; This is the signal variance parameter, expressed in squared mV·s. is the length scale parameter, with units of ℃·min, representing the characteristic scale of the combined temperature and time distance in the input space; and For each point in the input space, there are two dimensions: equilibrium temperature and equilibrium time. The unit for equilibrium temperature is °C and the unit for equilibrium time is min. The normalized Euclidean distance is calculated by dividing the temperature difference by the temperature range of 20°C and the time difference by the time range of 20 minutes, and is dimensionless. Improvements to the acquisition function are desired. The statement is as follows:
[0073] ;
[0074] In the formula, The desired improvement value is dimensionless. is the posterior mean of the Gaussian process, in mV·s; The value of the objective function is the currently known optimal value, expressed in mV·s. The cumulative distribution function of the standard normal distribution is dimensionless. represents the posterior standard deviation of the Gaussian process, in mV·s; is the probability density function of the standard normal distribution, which is dimensionless; The reference standard deviation is the average of the posterior standard deviations of all experimental points, expressed in mV·s. For standardized improvements, the quantity is dimensionless. Among them, For standardization improvement quantities, dimensionless, when hour, ,when hour, .
[0075] In the specific implementation of step S08, the linear regression equation of the standard addition method is expressed as follows:
[0076] ;
[0077] In the formula, The total peak area measured after spiking is expressed in mV·s. For reference peak area, the value is taken as the measured peak area of the unspecified water sample, in mV·s; The slope is dimensionless; The concentration of the added standard solution is expressed in μg / L. For reference concentration, the value is taken as the estimated concentration of the sample. The unit is μg / L; The original peak area of the target analyte in the water sample is given, in mV·s. The actual concentration of benzene series compounds in the water sample is also given. The intercept of the linear regression equation is calculated using the following formula:
[0078] ;
[0079] In the formula, This represents the actual concentration of benzene compounds in the water sample, in μg / L; other parameters have the same meaning as before. Estimated sample concentration. The concentration was estimated based on a calibration curve after preliminary measurements of the water sample under optimal split ratio, optimal equilibrium temperature, and optimal equilibrium time. The unit is μg / L. The low, medium, and high concentration levels are as follows: , , .
[0080] During the training of the split ratio optimization prediction model, the cross-entropy loss function The statement is as follows:
[0081] ;
[0082] In the formula, Cross-entropy loss, dimensionless; For batch size, the empirical value is 32, which is dimensionless; For the first The true feature importance label of each sample, with a value of 0 or 1, is dimensionless; For the first The predictive feature importance value for each sample, ranging from 0 to 1, is dimensionless; The normalization factor is dimensionless. Mean squared error loss function. The statement is as follows:
[0083] ;
[0084] In the formula, The mean squared error loss is dimensionless. The variance of the target value is taken as the variance of the separation parameter in the training set for separation prediction, and the variance of the peak area response value in the training set for response intensity prediction, in mV·s squared. For the first The true target value of each sample, in mV·s; For the first The predicted target value for each sample, in mV·s. Total loss function. The statement is as follows:
[0085] ;
[0086] In the formula, Total loss, dimensionless; The mean squared error loss for separation prediction is dimensionless. The mean square error loss for response intensity prediction is dimensionless. The separation loss weight is empirically set to 0.6, which is dimensionless. The response intensity loss weight is empirically set to 0.4, which is dimensionless. (Frobenius norm) The statement is as follows:
[0087] ;
[0088] In the formula, It is the Frobenius norm, dimensionless; For the first The separation matrix parameters for the next iteration; For the first The separation matrix parameters for the next iteration; and These are the number of rows and columns of the matrix, respectively, and are dimensionless. and The first Second and third The matrix of the nth iteration Line number Column elements, dimensionless. Iteration terminates when the norm is less than 0.0001 or the number of iterations reaches 200. Mean absolute error. The calculation formula is expressed as follows:
[0089] ;
[0090] In the formula, The mean absolute error is dimensionless. The number of samples in the test set is dimensionless. This is the average of the true values, in mV·s; For the first The true value of each sample, in mV·s; For the first Predicted values for each sample, in mV·s. Mean absolute percentage error. The calculation formula is expressed as follows:
[0091] ;
[0092] In the formula, This represents the mean absolute percentage error, expressed in %; other parameters have the same meaning as above.
[0093] In the blind source separation unit of independent component analysis, the covariance matrix of the fused feature matrix after centering is... The calculation formula is expressed as follows:
[0094] ;
[0095] In the formula, The covariance matrix is dimensionless. This is the fusion feature matrix before centralization, in mV·s; This is the mean vector of the fused feature matrix, in mV·s; This is the transpose of the centered feature matrix; The sample size is dimensionless. The reference variance for the fused features is taken as the variance of all elements in the centered feature matrix, expressed in mV·s squared. Covariance matrix Eigenvalue decomposition yields the eigenvalue diagonal matrix. and eigenvector matrix The fusion feature matrix after whitening transformation The statement is as follows:
[0096] ;
[0097] In the formula, is the whitened feature matrix, which is dimensionless; This is a diagonal matrix of eigenvalues of the covariance matrix, with the diagonal elements representing the eigenvalues. Dimensionless; The eigenvalue matrix is the negative 1 / 2 power of the eigenvalue matrix, with diagonal elements being... Dimensionless; The eigenvector matrix is dimensionless. This is the reciprocal of the reference standard deviation, expressed as the negative first power of mV·s. Separation vector. The update formula is expressed as follows:
[0098] ;
[0099] In the formula, For the first The separation vector of the next iteration is dimensionless; For the first The separation vector of the next iteration is dimensionless; The whitened feature vector is dimensionless. Since it is a nonlinear function, the hyperbolic tangent function is chosen. or cubic function Dimensionless; for The derivative of the hyperbolic tangent function For cubic functions Dimensionless; The mathematical expectation operator for the sample is calculated as follows: ,in For the first One whitened feature vector; The Euclidean norm is calculated as follows: ,in For vectors, For vector dimensions. Negative entropy. The statement is as follows:
[0100] ;
[0101] In the formula, It is negative entropy, dimensionless; The differential entropy of a homoscedastic Gaussian distribution is expressed in nat. For output signal The differential entropy, in units of nat; The reference entropy value is taken as the differential entropy of the standard normal distribution. The unit is nat.
[0102] The salting-out effect refers to the phenomenon where the solubility of nonpolar or weakly polar organic compounds in the aqueous phase decreases after the addition of inorganic salts to an aqueous solution, with more being transferred to the gas phase. The mechanism involves the hydration of water molecules by inorganic salt ions, reducing the number of free water molecules available to dissolve the organic compounds. Simultaneously, the presence of salt ions increases the ionic strength of the solution, disrupting the weak interactions between organic and water molecules. The salting-out effect leads to a decrease in the solubility of organic compounds in the aqueous phase. With salt concentration The relationship can be described by the Sechenov equation:
[0103] ;
[0104] In the formula, The solubility of the organic matter without added salt is expressed in mg / L. The solubility of the organic matter after salting is expressed in mg / L. is the Sechenov constant, which is dimensionless; This is the mass-volume ratio of salt, expressed in g / mL. For reference salt concentration, the empirical value is 1 g / mL.
[0105] The peak identification and deconvolution separation algorithm's processing flow involves: preprocessing the original chromatographic signal with wavelet denoising to remove baseline drift and high-frequency noise; calculating the first and second derivatives of the original chromatographic signal and identifying potential peak positions by combining the zero-crossing points of the second derivative with the sign change of the first derivative; using curvature analysis to evaluate the sharpness of each candidate peak and eliminating false peaks with curvature less than a curvature threshold; and establishing a multi-peak superposition model for the identified overlapping peak regions. The Gaussian function form is expressed as follows:
[0106] ;
[0107] In the formula, For the first Each single peak during retention time The signal strength at that location is expressed in mV. For the first The peak height of each peak is expressed in mV. Retention time, in minutes; For the first The peak position of each peak, in minutes; For the first The standard deviation of each peak, in min; The reference standard deviation is empirically defined as 1 minute. The exponentially corrected Gaussian function is expressed as follows:
[0108] ;
[0109] In the formula, For the first The asymmetry factor of each peak, in units of Other parameters have the same meaning as before. Multi-peak stacking model. The statement is as follows:
[0110] ;
[0111] In the formula, The sum of the signals is expressed in mV. The number of single peaks within the overlapping peak region is dimensionless. The Levenberg-Marquardt nonlinear least squares algorithm is used to optimize the parameters of the multi-peak stacking model, minimizing the sum of squared residuals between the fitted and measured curves. During the iteration process, physical constraints are applied to the multi-peak stacking model parameters to ensure that the peak height is positive and the peak width is within a reasonable range. After separating each single peak, numerical integration is performed on each single peak to calculate its area. The numerical integration employs either the trapezoidal rule or Simpson's rule.
[0112] To better understand and implement this invention, Example 2 of a specific application scenario is provided below: For an urban surface water environmental monitoring task, accurate quantitative analysis of benzene series pollutants in river water samples is required. Traditional gas chromatography methods encounter difficulties in selecting the split ratio parameter during the determination process. While a small split ratio improves sensitivity, the methanol solvent peak severely interferes with the benzene peak; while a large split ratio improves separation, it results in the inability to detect low-concentration samples. The technical team decided to use the optimized split ratio determination method of this invention to solve this technical problem.
[0113] The technical team first pretreated the samples. 100 mL of the river water sample was placed in a 250 mL Erlenmeyer flask, and 0.5 g of ascorbic acid was added as a reducing agent. Then, hydrochloric acid solution was added dropwise to adjust the pH to 1.8, ensuring an acidic environment to inhibit the loss of volatile organic compounds. 30 g of sodium chloride was added to the water sample at a mass-to-volume ratio of 0.3 g / mL. After thorough shaking to dissolve, the salting-out effect was utilized to enhance the transfer of benzene compounds from the aqueous phase to the gas phase. The pretreated water sample was then transferred to a 20 mL headspace vial and sealed for analysis.
[0114] During the column selection phase, the technical team prepared a mixed standard solution of eight benzene compounds (benzene, toluene, ethylbenzene, p-xylene, m-xylene, o-xylene, cumene, and styrene) at a concentration of 10.0 mg / L. Separation tests were performed using a DB-624 medium-polarity column and a DB-WAX strong-polarity column under the same chromatographic conditions. The resolution parameter for p-xylene and m-xylene on the DB-624 column was 1.32, while the resolution parameter for p-xylene and m-xylene on the DB-WAX column was 1.68. Based on the selection criterion of a resolution parameter greater than 1.5, the technical team chose the DB-WAX strong-polarity column as the column for subsequent analysis.
[0115] In the stage of establishing the split ratio response relationship dataset, the technical team selected six test points (5:1, 10:1, 15:1, 20:1, 25:1, and 30:1) within the split ratio range of 5:1 to 30:1. A mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L was prepared, and headspace gas chromatography analysis was performed at each split ratio parameter. Figure 2 As shown, there are significant differences in chromatographic separation performance under different split ratio parameters. The technical team recorded the resolution parameters of the methanol solvent peak and the benzene peak, as well as the peak area response values of the eight benzene series compounds, under each split ratio parameter. The specific data are shown in Table 1.
[0116] Table 1. Resolution parameters and peak area response values under different split ratio parameters.
[0117]
[0118] The technical team input the three-dimensional dataset of split ratio, separation parameter, and peak area response value from Table 1 into the split ratio optimization prediction model for processing. After receiving the data, the model input layer performs standardization, normalizing the separation parameter and peak area response value to the range of 0 to 1. The three parallel one-dimensional convolutional branches of the feature extraction layer extract feature patterns at different time scales using kernel sizes of 3, 5, and 7, respectively. Kernel size 3 captures locally rapidly changing features, kernel size 5 captures medium-scale trend features, and kernel size 7 captures globally slowly changing features. The feature fusion layer concatenates the outputs of the three convolutional branches into a 512-dimensional fused feature vector.
[0119] The independent component analysis (ICA) blind source separation unit reconstructs the fused feature vector into a fused feature matrix, centers it, and then performs whitening transformation using principal component analysis. The technical team initializes the separation matrix parameters as an identity matrix and updates them using a fast ICA iterative algorithm. During iteration, the hyperbolic tangent function is chosen as the nonlinear function, and the gradient ascent method is used to maximize the negative entropy of the output signal. After 87 iterations, the Frobenius norm of the separation matrix parameter changes between adjacent iterations decreases to 0.00008, satisfying the convergence condition. In the separated independent components, the interfering feature components exhibit time-domain waveform characteristics correlated with the methanol solvent peak, while the target feature components show an amplitude variation pattern positively correlated with the concentration of benzene series compounds. Figure 3 As shown, blind source separation effectively separates the aliased interference signal and the target response signal in the feature space.
[0120] Two parallel fully connected prediction branches receive interference and target feature components as inputs, respectively. The first fully connected prediction branch contains three fully connected layers with 256, 128, and 64 neurons, respectively, using the ReLU activation function, and has an output layer prediction separation parameter of 1.28. The second fully connected prediction branch also contains three fully connected layers with 256, 128, and 64 neurons, respectively, and has an output layer prediction response intensity parameter of 0.52. The multi-target weighted fusion layer calculates a comprehensive evaluation index based on a separation weight of 0.6 and a response intensity weight of 0.4. The standard separation threshold is set to 1.2, and the standard response intensity threshold is set to 0.45, resulting in a comprehensive evaluation index of 0.706. The model outputs an optimal split ratio parameter of 16:1 based on the principle of maximizing the comprehensive evaluation index. At this point, the prediction separation parameter of 1.28 is greater than the requirement of 1.2, and the peak area response value of benzene (171400) accounts for 60% of the initial peak area response value of 285600, which is greater than the requirement of 40%, satisfying the conditions for subsequent analysis.
[0121] In the headspace equilibrium condition optimization phase, the technical team set the initial equilibrium temperature to 60℃ and the initial equilibrium time to 30 min. A Gaussian process regression Bayesian optimization algorithm was used to optimize the parameters within the equilibrium temperature range of 50℃ to 70℃ and the equilibrium time range of 20 min to 40 min. The Gaussian process regression model selected the squared exponential covariance function, with the initial length scale parameter set to 5.0 and the initial signal variance parameter set to 1.0. By maximizing the marginal likelihood function, the length scale parameter was estimated to be 4.2 and the signal variance parameter to be 0.87. The aim was to improve the acquisition function by identifying candidate points with high predictive values and high uncertainty in the parameter space. After 15 iterations, the optimal equilibrium temperature was found to be 63℃ and the optimal equilibrium time to be 35 min, maximizing the total peak area response of the eight benzene series compounds. At this point, the total peak area response reached 1,486,500.
[0122] The technical team established a series of mixed standard working solutions with concentration gradients of 50 μg / L, 200 μg / L, 500 μg / L, 2000 μg / L, 4000 μg / L, and 5000 μg / L, totaling six levels. Each concentration level was measured in triplicate under optimal split ratio (16:1) and optimal equilibrium conditions. Due to partial overlap between p-xylene and m-xylene in the chromatograms, the team employed peak identification and deconvolution separation algorithms. First, the raw chromatographic signal underwent wavelet denoising preprocessing using the Daubechies wavelet function, with a decomposition layer of 5 layers to remove baseline drift and high-frequency noise. The first and second derivatives of the chromatographic signal were calculated, and 18 potential peak positions were identified by combining the zero-crossing points of the second derivative with the sign change of the first derivative. Curvature analysis was used to evaluate the sharpness of each candidate peak, with a curvature threshold of 0.05, eliminating four spurious peaks with curvatures less than the threshold. A multi-peak stacking model was established for the identified overlapping peak regions. It was assumed that each single peak conformed to an exponentially modified Gaussian function, and the model parameters included peak height, peak position, peak width, and asymmetry factor. The Levenberg-Marquardt nonlinear least squares algorithm was used to optimize the model parameters. After 42 iterations, the sum of squared residuals converged to a minimum value of 0.0023. After separating each single peak, numerical integration was performed on each single peak, and the peak area was calculated using Simpson's rule. Figure 4 As shown, the overlapping peaks were separated by deconvolution to obtain clear single-peak results. The technical team established calibration curves for eight benzene series compounds, and the linear correlation coefficients were all greater than 0.998, indicating that the calibration curves had a good linear relationship.
[0123] In the actual sample determination phase, the technical team performed gas chromatography analysis on pretreated river water samples according to the optimal split ratio of 16:1 and optimal equilibrium conditions. During the determination, the resolution parameter between the methanol solvent peak and the benzene peak was detected to be 1.31, meeting the requirement of greater than 1.0, thus requiring no adjustment of the split ratio parameter. Based on the peak area response values and calibration curves, the calculated mass concentrations of benzene, toluene, ethylbenzene, p-xylene, m-xylene, o-xylene, cumene, and styrene in the water sample were 126 μg / L, 238 μg / L, 94 μg / L, 167 μg / L, 143 μg / L, 178 μg / L, 52 μg / L, and 68 μg / L, respectively.
[0124] During the matrix effect correction phase, the technical team discovered that some river water samples contained high-molecular-weight organic matter, leading to matrix interference. For these samples, a standard addition method was used to establish a correction curve. First, preliminary measurements were performed on the water samples to estimate the predicted concentration. Then, standard solutions at low, medium, and high concentration levels were added to the samples. For example, with benzene as an example, the predicted concentration was 120 μg / L, and the three concentration levels were set at 60 μg / L, 120 μg / L, and 180 μg / L, respectively. Measurements were then performed on the water samples after adding the standard solutions, and the peak area response values were recorded. A linear regression equation was established between the added spiking amount and the response value. The intercept of the linear regression equation was used to calculate the actual concentration of benzene in the water samples, which was found to be 132 μg / L. The corrected result eliminated the influence of the matrix effect.
[0125] Compared to traditional gas chromatography methods that rely on empirical exploration or single-factor experiments to select split ratio parameters, this invention uses a split ratio optimization prediction model (structure as follows) Figure 5(As shown) The model achieves intelligent prediction of the optimal split ratio parameter. It employs independent component analysis (ICA) blind source separation technology to effectively separate the aliased methanol interference signal and benzene series response signal in the feature space, overcoming the shortcomings of traditional feature extraction methods that misclassify interference components as valid signals or submerge valid signals in the interference background. The separated interference feature components and target feature components are input into two parallel prediction branches for prediction, avoiding the loss of sensitivity due to improved separation in traditional methods. The multi-objective weighted fusion layer constructs a comprehensive evaluation index based on the relative importance of separation degree and response intensity, ensuring that the output optimal split ratio parameter retains the target response value to the maximum extent while ensuring effective separation of methanol and benzene. The Gaussian process regression Bayesian optimization algorithm constructs a probabilistic surrogate model of the objective function and uses expectation-improved acquisition function to balance the utilization of known optimal solutions and the exploration of unknown regions, efficiently approximating the optimal equilibrium condition within a limited number of experiments. The peak identification and deconvolution separation algorithm establishes a multi-peak superposition model and uses a nonlinear least squares algorithm to optimize model parameters, achieving mathematical separation of overlapping peaks and accurate calculation of single-peak area. This invention establishes a complete technical route from optimizing the split ratio parameter and finding the equilibrium condition to separating overlapping peaks, providing a scientific and effective solution for the accurate quantitative analysis of trace benzene compounds in environmental water samples.
[0126] It should be noted that the variables involved in this invention are explained in detail in Tables 2 and 3.
[0127] Table 2. Variable Explanation Table (Part 1)
[0128]
[0129] Table 3. Variable Explanation Table (Part Two)
[0130]
[0131] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A gas chromatographic method for determining benzene series compounds in water with optimized split ratio, characterized in that, The process includes: adjusting pH and adding sodium chloride in the sample pretreatment stage; selecting columns that meet the resolution parameter requirements in the column screening stage; establishing a three-dimensional dataset of split ratio, resolution parameter, and peak area response value; inputting the three-dimensional dataset into the split ratio optimization prediction model; the split ratio optimization prediction model extracts features through multi-scale convolution and then separates the fused feature vector into interference feature components and target feature components through an independent component analysis blind source separation unit; inputting the target feature components into two parallel fully connected prediction branches to predict the resolution parameter and response intensity parameter; generating a comprehensive evaluation index and outputting the optimal split ratio parameter through a multi-objective weighted fusion layer; using a Gaussian process regression Bayesian optimization algorithm to determine the optimal equilibrium temperature and optimal equilibrium time in the headspace equilibrium condition optimization stage; establishing a calibration curve; performing analysis according to the optimal split ratio parameter and optimal equilibrium conditions in the actual sample determination stage; and establishing a calibration curve using the standard addition method in the matrix effect correction stage. The optimal split ratio parameter is output as follows: the input layer of the split ratio optimization prediction model standardizes the three-dimensional dataset, normalizing the numerical range to between 0 and 1. The standardized data enters the feature extraction layer. The three parallel one-dimensional convolutional branches of the feature extraction layer use convolutional kernels with kernel sizes of 3, 5, and 7 to perform convolution operations on the input data to extract local and global feature patterns. The feature fusion layer concatenates the feature vectors output by the three one-dimensional convolutional branches along the feature dimension to form a fused feature vector. The fused feature vector is input to the independent component analysis blind source separation unit. The independent component analysis blind source separation unit decomposes the mixed feature signal into interference feature components and target feature components based on the statistical independence assumption. After separation... The two feature components are respectively input into the first fully connected prediction branch and the second fully connected prediction branch. The first fully connected prediction branch outputs the prediction separation parameter, and the second fully connected prediction branch outputs the response intensity parameter. The multi-objective weighted fusion layer calculates the comprehensive evaluation index by weighting the prediction separation parameter and the response intensity parameter with weight coefficients of 0.6 and 0.
4. Finally, the output layer selects the split ratio parameter that makes the comprehensive evaluation index reach its maximum value as the optimal split ratio parameter, and outputs the prediction separation parameter corresponding to the optimal split ratio parameter. When the prediction separation parameter is greater than 1.2 and the peak area response value of benzene is greater than 40% of the initial peak area response threshold, it is confirmed that the optimal split ratio parameter meets the separation requirements and sensitivity requirements. Among them, the interference feature components exhibit time-domain waveform characteristics related to the methanol solvent peak, while the target feature components show an amplitude variation pattern positively correlated with the concentration of benzene series compounds. Specifically, the independent component analysis blind source separation unit treats the fused feature vector as a mixed representation of multiple independent source signals, finds the unmixing matrix by maximizing the non-Gaussianity criterion, optimizes the separation matrix parameters using a fast independent component analysis iterative algorithm, and separates the interference feature components and the target feature components from the fused feature vector. Specifically, the calculation of the comprehensive evaluation index generated by the multi-objective weighted fusion layer is as follows: the comprehensive evaluation index is equal to the weighted sum of the quotient of the separation degree parameter divided by the standard separation degree threshold and the quotient of the response intensity parameter divided by the standard response intensity threshold. The separation degree weight is set to 0.6 and the response intensity weight is set to 0.
4.
2. The method according to claim 1, characterized in that, The sample pretreatment stage specifically involves adding ascorbic acid and hydrochloric acid solutions to the water sample to be tested to adjust the pH value to below 2, and then adding sodium chloride at a mass-volume ratio of 0.3 g / mL and dissolving it completely.
3. The method according to claim 2, characterized in that, The column screening stage specifically involves using a DB-624 medium polarity column and a DB-WAX strong polarity column to separate and test the benzene series standard solutions, recording the resolution parameters of p-xylene and m-xylene, and selecting columns with a resolution parameter greater than 1.5 as analytical columns.
4. The method according to claim 3, characterized in that, The steps to establish a three-dimensional dataset of split ratio-resolution parameter-peak area response value are as follows: select a test point every 5 units within the split ratio range of 5:1 to 30:1, perform headspace gas chromatography analysis on a mixed standard solution of benzene series compounds in methanol with a mass concentration of 1.0 mg / L, and record the resolution parameter between the methanol solvent peak and the benzene peak, as well as the peak area response values of the eight benzene series compounds.
5. The method according to claim 4, characterized in that, In the headspace equilibrium condition optimization stage, the initial equilibrium temperature is set to 60℃ and the initial equilibrium time is set to 30min. The Gaussian process regression Bayesian optimization algorithm is used to optimize the parameters in the range of equilibrium temperature from 50℃ to 70℃ and equilibrium time from 20min to 40min, with the objective function value being the maximization of the total peak area response value of the eight benzene series compounds.
6. The method according to claim 5, characterized in that, The application of the Gaussian process regression Bayesian optimization algorithm involves selecting the squared exponential covariance function to describe the correlation between different points in the input space, using the expectation improvement acquisition function to select the next experimental point, executing the experiment and updating the Gaussian process posterior distribution until the optimal equilibrium temperature and optimal equilibrium time are found.
Citation Information
Patent Citations
Detecting method of Benzene content
CN102495144A
Three-dimensional fluorescence spectrum benzene series identification method combining spectral line fitting and total variation denoising
CN120064233A