Rapid detection method for raffinate concentration parameter in hydrogen peroxide production process

By screening spectral wavelengths using Lasso and MBPCA and combining them with the random forest algorithm, a dual prediction model was constructed, which solved the problems of speed, accuracy, and stability in detecting residual concentration during hydrogen peroxide production. This model is suitable for rapid detection of residual concentration during hydrogen peroxide production.

CN121935592APending Publication Date: 2026-04-28CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA PETROLEUM & CHEMICAL CORP
Filing Date
2024-10-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In the existing technology, the detection methods for residual concentration in the hydrogen peroxide production process are cumbersome, slow, and prone to human error, making it difficult to achieve rapid, accurate, and stable analysis.

Method used

The Lasso algorithm was used to screen spectral wavelengths, and a dual prediction model was constructed by combining multi-block principal component analysis (MBPCA) and random forest algorithm. A method for rapid detection of residual concentration was established by combining spectral intensity variables and principal component scores.

Benefits of technology

This method enables rapid and accurate detection of raffinate concentration, improves the stability and repeatability of the method, is suitable for online analysis, reduces manual processing, and lowers operational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935592A_ABST
    Figure CN121935592A_ABST
Patent Text Reader

Abstract

The invention relates to the field of spectrum detection, and discloses a rapid detection method for raffinate concentration parameters in a hydrogen peroxide production process, and the method comprises the following steps: 1) preparing a standard raffinate; 2) obtaining the spectrum of the raffinate; 3) screening out a plurality of wavelength points through a Lasso algorithm; 4) taking the spectral intensity as an input variable I1, taking the raffinate concentration as an output variable, and adopting a random forest algorithm to construct a model M1; 5) adopting an MBPCA algorithm to obtain a global principal component score and a block score of the spectrum; (6) combining the variable I1 and the variable in the step (5) as an input variable I2, taking the raffinate concentration as an output variable, and constructing a model M2 by adopting a random forest algorithm; (7) obtaining a variable I1 of raffinate to be detected, substituting the variable I1 into the model M1 to obtain a combined variable I2, and substituting the combined variable I2 into the model M2 to obtain raffinate concentrations N1 and N2; and 8) taking the average value of N1 and N2. The method disclosed by the invention can be used for rapidly and accurately detecting the raffinate concentration with high repeatability and good stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spectral detection, specifically to a rapid detection method for the raffinate concentration parameter during hydrogen peroxide production. Background Technology

[0002] The anthraquinone process is currently the mainstream process for producing hydrogen peroxide. Its raw materials, intermediate products, and final products are almost all flammable, explosive, or combustion-supporting substances. Therefore, it is particularly necessary to accurately and promptly monitor the quality indicators of the hydrogen peroxide production process in order to provide early warnings for production safety.

[0003] Raffinate is the liquid remaining after hydrogen peroxide is extracted from the oxidation reaction during the anthraquinone process for hydrogen peroxide production. Its main components are anthraquinone and solvent, with a small amount of hydrogen peroxide also present. The concentration of hydrogen peroxide is the raffinate concentration. Raffinate concentration is an important analytical item in the anthraquinone process for hydrogen peroxide production, reflecting the quality of the hydrogen peroxide and closely related to production safety. The method for determining raffinate concentration generally involves taking a portion of the sample, adding potassium permanganate titrant, and calculating the corresponding raffinate concentration (i.e., the actual hydrogen peroxide content in the sample) based on the amount of titrant consumed at the titration endpoint, expressed in g / L. To ensure the quality indicators of hydrogen peroxide production and meet safety requirements, the analysis interval for raffinate concentration is very short, averaging every 2-4 hours. However, the currently mainstream manual titration method is cumbersome, slow, and requires highly skilled operators. High-frequency titration analysis is too costly in terms of manpower and resources and is prone to human error.

[0004] Therefore, it is essential to develop a method that can accurately, quickly, and stably obtain the raffinate concentration. Summary of the Invention

[0005] The purpose of this invention is to overcome the above-mentioned problems existing in the prior art and provide a rapid detection method for the raffinate concentration parameter in the hydrogen peroxide production process. This method is particularly suitable for the detection of raffinate concentration and is fast, accurate, highly repeatable and stable, with advantages in operation and accuracy. In addition, this method is applicable to both offline and online analysis.

[0006] To achieve the above objectives, the present invention provides a rapid detection method for the raffinate concentration parameter during hydrogen peroxide production, the method comprising:

[0007] (1) Obtain multiple standard raffinates with known and different raffinate concentrations;

[0008] (2) Obtain the spectrum of the standard raffinate and select the characteristic spectral bands;

[0009] (3) Using the minimum absolute shrinkage and selection algorithm, multiple wavelength points are selected from the above characteristic spectral bands;

[0010] (4) The spectral intensity variable I1 corresponding to the selected wavelength point is used as the input variable, and the corresponding raffinate concentration obtained in step (1) is used as the output variable. The random forest algorithm is used to construct the prediction model M1 between the input variable and the output variable.

[0011] (5) Use the multi-block principal component analysis algorithm to obtain the global principal component score of each characteristic spectral band; divide each characteristic spectral band into segments, and use the multi-block principal component analysis method to obtain the principal component score of each segment after segmentation, which is the block score of each segment after segmentation.

[0012] (6) Combine the spectral intensity variable I1 corresponding to the wavelength points selected in step (3) with the global principal component score and block score of each characteristic spectral band as variable I2, and use it as input variable. Use the corresponding raffinate concentration obtained in step (1) as output variable. Use the random forest algorithm to construct the prediction model M2 between the input variable and the output variable.

[0013] (7) Obtain the raffinate solution with unknown raffinate concentration, obtain the spectral intensity variable I1 corresponding to the screening wavelength point according to the method in steps (2) and (3), and substitute it into the prediction model M1 obtained in step (4) to obtain the raffinate concentration N1; obtain the corresponding combination variable I2 according to the method in steps (2), (3), (5) and (6), and substitute it into the prediction model M2 obtained in step (6) to obtain the raffinate concentration N2;

[0014] (8) Take the average value of the raffinate concentration N1 and the raffinate concentration N2 to obtain the final raffinate concentration of the raffinate to be tested.

[0015] The above technical solution enables rapid and accurate determination of the raffinate concentration. More importantly, the method's stability and reproducibility are further enhanced by averaging the results from two models, making it more suitable for scenarios requiring long-term stable operation, such as online analysis. Compared to existing technologies, the method provided by this invention avoids complex manual processing, is more efficient, greatly improves the analytical operating environment, and is applicable to both offline and online analysis, demonstrating significant advantages. Attached Figure Description

[0016] Figure 1 This is a correlation diagram of the measured values ​​and calculated values ​​obtained in Embodiment 1 of the present invention. Detailed Implementation

[0017] The endpoints and any values ​​of the ranges disclosed herein are not limited to the precise ranges or values, and these ranges or values ​​should be understood to include values ​​close to these ranges or values. For numerical ranges, the endpoint values ​​of the various ranges, the endpoint values ​​of the various ranges and individual point values, and individual point values ​​can be combined with each other to obtain one or more new numerical ranges, which should be considered as specifically disclosed herein.

[0018] This invention provides a rapid detection method for raffinate concentration parameters during hydrogen peroxide production, the method comprising:

[0019] (1) Obtain multiple standard raffinates with known and different raffinate concentrations;

[0020] (2) Obtain the spectrum of the standard raffinate and select the characteristic spectral bands;

[0021] (3) Using the minimum absolute shrinkage and selection algorithm, multiple wavelength points are selected from the above characteristic spectral bands;

[0022] (4) The spectral intensity variable I1 corresponding to the selected wavelength point is used as the input variable, and the corresponding raffinate concentration obtained in step (1) is used as the output variable. The random forest algorithm is used to construct the prediction model M1 between the input variable and the output variable.

[0023] (5) Use the multi-block principal component analysis algorithm to obtain the global principal component score of each characteristic spectral band; divide each characteristic spectral band into segments, and use the multi-block principal component analysis method to obtain the principal component score of each segment after segmentation, which is the block score of each segment after segmentation.

[0024] (6) Combine the spectral intensity variable I1 corresponding to the wavelength points selected in step (3) with the global principal component score and block score of each characteristic spectral band as variable I2, and use it as input variable. Use the corresponding raffinate concentration obtained in step (1) as output variable. Use the random forest algorithm to construct the prediction model M2 between the input variable and the output variable.

[0025] (7) Obtain the raffinate solution with unknown raffinate concentration, obtain the spectral intensity variable I1 corresponding to the screening wavelength point according to the method in steps (2) and (3), and substitute it into the prediction model M1 obtained in step (4) to obtain the raffinate concentration N1; obtain the corresponding combination variable I2 according to the method in steps (2), (3), (5) and (6), and substitute it into the prediction model M2 obtained in step (6) to obtain the raffinate concentration N2;

[0026] (8) Take the average value of the raffinate concentration N1 and the raffinate concentration N2 to obtain the final raffinate concentration of the raffinate to be tested.

[0027] The anthraquinone process is currently the mainstream hydrogen peroxide production process. It's understandable that the main reaction solutions involved generally include hydrogenation solution, oxidation solution, and raffinate. The main process of the anthraquinone process includes: preparing a working solution by mixing anthraquinone (generally 2-ethylanthraquinone) with an organic solvent (such as a mixture of C9-C10 heavy aromatics, trioctyl phosphate, and tetrabutylurea). Hydrogenation stage: Under pressure of 0.3 MPa or above, temperature of 40-80℃ and catalyst (such as Pd catalyst), anthraquinone in the working solution is hydrogenated and reduced by H2 to obtain anthraquinone or tetrahydroanthraquinone; Oxidation stage: The material after hydrogenation stage is oxidized with O2 under conditions of 30-60℃ and slight compression, so that anthraquinone and tetrahydroanthraquinone are oxidized to generate H2O2 and anthraquinone; Then, the material after oxidation stage is subjected to extraction (raffinate is the solution remaining after extracting H2O2), regeneration, purification and concentration to obtain a 20wt%-50wt% H2O2 aqueous solution.

[0028] The raffinate concentration refers to the concentration of hydrogen peroxide in the remaining liquid after hydrogen peroxide extraction, and this concentration is generally low. The inventors of this invention discovered that processing the raffinate from the anthraquinone process using the above method, particularly by using Lasso (Minimum Absolute Contraction) to screen variables and Multi-Block Principal Component Analysis (MBPCA) to extract global scores and block scores, can extract as much information related to the raffinate concentration from the spectrum as possible. This also largely eliminates the spectral influence of free water in the raffinate, resulting in higher analytical accuracy. Furthermore, by using Lasso to screen variables and directly combining them with a nonlinear regression random forest algorithm, and by combining variables extracted through Lasso and MBPCA with a random forest algorithm to establish a dual prediction model, and then averaging the predictions from the two models to determine the final analytical result, this method ensures accuracy while further improving the overall stability of the analysis when calculating unknown raffinate concentrations. This is particularly suitable for scenarios requiring long-term stable use, such as online analysis.

[0029] According to the present invention, preferably, in step (1), the number of standard raffinates is not less than 60 (for example, it can be 110, 120, 130, 180, 200, 220, 250, or any two of the above values ​​within a range). It is understood that the larger the sample size, the better the accuracy and stability. However, considering cost and operational feasibility, the number of standard raffinates is preferably 100-200. The inventors of the present invention have found that within the above range, the accuracy and stability of the results obtained by the prediction model can be further guaranteed.

[0030] According to the present invention, preferably, in step (1), the raffinate concentration of the standard raffinate is 0.03-1.2 g / L, more preferably 0.05-0.8 g / L (for example, it can be 0.05, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and values ​​within any two of the above values). It is understood that when selecting samples, in order to make the obtained model more accurate, the sample conditions generally need to be uniformly distributed within the possible range. Therefore, the raffinate concentration of the standard raffinate is preferably also uniformly distributed within the above range. For example, the difference between any two adjacent raffinate concentration values ​​is preferably in the range of 0.008-0.1 g / L.

[0031] According to the present invention, preferably, in step (2), the spectrum is a near-infrared spectrum.

[0032] Regarding the specific method for obtaining near-infrared spectra, this invention does not impose any particular limitations and can proceed according to conventional methods in the art. According to a preferred embodiment of this invention, the conditions for obtaining near-infrared spectra include: a temperature of 12-62℃, preferably 20-30℃ (for example, it can be 20, 22, 25, 28, 30, or any two of the above values ​​within a range); and a wavenumber range of 3500-12000 cm⁻¹. -1 Resolution is 2-16cm -1 Preferably, it is 6-12 (for example, it can be 6, 7, 8, 9, 10, 11, 12, or any two of the above values ​​within a range) cm. -1 The number of scans can be 64-128. Using the near-infrared spectroscopy testing conditions specified in this invention, clearer spectral results and more accurate spectral information can be obtained.

[0033] This invention is applicable to both offline and online near-infrared spectral acquisition methods in the laboratory. The instruments used for acquiring near-infrared spectra are not particularly limited and can be conventional choices in the field.

[0034] According to the present invention, preferably, the wavenumber range of the characteristic spectral band corresponding to the near-infrared spectrum is 4556-7596 cm⁻¹. -1 The inventors of this invention have further discovered that by selecting the above-mentioned characteristic spectral bands, the information related to the characteristic spectrum and the residual concentration can be more fully reflected.

[0035] According to the present invention, preferably, before performing step (3), the method further includes: preprocessing the characteristic spectral bands to reduce redundant information and / or noise in the spectrum.

[0036] According to the present invention, preferably, the pretreatment method is selected from second-order derivatives or first-order derivatives, more preferably first-order derivatives with a window width of 13-25 (e.g., 13, 15, 17, 19, 21, 23, 25, or any two of the above values ​​within a range). The inventors of the present invention have further discovered that using the above pretreatment method can better coordinate with subsequent MBPCA, ensuring sufficient information is obtained even if the raffinate concentration is low, thus further guaranteeing the accuracy of the results.

[0037] According to the present invention, in step (3), the number of wavelength points selected accounts for 2%-8% of the total number of wavelength points in the characteristic wavelength segment, preferably 3%-5%. By using the wavelength point range specified in the present invention, it is possible to further maximize the inclusion of effective spectral information while minimizing interference.

[0038] According to the present invention, the Lasso algorithm is selected to screen wavelength points in the feature spectrum. It is understood that the Lasso algorithm is a commonly used feature variable extraction method. Its basic principle is to introduce a norm penalty term into the least squares regression estimation, that is, to add an L1 regularization term. The L1 norm is generally used to calculate the sum of absolute errors between two vectors. Its essence is to calculate the absolute value. Therefore, the L1 norm has a natural advantage in sparse solution, which compresses the regression coefficients of some variables that do not contribute much to the model to 0, thereby removing useless information features and achieving the purpose of sparsification and feature selection. The Lasso algorithm is a common algorithm. Those skilled in the art are clear about its principle. For example, it can be seen in the explanation in Tibshirani R. Regression Shrinkage and Selection via the Lasso[J]. Journal of the Royal Statistical Society. Series B (Methodological), 1996, 58(1):267-288.

[0039] The formula for calculating the Lasso regression coefficients is as follows:

[0040]

[0041] Where t≥0 is the constraint constant; λ is the regularization parameter, also known as the penalty coefficient. i=1,2,…n, where n is the number of samples, j=1,2,…m, where m is the number of wavelength points in the spectrum. As λ increases, the optimal solution… The number of terms will decrease, and the coefficients of some independent variables will be compressed to 0, thereby achieving dimensionality reduction of high-dimensional or multivariate data, which can better solve many problems in modeling high-dimensional or multivariate data.

[0042] According to the present invention, in step (5), the number of segments of the characteristic spectrum is 2-4 segments, preferably 3 segments.

[0043] According to the present invention, more preferably, in step (5), the characteristic spectrum segmentation has 3 segments, and the wavenumber range of the first segment is 4556-5476 cm⁻¹. -1 The wavenumber range for the second segment is 5476-6396 cm⁻¹. -1 The wavenumber range for the third segment is 6396-7596 cm⁻¹. -1 .

[0044] According to the present invention, in step (5), the MBPCA algorithm is used to obtain the global principal component score of the characteristic spectral band and the principal component score of each segment after segmentation of the characteristic spectral band. It is understood that MBPCA is a commonly used data processing method through dimensionality reduction. However, the inventors of the present invention have found that, for raffinate, the above method, combined with the characteristic spectral band, segmentation and random forest algorithm as described above, can accurately determine the raffinate concentration in the raffinate. The MBPCA method is a common method, and those skilled in the art are clear about its principle. For example, it can be seen in the explanation in Westerhuis JA, Kourti T, MacGregor J F. Analysis of multiblock and hierarchical PCA and PLSmodels[J]. Journal of Chemometrics: A Journal of the Chemometrics Society, 1998, 12(5):301-321. The general principle of MBPCA is as follows:

[0045] For the B-block spectral matrix, block X 1 Block X 2 ... block X B The matrix X is formed by X = [X 1 X 2 …X B MBPCA is decomposed in two steps according to the following formula:

[0046] Full matrix decomposition: X = TP T +E

[0047] Block matrix decomposition: X b =T b P bT +E b b = 1, 2, ..., B

[0048] In the formula, T and P are the global score matrix and the load matrix, respectively. b P bThese are block matrices X b The score matrix and load matrix.

[0049] The steps of the MBPCA algorithm are as follows:

[0050] (1) Initialize the global score vector t

[0051] (2) Calculate the block score vector t b and block load vector p b b = 1, 2, ..., B

[0052] p b =X bT / t T t

[0053] p b =p b / ||p b ||

[0054] t b =X b p b

[0055] (3) Calculate the block global vector t and weight w, b = 1, 2, ..., B

[0056] T = [t] 1 t 2 …t B ]

[0057] w = T T t / t T t

[0058] w = w / ||w||

[0059] t = Tw

[0060] (4) Return to (2) and iterate to calculate t until convergence, to obtain t1, w1, t2 of the first principal component. b and p1 b

[0061] (5) Calculate the residual matrix X2 for each block. b b = 1, 2, ..., B, X2 b =X b -(tt T / t T t)X b

[0062] (6) Use X2 b Alternative X b Return to steps (1)-(4) and calculate t2, w2, and t2 of the second principal component. b and p2 b

[0063] (7) Following the above loop, calculate the t of all A principal components in sequence. i w i t i b and p i b , i = 1, 2, ..., A.

[0064] Generally, the number of global principal component scores for each characteristic spectral band and the number of block scores for each segmented band can each be independently greater than 6, for example, 8-12.

[0065] According to the present invention, the Random Forest (RF) algorithm mentioned in steps (4) and (6) is a fusion classification algorithm that includes many decision trees and voting strategies. It belongs to the ensemble algorithm and is also a natural nonlinear modeling tool that can be used for classification or regression analysis. The inventors of the present invention have found in their research that using the Random Forest algorithm to construct the prediction model of interest in the present invention has high accuracy, good tolerance to outliers and noise, and is not prone to overfitting. According to a preferred embodiment of the present invention, especially when combined with the MBPCA algorithm mentioned above, the model obtained can also ensure high accuracy for raffinate with generally low raffinate concentration. Regarding the Random Forest algorithm itself, it is a common algorithm, and those skilled in the art are aware of its principle. The above algorithm has been open-sourced and applied. For details, please refer to the description in Fang Kuangnan; Wu Jianbin; Zhu Jianping; Xie Bangchang. A Review of Random Forest Method Research [J]. Statistics and Information Forum, 2011, 26(03):32-38. The algorithm can be implemented in MATLAB.

[0066] In the Random Forest (RF) algorithm, the parameters ntree (generally the number of times the sample is trained) and mtry (generally the number of features selected in each sample training) can be controlled. In this invention, ntree can be set to 100-500, and the mtry value can be set to between 1 / 2 and 1 / 4 of the number of input variables.

[0067] According to a preferred embodiment of the present invention, a dual prediction model is established by organically combining two feature variable extraction methods, Lasso and MBPCA, with the random forest method. Specifically, a random forest algorithm that directly combines Lasso-screened variables with nonlinear regression is used, and a combination of variables extracted by Lasso screening and MBPCA is combined with the random forest algorithm to establish a dual prediction model. The final analysis result is then determined by averaging the prediction results of the two models. This ensures that the obtained prediction model has good accuracy in determining the raffinate concentration in the sample, while also exhibiting good stability and repeatability. It is particularly suitable for scenarios requiring long-term stable use, such as online analysis.

[0068] It is understandable that in steps (7) and (8), the raffinate solution with unknown raffinate concentration is obtained, and the spectrum and characteristic spectral bands are obtained in the manner of steps (2)-(6) (with corresponding preprocessing and segmentation), and the corresponding variables of the characteristic spectral bands are obtained. They are substituted into the corresponding prediction model, and the prediction results are averaged to obtain the final raffinate concentration result.

[0069] The present invention will be described in detail below through embodiments.

[0070] In the following examples, the raffinate was obtained during the anthraquinone process for producing hydrogen peroxide.

[0071] In step (1) of the following embodiments, the raffinate concentration of the standard raffinate is determined by potassium permanganate titration, and the raffinate concentration value obtained in this way is the measured value.

[0072] The instrument used to collect near-infrared spectra is a Fourier transform near-infrared spectrometer.

[0073] The operations following spectral acquisition are performed in the professional computing software MATLAB.

[0074] Example 1

[0075] (1) 200 raffinates were obtained and their corresponding measured values ​​were obtained. Their measured values ​​were distributed in the range of 0.06-0.51 g / L, and the difference between two adjacent raffinate concentrations was in the range of 0.008-0.1 g / L.

[0076] (2) Inject the samples into the sample cell (optical path length 2 mm) and perform spectral acquisition. The conditions for acquiring near-infrared spectra include: temperature 25℃, and wavenumber range of 3500-10000 cm⁻¹. -1 The resolution is 8cm. -1 The scan was performed 128 times. For the obtained near-infrared spectrum, the characteristic spectral band 4556-7596 cm⁻¹ was selected. -1 .

[0077] (3) Perform first-order differential preprocessing on the characteristic spectral bands obtained in step (2) and set the window width to 21.

[0078] Using the Lasso algorithm, 34 characteristic wavelength points were selected from the characteristic spectral bands after the above processing.

[0079] (4) Using the random forest algorithm, the spectral intensity variables corresponding to the 34 characteristic wavelength points selected by each sample through step (3) are used as input variables, and the measured values ​​of the raffinate concentration corresponding to each sample are used as output variables. ntree is set to 400 and mtry is set to 11 to construct the prediction model M1 between the input variables and the output variables.

[0080] (5) Multi-block principal component analysis is used for the characteristic spectral bands after preprocessing in step (3). The number of global principal component scores for the characteristic spectral bands is 9.

[0081] The characteristic spectral bands were then divided into three segments as follows: the wavenumber range of the first segment is 4556-5476 cm⁻¹. -1 The wavenumber range for the second segment is 5476-6396 cm⁻¹. -1 The wavenumber range for the third segment is 6396-7596 cm⁻¹. -1 Using multi-block principal component analysis, nine block scores are obtained for each segmented band, resulting in a total of 27 block scores for the three segments. Therefore, for each characteristic spectral band, there are nine global principal component scores and 27 block scores, for a total of 36 scores.

[0082] (6) Using the random forest algorithm, the spectral intensity variables corresponding to the 34 characteristic wavelength points selected by step (3) and the 36 score variables obtained by step (5) are combined as input variables, and the measured value of the raffinate concentration corresponding to each sample is used as output variables. ntree is set to 400 and mtry is set to 23 to construct the prediction model M2 between the input variables and the output variables.

[0083] (7) Take 39 raffinates with unknown raffinate concentrations as a validation set. According to the methods in steps (2)-(3), obtain the spectral intensity variable I1 corresponding to the screening wavelength point, and substitute it into the prediction model M1 obtained in step (4) to obtain the raffinate concentration N1. According to the methods in steps (2), (3), (5) and (6), obtain the corresponding combination variable I2, and substitute it into the prediction model M2 obtained in step (6) to obtain the raffinate concentration N2.

[0084] (8) Take the average of the raffinate concentrations N1 and N2 of each sample in the validation set to obtain the corresponding raffinate concentration of each sample (the calculated value).

[0085] To verify the accuracy of the prediction model, the measured values ​​of each of the 39 samples in the validation set were also obtained.

[0086] Table 1 shows the measured and calculated values ​​of the 39 samples in the validation set, as well as the deviation between the two (calculated value minus measured value). Furthermore, the performance of the model was evaluated using the root mean square error of prediction (RMSEP) and the correlation coefficient (R).

[0087]

[0088] Where n is the total number of validation set samples, y i Let be the measured value of the i-th sample. This is the calculated value for the i-th sample.

[0089]

[0090] Where n is the total number of validation set samples, y i Let be the measured value of the i-th sample. Let be the calculated value for the i-th sample. This represents the average of the measured values ​​of the validation set samples. The calculated R value is 0.990.

[0091] Table 1

[0092]

[0093]

[0094] Furthermore, a correlation plot was created using the measured values ​​(x-axis) and calculated values ​​(y-axis) of the 39 samples in the validation set, as shown in the figure. Figure 1 (where R is the correlation coefficient).

[0095] It is evident that the method provided by this invention can quickly and accurately determine the raffinate concentration in the anthraquinone process for producing hydrogen peroxide, and the prediction results for various concentrations are more accurate, demonstrating higher precision.

[0096] Example 2

[0097] Used to evaluate the repeatability of the method provided by this invention.

[0098] Two samples of the raffinate to be tested were taken, and the near-infrared spectra of each sample were measured three times in accordance with step (2) of Example 1. The three near-infrared spectra were processed as follows: characteristic spectral bands were obtained in accordance with the method in Example 1, and then Lasso screening variables and combination variables were obtained in accordance with the method in Example 1. The raffinate concentration (calculated value) was obtained by substituting it into the prediction model obtained in Example 1. Thus, three results were obtained for each sample in three parallel calculations, as shown in Table 2. The relative standard deviation was calculated as follows:

[0099]

[0100] Where S represents the standard deviation of the calculated value. The x represents the average of the calculated values ​​of the sample, i = 1, 2, ..., n, where n represents the number of samples. i This represents the calculated value of the i-th sample.

[0101] Table 2

[0102]

[0103] As shown in Table 2, the relative standard deviation of the three parallel calculations of the raffinate concentration for different samples using the method of the present invention is all below 2.0%. Therefore, the method provided by the present invention has good repeatability in determining the raffinate concentration of the sample to be tested. As explained above, the method provided by the present invention has high repeatability.

[0104] Example 3

[0105] Used to evaluate the long-term stability of the method provided by this invention.

[0106] Following step (2) in Example 1, near-infrared online analysis of the raffinate from a hydrogen peroxide production plant was continuously sampled for 20 minutes (from 1:30 AM to 1:50 AM on a certain day), obtaining near-infrared spectra of 19 samples. The spectra were then processed as follows: characteristic spectral bands were obtained as described in Example 1, and Lasso screening and combination variables were obtained as described in Example 1. These variables were then substituted into the prediction model obtained in Example 1 to obtain the raffinate concentration (calculated value) for each sample, as shown in Table 3.

[0107] Table 3

[0108] Online sampling time (h:min:s) Residue concentration calculated value (g / L) 01:30:03 0.10 01:31:12 0.14 01:32:20 0.09 01:33:29 0.10 01:34:37 0.09 01:35:46 0.09 01:36:55 0.12 01:38:05 0.14 01:39:14 0.14 01:40:23 0.13 01:41:33 0.09 01:42:41 0.14 01:43:50 0.14 01:44:58 0.14 01:46:06 0.14 01:47:15 0.14 01:48:24 0.14 01:49:32 0.14 01:50:41 0.14

[0109] As shown in Table 3, when the method of the present invention is used for online analysis of actual industrial samples, the analysis results can accurately reflect the stability and slight fluctuations of the material properties, and no obvious outliers are found. The results are consistent with the actual working conditions, indicating that the method of the present invention has high stability for continuous use.

[0110] Comparative Example 1

[0111] The method is the same as in Example 1, except that the random forest algorithm is replaced with a partial least squares algorithm to build the model. The results are shown in Table 4.

[0112] Table 4

[0113]

[0114]

[0115] By comparing Example 1 and Comparative Example 1, it can be seen that the raffinate concentration obtained by using the prediction model established in this invention is closer to the measured value, indicating that the prediction model constructed in this invention has outstanding effects in terms of accuracy and stability.

[0116] Although not shown, the inventors of this invention have found that averaging is more accurate and stable than using prediction model M1 and prediction model M2 alone.

[0117] The preferred embodiments of the present invention have been described in detail above; however, the present invention is not limited thereto. Within the scope of the inventive concept, various simple modifications can be made to the technical solutions of the present invention, including combinations of various technical features in any other suitable manner. These simple modifications and combinations should also be considered as the content disclosed in the present invention and are all within the protection scope of the present invention.

Claims

1. A rapid detection method for raffinate concentration parameters during hydrogen peroxide production, characterized in that, The method includes: (1) Obtain multiple standard raffinates with known and different raffinate concentrations; (2) Obtain the spectrum of the standard raffinate and select the characteristic spectral bands; (3) Using the minimum absolute shrinkage and selection algorithm, multiple wavelength points are selected from the above characteristic spectral bands; (4) The spectral intensity variable I1 corresponding to the selected wavelength point is used as the input variable, and the corresponding raffinate concentration obtained in step (1) is used as the output variable. The random forest algorithm is used to construct the prediction model M1 between the input variable and the output variable. (5) Use the multi-block principal component analysis algorithm to obtain the global principal component score of each characteristic spectral band; divide each characteristic spectral band into segments, and use the multi-block principal component analysis method to obtain the principal component score of each segment after segmentation, which is the block score of each segment after segmentation. (6) Combine the spectral intensity variable I1 corresponding to the wavelength points selected in step (3) with the global principal component score and block score of each characteristic spectral band as variable I2, and use it as input variable. Use the corresponding raffinate concentration obtained in step (1) as output variable. Use the random forest algorithm to construct the prediction model M2 between the input variable and the output variable. (7) Obtain the raffinate solution with unknown raffinate concentration, obtain the spectral intensity variable I1 corresponding to the screening wavelength point according to the method in steps (2) and (3), and substitute it into the prediction model M1 obtained in step (4) to obtain the raffinate concentration N1; obtain the corresponding combination variable I2 according to the method in steps (2), (3), (5) and (6), and substitute it into the prediction model M2 obtained in step (6) to obtain the raffinate concentration N2; (8) Take the average value of the raffinate concentration N1 and the raffinate concentration N2 to obtain the final raffinate concentration of the raffinate to be tested.

2. The method according to claim 1, wherein, In step (1), the number of standard raffinates shall not be less than 60.

3. The method according to claim 1 or 2, wherein, In step (1), the raffinate concentration of the standard raffinate is 0.03-1.2 g / L, preferably 0.05-0.8 g / L.

4. The method according to claim 1, wherein, In step (2), the spectrum is a near-infrared spectrum.

5. The method according to claim 4, wherein, The conditions for the near-infrared spectroscopy include: a temperature of 12-62℃ and a wavenumber range of 3500-12000 cm⁻¹. -1 Resolution is 2-16cm -1 ; Preferably, the conditions for the near-infrared spectroscopy include: a temperature of 20-30℃ and a resolution of 6-12 cm⁻¹. -1 ; And / or, the wavenumber range of the characteristic spectral band of the near-infrared spectrum is 4556-7596 cm⁻¹. -1 .

6. The method according to claim 4 or 5, wherein, Before performing step (3), the method further includes: preprocessing the characteristic spectral bands to reduce redundant information and / or noise in the spectrum.

7. The method according to claim 6, wherein, The preprocessing method is selected from second-order differential or first-order differential, and more preferably first-order differential with a window width of 13-25.

8. The method according to claim 1 or 5, wherein, In step (3), the number of wavelength points selected accounts for 2-8% of the total number of wavelength points in the characteristic wavelength segment, preferably 3-5%.

9. The method according to claim 1 or 5, wherein, In step (5), the number of segments in the characteristic spectrum segmentation is 2-4 segments, preferably 3 segments.

10. The method according to claim 9, wherein, In step (5), the characteristic spectrum is segmented into 3 segments, and the wavenumber range of the first segment is 4556-5476 cm⁻¹. -1 The wavenumber range for the second segment is 5476-6396 cm⁻¹. -1 The wavenumber range for the third segment is 6396-7596 cm⁻¹. -1 .