Rubber vulcanization accelerator quantitative detection method based on terahertz spectrum and data expansion strategy
By expanding the data and using the GA-SVR model, the problems of cumbersome traditional detection methods and model overfitting were solved, enabling efficient and environmentally friendly quantitative detection of rubber vulcanization accelerators and improving detection accuracy and speed.
Patent Information
- Application Number
- CN202610028680.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional detection methods are cumbersome and polluting. Terahertz spectroscopy is prone to model overfitting and mixed spectra overlap due to small samples, making it difficult to distinguish. Existing technologies cannot achieve efficient and environmentally friendly quantitative detection of rubber vulcanization accelerators.
A data augmentation strategy was adopted, combining the GA-SVR model and feature extraction. The sample size was expanded by least-squares Gaussian fitting and data fusion, and the data dimensionality was reduced by using the VISSA algorithm. Finally, an SVR model optimized by a genetic algorithm was constructed for detection.
It effectively expands small-sample terahertz spectral data, solves the problem of resolving spectral overlap in mixtures, has high detection accuracy, is environmentally friendly and fast, and has a single detection time of ≤5 minutes.
Smart Images

Figure CN121762487A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of terahertz spectroscopy detection technology, specifically to a quantitative detection method for rubber vulcanization accelerators. Background Technology
[0002] Vulcanization accelerators are key additives in rubber products, but their toxicity and pollution must be strictly controlled. Traditional detection methods (such as chromatography and chemical analysis) have drawbacks such as cumbersome operation, high pollution, and low efficiency; although terahertz spectroscopy has "fingerprint" characteristics, small sample spectral data can easily lead to model overfitting, and the spectra of mixtures overlap significantly, making direct differentiation difficult. Summary of the Invention
[0003] The purpose of this invention is to provide an efficient and environmentally friendly method for the quantitative detection of rubber vulcanization accelerators, which achieves accurate detection through data expansion, GA-SVR model and feature extraction. Detailed Implementation
[0004] This invention provides a quantitative detection method for rubber vulcanization accelerators based on terahertz spectroscopy and data augmentation strategies. It aims to solve the problems of cumbersome operation and high pollution associated with traditional detection methods, as well as the tendency for small samples in terahertz spectroscopy to lead to model overfitting and the difficulty in distinguishing overlapping spectra of mixtures. The specific steps are as follows:
[0005] S1: Sample preparation and spectral acquisition: Rubber mixture samples containing vulcanization accelerators were prepared according to the specified ratio. Time-domain spectra were acquired using a terahertz spectrometer, and the humidity of the sample chamber was controlled to be < 3% to reduce water vapor interference.
[0006] S2: Optical parameter extraction: The time-domain spectrum is Fourier transformed to obtain the frequency-domain spectrum, and the absorbance spectrum is calculated based on the formula.
[0007] S3: Data augmentation: The sample size is expanded through two strategies: least squares Gaussian fitting and data fusion, to solve the problem of small sample size.
[0008] S4: Model Construction and Optimization: The GA-SVR model was used to analyze the spectral data, and the model parameters were optimized using a genetic algorithm. The VISSA algorithm was used to extract features, reducing data dimensionality and analysis time.
[0009] S5: Quantitative detection: Through model training and prediction, quantitative analysis of vulcanization accelerator content is achieved.
[0010] S6: Analysis and Evaluation: The correlation coefficient (R) and root mean square error (RMSE) are used as model evaluation indicators. Attached Figure Description
[0011] Figure 1 This is a physical image of the terahertz spectrometer of the present invention.
[0012] Figure 2 This is a schematic diagram of the terahertz spectrometer of the present invention.
[0013] Figure 3 These are the average time-domain spectra of five substances in a specific embodiment of the present invention.
[0014] Figure 4 This is the average time-domain spectrum of a mixture sample according to a specific embodiment of the present invention.
[0015] Figure 5 This is the average absorbance spectrum of five substances in a specific embodiment of the present invention.
[0016] Figure 6 This is the average absorbance spectrum of a mixture sample according to a specific embodiment of the present invention.
[0017] Figure 7 This is a comparison chart of the original spectrum and the Gaussian model fitted spectrum of a specific embodiment of the present invention.
[0018] Figure 8 This is an extended data graph based on the data fusion method of a specific embodiment of the present invention.
[0019] Figure 9 This is a distribution map of the importance of data features in a specific embodiment of the present invention.
[0020] Figure 10 This is a diagram showing the model analysis results before feature extraction in a specific embodiment of the present invention.
[0021] Figure 11 This is a diagram showing the model analysis results after feature extraction according to a specific embodiment of the present invention.
[0022] The beneficial effects of this invention are as follows:
[0023] (1) It effectively expands small-sample terahertz spectral data and avoids model overfitting;
[0024] (2) It solves the problem of resolving spectral overlap in mixtures, resulting in high detection accuracy;
[0025] (3) No chemical reagents are required, the detection process is environmentally friendly and fast, and the single detection time is ≤5min;
[0026] (4) It can provide technical support for the optimization of rubber formulations and help the industry achieve green and sustainable development.
[0027] This invention provides a specific embodiment.
[0028] Nitrile butadiene rubber (NBR), silica, antioxidant N-phenyl-N'-(1,3-dimethylbutyl)-p-phenylenediamine (44S), toxic vulcanization accelerator ethyl thiourea (ETU), and non-toxic vulcanization accelerator tetrabenzyl thiuram disulfide (TBzTD) were selected as raw materials for the mixture, while polyethylene was selected as a sample preparation aid.
[0029] The above raw materials were crushed separately using a high-speed pulverizer. The crushed raw material powder was then sieved using a 200-mesh sieve to select powder with uniform particle size for later use.
[0030] Weigh the raw material powders according to the preset mass ratio, where nitrile rubber accounts for 60%, silica for 20%, antioxidant 44S for 5%, and polyethylene for 5%. The total mass fraction of vulcanization accelerators ETU and TBzTD is 10%, and they are mixed in gradient ratios (10:0, 8:2, 6:4, 4:6, 2:8, 0:10), preparing a total of 6 different mixtures. Pour the weighed powders into a ceramic mortar and grind them thoroughly until uniformly mixed. Please refer to Table 1 for specific experimental parameters.
[0031]
[0032] The uniformly mixed powder was poured into a mold and pressed into a smooth, circular sample with a diameter of 13 mm and a thickness of 1 mm. Simultaneously, samples of five pure substances, namely nitrile rubber, silica, antioxidant 44S, ETU, and TBzTD, were prepared. 5% polyethylene was also added during the preparation of the pure substance samples to improve the tableting effect.
[0033] All samples were placed in a constant temperature drying oven and dried for 1.5 hours before being taken out for testing. 36 samples were prepared for each pure substance and each mixture in each proportion, for a total of 396 samples.
[0034] A terahertz time-domain spectrometer with a spectral range of 0.05–5 THz was used. This instrument is equipped with a femtosecond pulsed laser (pulse repetition frequency 80 MHz, center wavelength 780 nm, pulse width less than 100 fs), a terahertz radiation generator, a detector, and a delay device. To eliminate the absorption interference of water vapor in the air on the terahertz waves, the sample was placed in a sealed sample chamber, and dry nitrogen gas was continuously introduced into the chamber to maintain the relative humidity below 3%. A picture of the terahertz spectrometer is shown below. Figure 1 The schematic diagram of a terahertz spectrometer is as follows: Figure 2 .
[0035] Acquire the terahertz time-domain spectrum under no-load conditions as the reference signal E ref(t) Place the sample to be tested into the sample chamber and collect the terahertz time-domain spectrum of the sample as the sample signal E. sam (t); To reduce random errors, spectral data were collected three times on both sides of each sample, and the average value was taken as the final time-domain spectral data for that sample; 180 time-domain spectral data points were collected for five pure substances and 216 time-domain spectral data points for six proportional mixtures. The average time-domain spectra of the five substances are shown below. Figure 3 Average time-domain spectrum of the mixture sample Figure 4 .
[0036] For reference signal E ref (t) and sample signal E sam (t) Perform Fast Fourier Transform (FFT) on each signal to convert the time-domain signal into the frequency-domain signal E. ref (w) and E sam (w).
[0037] Based on the optical parameter extraction model, the terahertz absorbance A of the sample is calculated using the formula (1): (1)
[0038] In the formula, ω is the angular frequency, and Eref(w) and Esam(w) are the amplitudes of the sample and reference frequency domain signals, respectively. Absorbance data in the 0.2-2.0 THz frequency band were selected as the basis for subsequent analysis to avoid signal saturation and noise interference above 2.0 THz. The average absorbance spectra of the five substances are shown below. Figure 5 The average absorbance spectrum of the mixture sample Figure 6 .
[0039] To address the issue of overfitting in terahertz spectra due to small sample sizes, two strategies—least square Gaussian fitting and data fusion—were employed to augment the absorbance spectral data.
[0040] The least squares Gaussian fitting method is used to augment the absorbance spectral data:
[0041] A Gaussian function fitting model was constructed, using the original absorbance spectral data as the fitting object. The Gaussian function parameters were optimized using the least squares method to minimize the sum of squared errors between the fitted curve and the original spectrum. The average of the two fitted curves was taken to generate new expanded spectral data. The original spectral data and the expanded data were merged to achieve a 1.5-fold increase in sample size.
[0042] Verify the validity of the augmented data: Compare the absorption peak positions and intensities of the original spectrum and the fitted spectrum to ensure that no spurious features are introduced into the fitted spectrum. A comparison of the original spectrum and the Gaussian model-fitted spectrum is shown in the figure below. Figure 7 .
[0043] Data fusion methods are used to augment absorbance spectral data:
[0044] The Savitzky-Golay algorithm was used to smooth the original absorbance spectral data, removing spectral noise while retaining the core characteristic peaks of the vulcanization accelerator. The preprocessed spectral data was then fused with the original absorbance spectral data according to sample dimensions, ensuring a one-to-one correspondence between original and preprocessed samples to form a new fused dataset. This fusion method doubled the sample size while maintaining the same data dimensions as the original spectra. The data fusion method for expanding the data is as follows: Figure 8 .
[0045] The Kennard-Stone (KS) algorithm was used to partition the original spectral dataset, the least-squares Gaussian fitting augmented dataset, and the data fusion augmented dataset. The specific steps are as follows:
[0046] The Euclidean distance between any two samples in the dataset is calculated using the formula (2):
[0047] (2)
[0048] In the formula x p and x q This represents two different samples, where N represents... Total number of samples .
[0049] Train the model using three types of data before feature extraction:
[0050] The datasets are divided into a calibration set (for model training) and a test set (for model validation) in a 2:1 ratio.
[0051] To address the issues of high dimensionality and time-consuming model analysis after expansion, the Variable Space Iterative Shrinkage (VISSA) algorithm is used to extract a subset of spectral features. The steps are as follows:
[0052] Step 1: First, standardize the three sets of data according to formula (3) to obtain the initial data, denoted as S={x1,x2···x 230}
[0053] (3)
[0054] Step 2: Set the number of iterations and conditions, and perform the following operations: Randomly select a certain number of variables from the current variable set. Based on the selected variables, use PLS to construct a regression model and calculate the importance VIP value of the variables. The calculation formula is shown in equation (4). Sort the variables according to the VIP value and retain the variables whose VIP value is greater than the set threshold.
[0055] (4)
[0056] in H is the total number of variables, H is the number of principal components in the PLS model, and SS is the total number of variables. h The variance contribution of each principal component, W 2 hj This is the square of the weight of the j-th variable in the h-th principal component. Set VIP=0.5 and retain variables with a value greater than this threshold until the feature importance no longer changes (to 0 or 1).
[0057] Step 3: Based on the conditions in Step 2, m features are retained, denoted as S′={x1,x2···x}. m}, which constitute the spectral data X after feature extraction.
[0058] Step 4: Evaluate the model's performance through cross-validation to ensure the effectiveness of feature extraction. The coefficients of the PLS regression model are used to assess the importance of variables, and their calculation formula is shown in equation (5):
[0059] (5)
[0060] Where X is the input variable matrix and y is the target variable vector.
[0061] The model performance is evaluated by the root mean square error of cross-validation, and the calculation formula is shown in equation (6):
[0062] (6)
[0063] In the formula, observed t The actual value, predicted t These are predicted values.
[0064] The distribution of feature importance after feature extraction using the VISSA algorithm on the original spectral data, data augmented by data fusion, and data augmented by least squares fitting are shown in the figure. Figure 9 .
[0065] After VISSA feature extraction, 66 features were extracted from the original spectral data, reducing the dimensionality by 71.30%; 91 features were extracted from the data fusion augmented dataset, reducing the dimensionality by 60.43%; and 54 features were extracted from the least squares Gaussian fitting augmented dataset, reducing the dimensionality by 76.52%.
[0066] Using the feature-extracted absorbance spectral data as input variables and the contents of vulcanization accelerators ETU and TBzTD as output variables, an SVR regression model was constructed. The model introduced an insensitive loss function and slack variables to handle nonlinear regression problems.
[0067] To address the difficulty in determining the penalty coefficient C and kernel width γ in the SVR model, a genetic algorithm is used for global optimization. The steps are as follows:
[0068] Encoding: C and γ are used as optimization parameters and binary encoded to generate the initial population;
[0069] Fitness function: The reciprocal of the root mean square error (RMSE) between the model's predicted values and the actual values is used as the fitness function to evaluate the quality of individual models.
[0070] Genetic manipulation: Selecting, crossover, and mutation operations are performed sequentially on the population to select individuals with high fitness;
[0071] Iteration Termination: Repeat the genetic operation until the preset number of iterations is reached, and output the optimal C and γ parameters.
[0072] The calibration set data is used to train the model, and the test set data is used to verify the model performance. The correlation coefficient (R) and root mean square error (RMSE) are used as model evaluation metrics, as shown in the following formulas:
[0073]
[0074]
[0075] When evaluating the performance of a prediction model, Rc (training set) and Rp (prediction set) are key parameters. Rc and Rp values approaching 1 indicate a better fit between the model and the calibration and prediction sets. Similarly, RMSEC and RMSEP values closer to zero signify stronger modeling performance and predictive ability.
[0076] The model analysis results before feature extraction are shown in the figure below. Figure 10 Model analysis results after feature extraction Figure 11 .
[0077] To more intuitively differentiate the model's effectiveness in analyzing different data, the parameters involved in the model analysis process before and after feature extraction were summarized. Table 2 shows the relevant parameters for model analysis before feature extraction, and Table 3 shows the relevant parameters for model analysis after feature extraction.
[0078]
[0079]
[0080] Comparing Tables 2 and 3, it is evident that feature extraction improved the analytical performance of the GA-SVM model across all three datasets. When analyzing the original spectral data, the parameter Rp showed the largest increase, while RMSEP decreased the most. When analyzing data augmented by data fusion and least squares Gaussian fitting, the changes in Rp and RMSEP were smaller. This indicates that the original spectral data contained numerous negative or useless features, affecting the model's analytical performance. It also suggests that the data fitted by data fusion and least squares Gaussian fitting can, to some extent, mitigate the impact of negative or useless features. Among these, the GA-SVM model performed best in analyzing data augmented by least squares Gaussian fitting, both before and after feature extraction. This demonstrates that least squares Gaussian fitting effectively introduces new features characterizing sample information through fitting the original spectra.
[0081] In summary, this example presents a novel method for rapidly detecting the vulcanization accelerator content in rubber mixtures using terahertz time-domain spectroscopy, data augmentation strategies, and chemometrics. This provides a fast, accurate, safe, environmentally friendly, and energy-efficient reference method for optimizing rubber product formulations and assessing their safety.
[0082] It should be noted that the above specific embodiments are only for detailed explanation of the technical solution of the present invention, and should not be used to limit the scope of the present invention. Those skilled in the art should understand all and part of the process of the above embodiments, and equivalent substitutions made in accordance with the claims of the present invention still fall within the scope of the present invention.
Claims
1. A quantitative detection method for rubber vulcanization accelerators based on terahertz spectroscopy and data augmentation strategy, characterized in that, Includes the following steps: S1: Sample preparation and spectral acquisition: Rubber mixture samples containing vulcanization accelerators were prepared according to the preset mass ratio. The time-domain spectral data of the samples were acquired using a terahertz time-domain spectrometer. During the acquisition process, the relative humidity of the sample chamber was controlled to be <3%. S2: Optical parameter extraction: The collected time-domain spectral data is subjected to fast Fourier transform to convert the time-domain signal into a frequency-domain signal. Then, the absorbance spectrum of the sample is calculated based on the optical parameter extraction model, and the absorbance data in the 0.2~2.0THz frequency band is selected as the data for subsequent analysis. S3: Data augmentation: The absorbance spectral data are augmented using the least squares Gaussian fitting method and / or data fusion method to obtain an augmented dataset; S4: Dataset partitioning: The Kennard-Stone algorithm was used to partition the original absorbance spectrum dataset and the augmented dataset into a calibration set and a test set at a ratio of 2:1, respectively. S5: Feature Extraction: The Variable Space Iterative Shrinkage (VISSA) algorithm is used to extract features from each dataset, and the feature variables with VIP values > 0.5 are selected to obtain the optimal feature subset; S6: Model Construction and Optimization: Using the optimal feature subset as the input variable and the sulfurization accelerator content as the output variable, a support vector regression (SVR) model is constructed. A genetic algorithm is used to globally optimize the penalty coefficient C and kernel function width γ of the SVR model to obtain the GA-SVR quantitative analysis model. S7: Quantitative detection and evaluation: Input the calibration set data into the GA-SVR quantitative analysis model for training, use the test set data to verify the model performance, and use the correlation coefficient (R) and root mean square error (RMSE) as model evaluation indicators to achieve quantitative detection of rubber vulcanization accelerator content.
2. The method according to claim 1, characterized in that, The raw materials of the rubber mixture sample mentioned in step S1 include nitrile rubber (NBR), silica, antioxidant N-phenyl-N'-(1,3-dimethylbutyl)-p-phenylenediamine (44S), vulcanization accelerator and polyethylene. The mass ratio of each raw material is: 60% nitrile rubber, 20% silica, 5% antioxidant 44S, and 5% polyethylene. The vulcanization accelerator is ethyl thiourea (ETU) and tetrabenzyl thiuram disulfide (TBzTD), with a total mass fraction of 10%, and they are blended in gradient ratios of 10:0, 8:2, 6:4, 4:6, 2:8 and 0:
10.
3. The method according to claim 1, characterized in that, The specific process of sample preparation in step S1 is as follows: each raw material is crushed and sieved through a 200-mesh sieve, weighed according to the ratio, and then thoroughly ground and mixed evenly. The mixture is poured into a mold and pressed into a round sample with a smooth surface, a diameter of 13 mm, and a thickness of 1 mm. The sample is dried for 1.5 hours before testing. 36 sample pieces are prepared for each pure substance and each proportion of the mixture.
4. The method according to claim 1, characterized in that, The terahertz time-domain spectrometer mentioned in step S1 has a spectral range of 0.05~5THz and is equipped with a femtosecond pulsed laser with a pulse repetition frequency of 80MHz, a center wavelength of 780nm, and a pulse width of less than 100fs. During spectral acquisition, the time-domain spectrum under no-load conditions is first acquired as a reference signal E(t), and then the time-domain spectrum of the sample is acquired as the sample signal E(t). Each sample is acquired 3 times on both sides, and the average value is taken as the final time-domain spectral data.
5. The method according to claim 1, characterized in that, The formula for calculating the absorbance spectrum in step S2 is as follows: In the formula, A(ω) is the absorbance, ω is the angular frequency, E(ω) is the amplitude of the sample's frequency domain signal, and E(ω) is the amplitude of the reference signal's frequency domain signal.
6. The method according to claim 1, characterized in that, The specific process of the least squares Gaussian fitting method described in step S3 is as follows: using the original absorbance spectral data as the fitting object, a Gaussian function fitting model is constructed. The model parameters are optimized by the least squares method to minimize the sum of squared errors between the fitted curve and the original spectrum. The average value of the two fitted curves is taken to generate expanded data, which is then merged with the original data to achieve a 1.5-fold expansion of the sample size. Furthermore, the absorption peak positions and intensities of the fitted spectrum and the original spectrum are matched, and no false features are introduced.
7. The method according to claim 1, characterized in that, The specific process of the data fusion method described in step S3 is as follows: the Savitzky-Golay algorithm is used to perform smooth preprocessing on the original absorbance spectral data, and the preprocessed data is fused with the original absorbance spectral data according to the sample dimensions to form a fused dataset, thereby expanding the sample size by 2 times, and the dimensions of the fused data are consistent with the original spectral data.
8. The method according to claim 1, characterized in that, In step S4, when the Kennard-Stone algorithm partitions the dataset, the sample distribution is determined by calculating the Euclidean distance between any two samples in the dataset. The formula for calculating the Euclidean distance is: In the formula, d is the Euclidean distance between the p-th sample and the q-th sample, x and x' are the absorbance values of the p-th and q-th samples at the i-th wavelength, respectively, and N is the total number of samples.
9. The method according to claim 1, characterized in that, The specific process of VISSA algorithm for feature extraction in step S5 is as follows: First, the dataset is standardized, then variables are selected iteratively to construct a partial least squares (PLS) regression model and calculate the VIP value of the variables. Variables with VIP values > 0.5 are retained until the number of features in the feature set is stable. After feature extraction, the dimensionality of the original spectral data was reduced by 71.30%, the dimensionality of the data fusion-enlarged dataset was reduced by 60.43%, and the dimensionality of the least squares Gaussian fitting-enlarged dataset was reduced by 76.52%.
10. The method according to claim 1, characterized in that, The specific process of genetic algorithm optimization in step S6 is as follows: the penalty coefficient C and the kernel function width γ are binary encoded to generate an initial population. The reciprocal of the root mean square error (RMSE) between the model prediction value and the true value is used as the fitness function. The optimal individuals are iteratively screened through selection, crossover, and mutation genetic operations until the preset number of iterations is reached, and the optimal C and γ parameters are output. When the GA-SVR quantitative analysis model analyzes the least squares Gaussian fitting augmented dataset, the correlation coefficient Rp ≥ 0.9826 and the root mean square error RMSEP ≤ 0.0023 on the test set.