Fusion decision-making system design based on pesticide Raman spectrum

By constructing a Raman spectral database of carbamate pesticides and combining it with a fusion decision system of multiple machine learning algorithms, the problem of low accuracy in identifying pesticides with similar chemical structures was solved, and high-precision pesticide detection was achieved.

CN121237267APending Publication Date: 2025-12-30JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511361701.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-23
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing technologies suffer from difficulties in identifying pesticides with similar chemical structures using Raman spectra, and the accuracy is low. This is especially true for the detection of carbamate pesticides, where machine learning models lack versatility.

Method used

A Raman spectral database of carbamate pesticides was constructed using data augmentation methods, and dimensionality was reduced using principal component analysis (PCA). A fusion decision system was built by combining three machine learning algorithms: random forest (RF), support vector machine (SVM), and backpropagation neural network (BP neural network). The model output was integrated using a soft voting method to determine the final probability.

Benefits of technology

It achieves an accuracy of 65.43% at a signal-to-noise ratio of 5dB and 98.99% at a signal-to-noise ratio of 30dB, significantly improving the recognition accuracy and system robustness of carbamate pesticides.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121237267A_ABST
    Figure CN121237267A_ABST
Patent Text Reader

Abstract

The invention discloses a pesticide Raman spectrum-based fusion decision-making system design, and particularly relates to the technical field of pesticide Raman spectrums.The pesticide Raman spectrum-based fusion decision-making system design comprises the following steps: S1, collecting Raman spectrum data of four carbamate pesticides; s2, carrying out preprocessing on the original Raman spectrum data by utilizing normalization processing, wavelet transform denoising and a self-adaptive iteration reweighted penalty least square method airPLS; and S3, carrying out dimension reduction on the preprocessed data by using a principal component analysis (PCA). The fusion decision-making system design based on the pesticide Raman spectrum is applied to a Raman detection technology in the field of carbamate pesticide identification, and a Raman spectrum database of carbamate pesticides is constructed through a data enhancement method. The algorithm input redundancy is greatly reduced through PCA dimension reduction, classification research is performed on the algorithm input redundancy based on three types of machine learning of a random forest, a support vector machine and a BP neural network, and finally a final recognition result is obtained through a soft voting decision fitting method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of pesticide Raman spectroscopy, and particularly relates to a fusion decision system design based on pesticide Raman spectroscopy. BACKGROUND

[0002] Pesticide detection is not only a key issue to ensure food safety and human health, but also an important measure to promote agricultural modernization and sustainable development. Compared with traditional pesticide detection methods, Raman spectroscopy does not require complex pretreatment, and its molecular vibration spectral characteristics can accurately identify the structural information of different pesticides. Combined with surface-enhanced Raman scattering (SERS) technology, the detection sensitivity can be greatly enhanced. However, in the Raman spectrum recognition work for similar chemical structures, manual spectrum recognition has the problems of spectrum recognition difficulty and low accuracy.

[0003] In recent years, the method of combining Raman spectroscopy technology with machine learning has brought revolutionary breakthroughs in the field of pesticide detection. This method not only improves the sensitivity and accuracy of pesticide detection, but also effectively improves the detection speed and detection ability in complex environments. Principal component analysis is widely used in dimensionality reduction of Raman spectrum data to improve recognition accuracy. In order to deal with the problems of Raman data scarcity and weak model generality, many researchers have carried out a series of research on fusion decision and made significant progress. SUMMARY

[0004] The main purpose of the present application is to provide a fusion decision system design based on pesticide Raman spectroscopy, which can effectively solve the problems in the background art.

[0005] To achieve the above-mentioned purpose, the technical scheme adopted by the present application is as follows:

[0006] A fusion decision system design based on pesticide Raman spectroscopy, comprising the following steps:

[0007] S1, collecting Raman spectrum data of four kinds of carbamate pesticides;

[0008] S2, using normalization processing, wavelet transform denoising and adaptive iterative reweighted penalized least squares (airPLS) to pretreat the original Raman spectrum data, and constructing a Raman spectrum database of carbamate pesticides through a data enhancement method, the data enhancement method including numerical offset, linear superposition, filtering processing and adding Gaussian white noise to the spectrum data;

[0009] S3, using principal component analysis (PCA) to reduce the dimension of the preprocessed data, and based on three machine learning algorithms of random forest (RF), support vector machine (SVM) and BP neural network, a fusion decision system is constructed, the fusion decision system uses soft voting method to integrate the output probabilities of the three models, and the class with the maximum sum of final probabilities is taken as the recognition result.

[0010] Preferably, in the step S1, the four kinds of carbamate pesticides are: 1-naphthyl-N-methyl carbamate, o-isopropyl phenyl methyl carbamate and 2-sec-butyl phenyl-N-methyl carbamate.

[0011] Preferably, in the step S2, the normalization processing adopts the minimum-maximum standardization method to linearly transform the original data to the interval [0, 1];

[0012] The wavelet transform denoising adopts db4 wavelet basis function for 5-layer decomposition, adopts the default threshold rule and selects the soft threshold processing mode, and the threshold factor is set to 2;

[0013] The airPLS algorithm is used for baseline correction, and the smoothness parameter λ is set to 10 5 , and the iteration stop condition is that the weight change is less than 0.001% and the maximum iteration number is 20 times.

[0014] Preferably, in the data enhancement method, the numerical offset is to shift the entire spectrum by 1 to 20 wave number units to the left and right on the wave number axis;

[0015] Linear superposition is to add two and multiple spectra of the same category of pesticides by weighting with a proportional coefficient k (0.1≤k≤0.9), and the sum of the proportional coefficients is ensured to be 1;

[0016] The filtering processing includes sliding average filtering (window size is 5), exponential weighted moving average filtering (smoothing factor α=0.3) and Savitzky-Golay filtering (window size is 5, polynomial order is 2);

[0017] Adding Gaussian white noise is to calculate and add noise with corresponding power according to formula (5) in the range of 5dB to 30dB according to the target signal-to-noise ratio (SNR).

[0018] Preferably, in the step S3, the principal component analysis (PCA) retains the first 9 principal components, and the cumulative contribution rate is greater than 95%;

[0019] The random forest model contains 100 decision trees, and the node splitting criterion is "gini";

[0020] The support vector machine model uses radial basis function (RBF) as kernel function, with penalty parameter C = 1.0 and kernel function parameter gamma = "scale";

[0021] The BP neural network is a three-layer network structure containing one hidden layer with 12 nodes, the activation function is ReLU, the optimizer is Adam, and the learning rate is set to 0.001.

[0022] The soft voting method involves summing the category probability vectors output by the three machine learning models and taking the category with the largest sum of probabilities as the final prediction result.

[0023] Compared with the prior art, the present invention has the following beneficial effects:

[0024] 1. This invention applies Raman spectroscopy technology to the identification of carbamate pesticides. A Raman spectral database of carbamate pesticides is constructed using data augmentation methods. Dimensionality reduction via PCA significantly reduces the redundancy of the algorithm input. Three machine learning methods—random forest, support vector machine, and backpropagation neural network—are used for classification studies. Finally, a soft-voting decision fitting method is employed to obtain the final identification result. The accuracy reaches 65.43% at a signal-to-noise ratio of 5dB and 98.99% at 30dB. Compared to general machine learning methods, this method not only effectively improves the identification accuracy but also possesses superior system robustness. This method provides a new approach for pesticide Raman spectroscopy detection, contributing to the classification and identification of carbamate pesticides in agricultural production processes. Attached Figure Description

[0025] Figure 1 This is a flowchart of the pesticide Raman classification system of the present invention;

[0026] Figure 2 This is a diagram illustrating the Raman data preprocessing process of the present invention.

[0027] Figure 3 This is a Raman data enhancement diagram of the present invention;

[0028] Figure 4 This is a noisy image of the Raman data in this invention;

[0029] Figure 5 This is a confusion matrix diagram of the three machine learning methods of this invention;

[0030] Figure 6 This is a diagram showing the PCA principal component Pareto plot and classification system of the present invention under different noise conditions.

[0031] Figure 1(a) Flowchart of the pesticide Raman classification system; (b) Chemical structure diagrams of four pesticides; (c) Raman data of the four pesticides at [200, 1800] cm⁻¹.

[0032] Figure 2 (a) Original Raman data; (b) Normalized Raman data; (c) Wavelet-denoised Raman data; (d) Baseline-corrected Raman data;

[0033] Figure 3 (a) Raman image shifted 20 units to the left; (b) Raman data after linear superposition; (c) moving average filtering; (d) weighted filtering; (e) exponential filtering; (f) SG filtering;

[0034] Figure 4 Medium: (a)5dB; (b)10dB; (c)15dB; (d)20dB; (e)25dB; (f)30dB;

[0035] Figure 5 (a) Random Forest training set; (b) Support Vector Machine training set; (c) Backpropagation Neural Network training set; (d) Random Forest test set; (e) Support Vector Machine test set; (f) Backpropagation Neural Network test set;

[0036] Figure 6 (a) Principal component Pareto plot of the FOM double Raman spectral data; (b) showing the classification effect of convergence decision under different signal-to-noise ratios. Detailed Implementation

[0037] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0038] Example 1: Raman spectra of four carbamate pesticides were tested, with a test range of 200-1800 cm⁻¹. -1 The sample to be tested is placed on a test plate. Solid samples are spread flat on a glass slide, while solution samples need to form hemispherical droplets and then dry at room temperature. Each sample is tested ten times and the data is recorded. Based on the assignment of various pesticide Raman vibration modes, the characteristic peaks of different types of Raman are compared as key information for the qualitative analysis of the samples.

[0039] The four pesticides are: disulfide, 1-naphthyl-N-methylcarbamate, o-isopropylphenylmethylcarbamate, and 2-sec-butylphenyl-N-methylcarbamate.

[0040] I. Experimental Samples

[0041] Thiram: Dimethyl disulfide (thiocarbonyl dimethylamine (C6H12N2S4, analytical grade));

[0042] Sevin: 1-Naphthyl-N-methylcarbamate (C12H11NO2, 97%);

[0043] Isopropylcarbamate: o-isopropylphenyl methyl carbamate (C11H15NO2, analytical grade);

[0044] sec-Butylcarbamate: 2-sec-butylphenyl-N-methylcarbamate (C12H17NO2, 97%).

[0045] Example 2: A standardized method is used to perform a linear transformation on the original Raman data, so that the data falls within the [0,1] interval, which is more conducive to the training accuracy of the classification system. The normalization process reference formula is as follows:

[0046]

[0047] Where xN represents the normalized data, xmin is the minimum value of the test data, and xmax is the maximum value of the test data.

[0048] The wavelet transform denoising process generally consists of three steps: multi-scale decomposition using wavelet transform, denoising at each scale, and signal reconstruction. Through wavelet denoising, the characteristic peak signal-to-noise ratio of Raman data can be improved by 4.2 dB.

[0049] Background noise is removed using the airPLS adaptive iterative reweighted penalized least squares method. This algorithm effectively removes noise while preserving the effective information of the Raman spectrum to the maximum extent. The algorithm mainly consists of two steps. First, a curve smoothing method is used to balance the accuracy of the data and the roughness of the fitted data. The calculation formulas for both are shown below:

[0050]

[0051] The original peak value of the spectrum is x, the length of the abscissa of the spectral data is m, z is the fitted spectral peak value, F is the accuracy of z relative to x, and R is the roughness of z. As can be seen from the formula, the larger F is, the better the fit of the curve; the larger R is, the smoother the curve, but at the same time, it deviates more from the true value. These two are contradictory and can be adjusted using the balance parameter Q, as shown in the following formula:

[0052] Q=F+λR

[0053] A larger λ indicates a greater proportion of roughness, resulting in a smoother fitted curve. The adaptive iterative reweighting method is mainly used to calculate data weights and add penalty values ​​to control the smoothness of the fitted baseline. The calculation formula is as follows:

[0054]

[0055] Where t is the iteration number, and the formula for the weighted vector ω is as follows:

[0056]

[0057] dt is composed of the part where the difference between the vector x and zt-1 in the t iterations is less than 0. If the value of the i-th point is greater than the reference value, it is considered to be part of the peak, and its weight is set to 0, and it is ignored in the next iteration calculation.

[0058] Step 3: During the noise addition process, considering the uncertainty of the ratio between the signal power and the added Gaussian white noise power, a method of setting the signal-to-noise ratio (SNR) is introduced for noise addition. The SNR formula is as follows:

[0059]

[0060] Wherein, SNR is the signal-to-noise ratio, measured in dB; Ps is the signal power; and Pn is the noise power. The larger the SNR value, the smaller the relative noise power.

[0061] Moving average filtering obtains the filtering result at the current time by comparing the sample observations at previous and subsequent time points and calculating their average. The calculation formula is shown below:

[0062]

[0063] pi represents the filtered result at time t. gt represents the actual measured value, and n is the sliding window radius (n is set to 10).

[0064] Weighted sliding filtering adjusts the influence of each observation on the filtering result according to different weights. The calculation formula is as follows:

[0065]

[0066] Here, n is set to 7, and the weighting coefficient vector ω is [1,2,3,4,3,2,1].

[0067] The weights of the exponential moving average filter decay over time, and the calculation formula is as follows:

[0068] p t =ω·g t +(1-ω)·p t-1

[0069] ω is the decay weight, which is usually taken as 0.9.

[0070] SavitzkyGolay (SG) filtering is based on the linear least squares method and achieves smooth filtering without changing the signal's direction and width. The specific formula is as follows:

[0071]

[0072] Set the polynomial degree to 4 and the filter window to 9.

[0073] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A design of a fusion decision system based on pesticide Raman spectroscopy, characterized in that, The method comprises the following steps: S1, collecting Raman spectrum data of four kinds of carbamate pesticides; S2, preprocessing the original Raman spectrum data by using normalization processing, wavelet transform denoising and adaptive iterative reweighted penalized least squares (airPLS), and constructing a Raman spectrum database of carbamate pesticides by a data enhancement method, the data enhancement method comprising numerical offset, linear superposition, filter processing and adding Gaussian white noise to the spectrum data; S3, reducing the dimension of the preprocessed data by using principal component analysis (PCA), and constructing a fusion decision system based on three machine learning algorithms, namely random forest (RF), support vector machine (SVM) and BP neural network, the fusion decision system integrating the output probabilities of the three models by using soft voting method, and taking the class with the maximum sum of probabilities as the recognition result.

2. The system according to claim 1, characterized in that it comprises: In the step S1, the four kinds of carbamate pesticides are 1-naphthyl-N-methyl carbamate, o-isopropyl phenyl methyl carbamate and 2-sec-butyl phenyl-N-methyl carbamate.

3. The system as claimed in claim 1, wherein the system is designed based on the fusion decision of the Raman spectrum of the pesticide. In the step S2, the normalization processing adopts the minimum-maximum standardization method to linearly transform the original data to the interval [0, 1]; The wavelet transform denoising adopts db4 wavelet basis function for 5-layer decomposition, adopts the default threshold rule and selects the soft threshold processing mode, and the threshold factor is set to 2; The airPLS algorithm was used for baseline correction with smoothing parameter λ set to 10 5 and the iteration stopping condition was a weight change of less than 0.001% and a maximum of 20 iterations.

4. The system as claimed in claim 1, wherein the system is designed based on the fusion decision of the Raman spectrum of the pesticide. In the data enhancement method, the numerical offset is to shift the entire spectrum by 1 to 20 wave number units to the left and right on the wave number axis; The linear superposition is to add two or more spectra of the same class of pesticides by weighting according to the proportion coefficient k (0.1≤k≤0.9), and the sum of the proportion coefficients is ensured to be 1; The filter processing includes sliding average filtering (window size of 5), exponential weighted moving average filtering (smoothing factor α=0.3) and Savitzky-Golay filtering (window size of 5, polynomial order of 2); The adding of Gaussian white noise is to calculate and add the noise with the corresponding power according to the target signal-to-noise ratio (SNR) in the range of 5dB to 30dB according to formula (5).

5. The system as claimed in claim 1, wherein the said system is designed based on the fusion decision of the Raman spectroscopy of the pesticides. In the step S3, the principal component analysis (PCA) retains the first 9 principal components, and the cumulative contribution rate is greater than 95%; The random forest model contains 100 decision trees, and the node splitting criterion is "gini"; The support vector machine model uses a radial basis function (RBF) as the kernel function, the penalty parameter C=1.0, and the kernel function parameter gamma="scale"; The BP neural network is a three-layer network structure containing one hidden layer, the number of hidden layer nodes is 12, the activation function is ReLU, the optimizer is Adam, and the learning rate is set to 0.001; The soft voting method is to sum the class probability vectors output by the three machine learning models, and take the class with the maximum probability sum as the final prediction result.