DebNet intelligent classification algorithm based on fluorescence method pesticide residue detection

Through DebNet intelligent classification algorithm, combined with fluorescence spectroscopy technology and deep learning technology, the problems of low classification accuracy and high cost in traditional methods in pesticide residue detection are solved, and efficient, accurate and low-cost pesticide residue detection are achieved.

CN120107686AActive Publication Date: 2025-06-06SHANDONG INST OF BUSINESS & TECH

Patent Information

Application Number
CN202510236068.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-06
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

Traditional fluorescence spectroscopy analysis methods have problems such as low classification accuracy, insufficient sample data, overfitting and high data processing complexity in pesticide residue detection, and are costly and are not suitable for large-scale rapid detection.

Method used

The DebNet intelligent classification algorithm based on fluorescence method is used to detect pesticide samples through a fluorescence spectrometer, pre-process and data enhancement of spectral data, and build a hybrid neural network model DebNet, including a convolutional layer, a pooling layer, an LSTM layer, a self-attention module and a fully connected layer, for model training and classification recognition.

Benefits of technology

It improves the accuracy of pesticide classification, reduces sample demand, simplifies data processing process, reduces detection costs, and is suitable for large-scale rapid testing to meet the needs of modern rapid monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107686A_ABST
    Figure CN120107686A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of pesticide residue detection, and particularly relates to a DebNet intelligent classification algorithm based on fluorescence method pesticide residue detection, which comprises the following steps: detecting a pesticide sample by using a fluorescence spectrometer to obtain one-dimensional spectral data of the pesticide sample; preprocessing the spectral data; performing data enhancement on the preprocessed spectral data, wherein an enhancement mode comprises linear interpolation, noise injection, spectrum translation and spectrum cutting; building a hybrid neural network model comprising an input layer, a convolutional layer, a pooling layer, an LSTM layer, a self-attention module, a full connection layer and an output layer, namely a DebNet model; training a DebNet model based on the enhanced spectral data, wherein the training process comprises forward propagation and back propagation; and utilizing the trained DebNet model to classify and identify the spectral data of a to-be-detected sample. The method can improve the classification accuracy, reduce the sample requirements, simplify the data processing flow, and reduce the detection cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of pesticide residue detection, and specifically relates to a DebNet intelligent classification algorithm based on fluorescence pesticide residue detection. Background Art

[0002] The widespread use of pesticides is of great significance to improving agricultural production efficiency, but it also poses a threat to the environment and human health. Therefore, the detection and classification of pesticide residues are crucial to ensure food safety and environmental protection. Traditional pesticide detection methods, such as gas chromatography and high-performance liquid chromatography, are accurate, but cumbersome, time-consuming and sample-destructive. In recent years, fluorescence spectroscopy technology has been widely used in the field of pesticide residue detection as a rapid, non-destructive and low-cost detection method.

[0003] However, traditional fluorescence spectroscopy analysis methods still face problems such as low classification accuracy, insufficient sample data, and complex hyperspectral data processing, which are mainly reflected in: Insufficient classification accuracy: The feature overlap and similarity of fluorescence spectral data lead to frequent misjudgments in the classification process of traditional methods, making it difficult to effectively distinguish different types of pesticides; Overfitting problem: Existing machine learning methods, especially when the sample size is small, are prone to overfitting, resulting in poor model generalization ability; High data processing complexity: Traditional methods have high requirements for the preprocessing of spectral data, and most methods require complex manual intervention, which limits their automated application; High cost: Traditional pesticide residue detection methods usually require expensive equipment and long detection cycles, which are not suitable for large-scale rapid detection. Summary of the invention

[0004] In view of the above deficiencies in the prior art, the purpose of the present invention is to provide a DebNet intelligent classification algorithm based on fluorescence pesticide residue detection, which can improve classification accuracy, reduce sample requirements, simplify data processing procedures, and reduce detection costs.

[0005] To achieve the above objectives, the present invention provides a DebNet intelligent classification algorithm based on fluorescence pesticide residue detection, comprising the following steps: S1. Using a fluorescence spectrometer to detect pesticide samples and obtain one-dimensional spectrum data of the pesticide samples; S2, preprocessing the spectral data, the preprocessing method includes filtering and denoising, correcting the baseline drift of the spectrum and normalization; S3, performing data enhancement on the preprocessed spectral data, wherein the enhancement methods include linear interpolation, noise injection, spectral shift and spectrum clipping; S4. Build a hybrid neural network model including input layer, convolution layer, pooling layer, LSTM layer, self-attention module, fully connected layer and output layer, which is the DebNet model; S5. Train the DebNet model based on the enhanced spectral data. The training process includes forward propagation and back propagation. S6. Use the trained DebNet model to classify and identify the spectral data of the sample to be tested.

[0006] As a preferred embodiment of the present invention, in the S1, the measurement range of the fluorescence spectrometer used is 300-500nm, the sampling interval is 0.5nm, the excitation wavelength is set to 280nm, and the one-dimensional spectral data is in the form of an array of wavelength-fluorescence intensity.

[0007] As a preferred embodiment of the present invention, in S2, the preprocessing process is to first calculate the signal-to-noise ratio SNR of the spectral data and shield the wavelength range where SNR is less than 3, then use a Savitz-Ky-Golay filter to denoise the spectral data to remove noise and smooth the spectrum, and then apply an adaptive iterative reweighted penalized least squares algorithm to correct the baseline drift of the spectrum, and finally use Min-max normalization to standardize all data to the range of [0, 1] to ensure that the fluorescence intensity between different samples will not interfere with the DebNet model.

[0008] As a preferred solution of the present invention, in the preprocessing process, an adaptive spectral segment weighted fusion method is set after the normalization step to further perform preprocessing, specifically including: Step 1: Divide the normalized spectral data into several sub-bands of fixed length, each sub-band is 10 nm long; Step 2: Calculate the local signal-to-noise ratio (LSNR) for each sub-band and dynamically assign weights based on the LSNR. The weight calculation formula is: ; In the formula, is the weight of the z-th sub-band; is a learnable parameter; U is the total number of sub-bands, and u is the index of the sub-band; , are the LSNR of the zth and uth sub-bands respectively; Step 3: Perform feature fusion on the weighted sub-bands through a one-dimensional convolution layer, with a convolution kernel size of 1×1 and an output channel number of 32; Step 4: Perform residual connection on the fused features and the original spectral data to generate the final optimized preprocessed data.

[0009] As a preferred solution of the present invention, in S3, the data enhancement method is specifically: Linear interpolation: Linear interpolation is performed on samples with the same label. Two spectral data sets that share the same label are randomly selected for each operation. Interpolation is implemented by generating two random numbers ranging from 0 to 1 and ensuring that their sum is equal to 1. These two random numbers are then multiplied by the corresponding spectral data and added to obtain new interpolated samples. Noise injection: Add a Gaussian signal to the original spectrum to simulate random noise; Spectral shift: The shift direction is randomly selected, including positive or negative, and the shift amount is uniformly sampled within the range of ±0.2nm. The shifted spectrum is resampled using the cubic spline interpolation algorithm. Spectrum clipping: Randomly select wavelength positions and set the corresponding spectral signal intensity value to zero at the selected wavelength position to simulate spectral missing or abnormal conditions in the measurement.

[0010] As a preferred solution of the present invention, in the S4, the specific architecture of the DebNet model includes 4 convolutional layers, 4 pooling layers, 1 LSTM layer, 1 self-attention module, 2 fully connected layers, and the final output layer uses the Softmax function for multi-classification; Among them, the LeakyReLU activation function is used after the convolution layer to prevent the gradient vanishing problem; the maximum pooling layer is set after the first three convolution layers, and the global average pooling layer is set after the last convolution layer; the LSTM layer contains 100 hidden units to capture the time series characteristics of the spectral data and generate a global time-dependent representation; the self-attention module is used to optimize feature selection and enhance the ability to focus on key features; the two fully connected layers have 128 and 64 neurons respectively to further extract features, and regularization techniques are used to prevent overfitting; A batch normalization layer is set after each convolutional layer and fully connected layer to improve training speed and stability; the self-attention module includes a fully connected layer 1, a ReLU activation layer, a fully connected layer 2, and a Softmax layer, which are set in sequence.

[0011] As a preferred embodiment of the present invention, the process of obtaining the output of the DebNet model is as follows: After the enhanced spectral data is input into the convolution layer, the LeakyReLU activation function is used, which is expressed as: ; In the formula, represents the LeakyReLU activation function; x represents the input signal; a is a coefficient between 0 and 1, which is used to control the output slope when x is negative; The calculation formula of the one-dimensional structure convolution kernel used in the convolution layer is: ; In the formula, Represents the value of the feature map output after convolution at position i; is the value of the input signal at position i+m; is the weight parameter of the convolution kernel at position m; b is the bias term used to adjust the convolution output; M is the size of the convolution kernel; The maximum pooling layer is set after the first three convolutional layers, and the maximum pooling method is used for sampling, which is expressed as: ; In the formula, It represents the value of the jth feature map at position l after the hth convolution layer and the maximum pooling layer; l is the size of the convolution kernel; max represents the maximum pooling operation, which selects the maximum value from the given input; and Represents the two adjacent values ​​at position j in the feature map output by the hth convolutional layer; A global average pooling layer is set after the last convolutional layer, and its calculation formula is: ; Where y represents the output value of the global average pooling layer; H and W represent the height and width of the feature map respectively; Represents the value of the position (p, q) in the feature map, p represents the index on the height dimension of the feature map, and q represents the index on the width dimension of the feature map; After multiple layers of convolution and maximum pooling operations, the extracted sample features are processed by the global average pooling layer to reduce the dimension of the feature map and convert it into a vector representation of a fixed size. The feature vector is then flattened into a one-dimensional vector and input into the LSTM layer. The features output from the LSTM layer are further input into the self-attention module. The fully connected layer 1 of the self-attention module maps the features output by the LSTM to a low-dimensional attention space. The number of neurons is set to 100. Then, nonlinear mapping is introduced through the ReLU activation function to improve the flexibility of feature selection. Feature weights are generated from the low-dimensional attention space through the fully connected layer 2. The number of neurons is also 100. Finally, the generated weights are normalized by the Softmax function through the Softmax layer so that the sum of the weights is 1. The importance of different features is weighted. The generated attention weights are multiplied element by element with the features output by the LSTM to highlight key features and suppress irrelevant information. The features optimized by the self-attention module are input into two fully connected layers with 128 and 64 neurons respectively, and the Dropout ratios are 0.5 and 0.3 respectively to prevent overfitting; The Softmax activation function is used in the final output layer to realize the probability prediction of pesticide categories, and the number of neurons in the output layer is 4.

[0012] As a preferred solution of the present invention, in the DebNet model, a multi-scale time feature extraction unit is further provided between the LSTM layer and the self-attention module, and the unit is composed of a parallel time convolution network TCN branch and a bidirectional gated recurrent unit BiGRU branch, wherein: The TCN branch adopts a dilated causal convolution structure, which includes three convolutional layers with dilation rates of 1, 2, and 4, and the convolution kernel size of each layer is 3, which is used to capture the multi-scale local temporal dependencies of the spectral sequence; The BiGRU branch contains 50 forward GRU units and 50 reverse GRU units, which are used to extract bidirectional long-range temporal features; The output features of TCN and BiGRU are dynamically integrated through the gated fusion mechanism. The fusion formula is: ; In the formula, Represents the fused feature output; Represents the Sigmoid function; is a learnable parameter matrix used to adjust the weights of input features; Represents the output features of TCN; Represents the output features of BiGRU; Indicates concatenating the output features of TCN and BiGRU; Represents element-wise multiplication.

[0013] As a preferred embodiment of the present invention, in S5, the training process is: in the forward propagation, convolution, pooling, LSTM, self-attention module and fully connected layer are calculated in sequence, and finally a predicted value is generated, and the loss function is calculated using the true value; in the back propagation, the weights of each layer of the DebNet model are updated by calculating the gradient; the training continues until the loss function value converges to the minimum, and the final training result is output; the dynamic hard example mining loss function is used in the back propagation, which is expressed as: ; In the formula, is the dynamic hard example mining loss value; N is the total number of samples in the batch, n represents one of the samples; K is the total number of pesticide categories, k represents one of the categories; is the probability that the nth sample is predicted to be the kth class; Indicates the historical batch The exponential moving average of is the focusing factor; is the true label predicted by the nth sample; ln represents the natural logarithm.

[0014] As a preferred solution of the present invention, in S5, during training, the loss function is reduced by using the Adam optimization algorithm, and the cosine annealing scheduler is integrated to dynamically adjust the learning rate, wherein the Adam optimization algorithm parameters are set to: the exponential decay rate of the first-order moment estimate The exponential decay rate of the second-order moment estimate is 0.9 is 0.999, the numerical stability parameter For 10 -7 ; Learning rate is 0.001; Cosine annealing learning rate The calculation formula is: ; In the formula, the number of cycles t is set to 100; the minimum learning rate is set to ; The maximum learning rate is set to ; To speed up the convergence of the model, the training samples were divided into multiple batches, and the number of batch samples was set to 128. The training samples, i.e. the enhanced spectral data, were randomly divided into three parts: 70% of the spectral data was used as the training set; 10% of the spectral data was used as the validation set to adjust the neuron weight parameters during the back-propagation training process; and 20% of the spectral data was used as the test set to test the performance of the trained network model. After the training is completed, four evaluation indicators, namely precision, recall, F1 score and accuracy, are used to measure the classification performance of the DebNet model.

[0015] The algorithm involved in the present invention can be executed by an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the above algorithm calculation is implemented by executing the software by the processor.

[0016] The beneficial effects of the present invention are: High efficiency: Compared with traditional pesticide residue detection methods (such as gas chromatography, liquid chromatography, etc.), this invention combines fluorescence spectroscopy and deep learning technology to quickly complete the classification and identification of pesticides, greatly improving the detection efficiency. By introducing a multi-layer convolutional network (CNN) and a time series analysis module (LSTM), the feature extraction and processing capabilities are further optimized, which is suitable for rapid detection of large-scale samples, significantly improving timeliness and meeting the needs of modern rapid monitoring.

[0017] Nondestructive testing: Fluorescence spectroscopy is a nondestructive testing method that does not damage pesticide samples and is suitable for multiple analyses and sample preservation. This feature makes this method very suitable for applications that require sample integrity. At the same time, the acquisition of spectral data does not rely on complex chemical reagents or expensive equipment, effectively reducing the operating costs of the experiment.

[0018] High classification accuracy: By adopting a hybrid model of one-dimensional convolutional neural network (1D-CNN) and LSTM, the present invention can extract multidimensional features from fluorescence spectral data, and combine the self-attention module to perform weighted optimization on key features, significantly improving the accuracy of pesticide classification. Especially in the case of overlapping spectral data and high classification difficulty, the hybrid model can effectively distinguish different types of pesticides, solving the problem of insufficient classification accuracy in traditional methods.

[0019] Strong adaptability: The present invention uses data enhancement techniques (including linear interpolation, noise injection, and spectrum clipping) to expand the sample size and enhance the robustness of the model. These techniques can not only improve the generalization ability of the model, but also combine with the self-attention module to enable it to adapt to different types and newly emerging pesticide samples, with excellent flexibility and scalability.

[0020] Low cost: Compared with traditional chemical analysis methods, the equipment required by this invention is simple and mainly relies on fluorescence spectrometers and computer equipment, thus greatly reducing the initial investment and operating costs of detection. In addition, the automated data processing and model training process reduces manual intervention and operation time, further reducing costs.

[0021] Broad application prospects: The present invention is not only applicable to the detection of various pesticide residues, but also has strong scalability and can be applied to the classification and identification of other chemical substances, especially those compounds with similar spectral characteristics. By introducing time series feature extraction (LSTM) and attention mechanism in deep learning, the present invention has broad application potential in complex sample classification and large-scale detection scenarios, and can be widely used in environmental monitoring, food safety and other fields. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the process of the present invention; Figure 2 It is a schematic diagram of the process of reverse propagation in the present invention; Figure 3 It is a structural diagram of the DebNet model in the present invention; Figure 4 It is a schematic diagram of the confusion matrix during verification of the present invention; Figure 5 It is a schematic diagram of the accuracy curve during verification of the present invention; Figure 6 It is a schematic diagram of the loss curve during verification of the present invention. DETAILED DESCRIPTION

[0023] The embodiments of the present invention are further described below in conjunction with the accompanying drawings: Example 1 like Figure 1 As shown, a DebNet intelligent classification algorithm based on fluorescence pesticide residue detection includes the following steps: S1. Using a fluorescence spectrometer to detect pesticide samples and obtain one-dimensional spectrum data of the pesticide samples; S2, preprocessing the spectral data, the preprocessing method includes filtering and denoising, correcting the baseline drift of the spectrum and normalization; S3, performing data enhancement on the preprocessed spectral data, wherein the enhancement methods include linear interpolation, noise injection, spectral shift and spectrum clipping; S4. Build a hybrid neural network model including input layer, convolution layer, pooling layer, LSTM layer, self-attention module, fully connected layer and output layer, which is the DebNet model; S5. Train the DebNet model based on the enhanced spectral data. The training process includes forward propagation and back propagation. S6. Use the trained DebNet model to classify and identify the spectral data of the sample to be tested.

[0024] In S1, the measurement range of the fluorescence spectrometer used is 300-500 nm, the sampling interval is 0.5 nm, the excitation wavelength is set to 280 nm, and the one-dimensional spectral data is in the array form of wavelength-fluorescence intensity.

[0025] In S2, the preprocessing process is to first calculate the signal-to-noise ratio (SNR) of the spectral data and screen out the low-quality wavelength range (such as 300-320 nm) with SNR < 3. Then, the spectral data is denoised using a Savitz-Ky-Golay filter (polynomial degree = 3, filter length = 7) to remove noise and smooth the spectrum. Then, an adaptive iterative reweighted penalized least squares algorithm is applied to correct the baseline drift of the spectrum. Finally, Min-max normalization is used to standardize all data to the range of [0, 1] to ensure that the fluorescence intensity between different samples does not interfere with the DebNet model.

[0026] The SNR is calculated using a dynamic wavelength selection algorithm, a standard deviation method, a peak height method, a wavelet transform, a Fourier transform or a filtering method, preferably a dynamic wavelength selection algorithm.

[0027] Sufficient training data can enable the neural network to fully learn the internal features of the data, enhance the generalization ability and robustness of the model, and avoid overfitting as much as possible. It is difficult to obtain a large number of fluorescence spectrum samples by manpower alone, so this embodiment uses a data enhancement method to expand the number of samples.

[0028] In S3, the data enhancement method is as follows: Linear interpolation: Linear interpolation is performed on samples with the same label. Two spectral data sets that share the same label are randomly selected for each operation. Interpolation is implemented by generating two random numbers ranging from 0 to 1 and ensuring that their sum is equal to 1. These two random numbers are then multiplied by the corresponding spectral data and added to obtain new interpolated samples. Noise injection: A Gaussian signal is added to the original spectrum to simulate random noise, with a mean of 0 and a standard deviation of 0.01; Spectral shift: To address the wavelength calibration deviation problem of the fluorescence spectrometer, the original spectral data is subjected to wavelength shift operation; the shift direction is randomly selected, including positive or negative, and the shift amount is uniformly sampled within the range of ±0.2nm (such as +0.15nm or -0.18nm), and the shifted spectrum is resampled using the cubic spline interpolation algorithm; Spectrum clipping: Randomly select wavelength positions and set the corresponding spectral signal intensity value to zero at the selected wavelength position to simulate spectral missing or abnormal conditions in the measurement.

[0029] Data augmentation improves the model's adaptability to different detection conditions and incomplete data.

[0030] In S4, the specific architecture of the DebNet model includes 4 convolutional layers, 4 pooling layers, 1 LSTM layer, 1 self-attention module, and 2 fully connected layers. The final output layer uses the Softmax function for multi-classification. The convolution kernel sizes used in the 4 convolutional layers are 7x1, 3x1, 3x1, and 3x1, respectively, and the number of convolution kernels in each layer is 32, 64, 128, and 256, respectively. Among them, the LeakyReLU activation function is used after the convolution layer to prevent the gradient disappearance problem; the maximum pooling layer is set after the first three convolution layers, and the global average pooling layer is set after the last convolution layer; the LSTM layer contains 100 hidden units to capture the time series characteristics of spectral data and generate a global time-dependent representation; the self-attention module is used to optimize feature selection and enhance the ability to focus on key features; the two fully connected layers have 128 and 64 neurons respectively to further extract features, and regularization techniques are used to prevent overfitting.

[0031] The structural diagram of the DebNet model is as follows Figure 3 As shown, the output part corresponds to Table 1.

[0032] In the DebNet model, a batch normalization layer is set after each convolutional layer and fully connected layer to improve training speed and stability; the self-attention module includes a fully connected layer 1, a ReLU activation layer, a fully connected layer 2, and a Softmax layer, which are set in sequence.

[0033] The process of obtaining output of DebNet model is as follows: After the enhanced spectral data is input into the convolution layer, the LeakyReLU activation function is used, which is expressed as: ; In the formula, represents the LeakyReLU activation function; x represents the input signal; a is a coefficient between 0 and 1, which is used to control the output slope when x is negative; The calculation formula of the one-dimensional structure convolution kernel used in the convolution layer is: ; In the formula, Represents the value of the feature map output after convolution at position i; is the value of the input signal at position i+m; is the weight parameter of the convolution kernel at position m; b is the bias term used to adjust the convolution output; M is the size of the convolution kernel; The maximum pooling layer is set after the first three convolutional layers, and the maximum pooling method is used for sampling, which is expressed as: ; In the formula, It represents the value of the jth feature map at position l after the hth convolution layer and the maximum pooling layer; l is the size of the convolution kernel; max represents the maximum pooling operation, which selects the maximum value from the given input; and Represents the two adjacent values ​​at position j in the feature map output by the hth convolutional layer; A global average pooling layer is set after the last convolutional layer, and its calculation formula is: ; Where y represents the output value of the global average pooling layer; H and W represent the height and width of the feature map respectively; Represents the value of the position (p, q) in the feature map, p represents the index on the height dimension of the feature map, and q represents the index on the width dimension of the feature map; After multiple layers of convolution and maximum pooling operations, the extracted sample features are processed by the global average pooling layer to reduce the dimension of the feature map and convert it into a vector representation of a fixed size. The feature vector is then flattened into a one-dimensional vector and input into the LSTM layer. The features output from the LSTM layer are further input into the self-attention module. The fully connected layer 1 of the self-attention module maps the features output by the LSTM to a low-dimensional attention space. The number of neurons is set to 100. Then, nonlinear mapping is introduced through the ReLU activation function to improve the flexibility of feature selection. Feature weights are generated from the low-dimensional attention space through the fully connected layer 2. The number of neurons is also 100. Finally, the generated weights are normalized by the Softmax function through the Softmax layer so that the sum of the weights is 1. The importance of different features is weighted. The generated attention weights are multiplied element by element with the features output by the LSTM to highlight key features and suppress irrelevant information. The features optimized by the self-attention module are input into two fully connected layers with 128 and 64 neurons respectively, and the Dropout ratios are 0.5 and 0.3 respectively to prevent overfitting; The Softmax activation function is used in the final output layer to realize the probability prediction of pesticide categories, and the number of neurons in the output layer is 4.

[0034] In S5, the training process is to calculate the convolution, pooling, LSTM, self-attention module and fully connected layer in sequence in the forward propagation, finally generate the predicted value, and use the true value to calculate the loss function; Figure 2 As shown in the figure, in the back propagation, the weights of each layer of the DebNet model are updated by calculating the gradient; the training continues until the loss function value converges to the minimum, and the final training result is output; the dynamic hard example mining loss function is used in the back propagation, which is expressed as: ; In the formula, is the dynamic hard example mining loss value; N is the total number of samples in the batch, n represents one of the samples; K is the total number of pesticide categories, k represents one of the categories; is the probability that the nth sample is predicted to be the kth class; Indicates the historical batch The exponential moving average of is the focusing factor; is the true label predicted by the nth sample (in one-hot encoding form, with a value of 0 or 1); ln represents the natural logarithm; During training, the loss function is reduced using the Adam optimization algorithm, and the cosine annealing scheduler is integrated to dynamically adjust the learning rate. The parameters of the Adam optimization algorithm are set to: the exponential decay rate of the first-order moment estimate The exponential decay rate of the second-order moment estimate is 0.9 is 0.999, the numerical stability parameter For 10 -7 ; Learning rate is 0.001; Cosine annealing learning rate The calculation formula is: ; In the formula, the number of cycles t is set to 100; the minimum learning rate is set to ; The maximum learning rate is set to ; To speed up the convergence of the model, the training samples were divided into multiple batches, and the number of batch samples was set to 128. The training samples, i.e. the enhanced spectral data, were randomly divided into three parts: 70% of the spectral data was used as the training set; 10% of the spectral data was used as the validation set to adjust the neuron weight parameters during the back-propagation training process; and 20% of the spectral data was used as the test set to test the performance of the trained network model. After the training is completed, four evaluation indicators, namely precision, recall, F1 score and accuracy, are used to measure the classification performance of the DebNet model.

[0035] After prediction on the test set, the model performance is evaluated using four evaluation indicators: precision, recall, F1 score and accuracy, which can fully reflect the effectiveness of the model in identifying different types of pesticides and ensure that the model has high classification accuracy and strong generalization ability in practical applications. For example, in a specific detection task, according to the intelligent classification algorithm of this embodiment, the final performance data is as follows: Table 1 Performance data table

[0036] At the same time, the loss value curve, accuracy curve and confusion matrix are plotted to evaluate the performance of the model in the pesticide classification task, such as Figure 4-Figure 6 As shown, it can be seen that this model has a good classification effect.

[0037] Example 2 This embodiment further includes the following improvements based on Embodiment 1: In the preprocessing process, an adaptive spectral segment weighted fusion method is set after the normalization step for further preprocessing, specifically including: Step 1: Divide the normalized spectral data into several sub-bands of fixed length, each sub-band is 10 nm long; Step 2: Calculate the local signal-to-noise ratio (LSNR) for each sub-band and dynamically assign weights based on the LSNR. The weight calculation formula is: ; In the formula, is the weight of the z-th sub-band; is a learnable parameter; U is the total number of sub-bands, and u is the index of the sub-band; , are the LSNR of the zth and uth sub-bands respectively; Step 3: Perform feature fusion on the weighted sub-bands through a one-dimensional convolution layer, with a convolution kernel size of 1×1 and an output channel number of 32; Step 4: Perform residual connection on the fused features and the original spectral data to generate the final optimized preprocessed data.

[0038] The sub-band weights are automatically optimized through end-to-end training to suppress interference in low signal-to-noise ratio areas while enhancing feature expression in high signal-to-noise ratio areas.

[0039] In the DebNet model, a multi-scale temporal feature extraction unit is set between the LSTM layer and the self-attention module. The unit consists of a parallel temporal convolutional network TCN branch and a bidirectional gated recurrent unit BiGRU branch, where: The TCN branch adopts a dilated causal convolution structure, which includes three convolutional layers with dilation rates of 1, 2, and 4, and the convolution kernel size of each layer is 3, which is used to capture the multi-scale local temporal dependencies of the spectral sequence; The BiGRU branch contains 50 forward GRU units and 50 reverse GRU units, which are used to extract bidirectional long-range temporal features; The output features of TCN and BiGRU are dynamically integrated through the gated fusion mechanism. The fusion formula is: ; In the formula, Represents the fused feature output; Represents the Sigmoid function; is a learnable parameter matrix used to adjust the weights of input features; Represents the output features of TCN; Represents the output features of BiGRU; Indicates concatenating the output features of TCN and BiGRU; Represents element-wise multiplication.

[0040] The fused features are input into the self-attention module, which improves the model's ability to identify complex spectral patterns by jointly optimizing multi-scale temporal features and attention mechanism.

Claims

1. A DebNet intelligent classification algorithm based on fluorescence pesticide residue detection, characterized in that The following steps are involved: S1. Using a fluorescence spectrometer to detect pesticide samples and obtain one-dimensional spectrum data of the pesticide samples; S2, preprocessing the spectral data, the preprocessing method includes filtering and denoising, correcting the baseline drift of the spectrum and normalization; S3, performing data enhancement on the preprocessed spectral data, wherein the enhancement methods include linear interpolation, noise injection, spectral shift and spectrum clipping; S4. Build a hybrid neural network model including input layer, convolution layer, pooling layer, LSTM layer, self-attention module, fully connected layer and output layer, which is the DebNet model; S5. Train the DebNet model based on the enhanced spectral data. The training process includes forward propagation and back propagation. S6. Use the trained DebNet model to classify and identify the spectral data of the sample to be tested.

2. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In the above-mentioned S1, the measurement range of the fluorescence spectrometer used is 300-500nm, the sampling interval is 0.5nm, the excitation wavelength is set to 280nm, and the one-dimensional spectrum data is in the form of an array of wavelength-fluorescence intensity.

3. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In the S2 described above, the preprocessing process is to first calculate the signal-to-noise ratio (SNR) of the spectral data and shield the wavelength range where SNR < 3, then use the Savitz-Ky-Golay filter to denoise the spectral data to remove noise and smooth the spectrum, then apply the adaptive iterative reweighted penalized least squares algorithm to correct the baseline drift of the spectrum, and finally use Min-max normalization to standardize all data to the range of [0, 1] to ensure that the fluorescence intensity between different samples will not interfere with the DebNet model.

4. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 3 is characterized in that: In the preprocessing process, an adaptive spectral segment weighted fusion method is set after the normalization step for further preprocessing, specifically including: Step 1: Divide the normalized spectral data into several sub-bands of fixed length, each sub-band is 10 nm long; Step 2: Calculate the local signal-to-noise ratio (LSNR) for each sub-band and dynamically assign weights based on the LSNR. The weight calculation formula is: ; In the formula, is the weight of the z-th sub-band; is a learnable parameter; U is the total number of sub-bands, and u is the index of the sub-band; , are the LSNR of the zth and uth sub-bands respectively; Step 3: Perform feature fusion on the weighted sub-bands through a one-dimensional convolution layer, with a convolution kernel size of 1×1 and an output channel number of 32; Step 4: Perform residual connection on the fused features and the original spectral data to generate the final optimized preprocessed data.

5. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In the above-mentioned S3, the data enhancement method is specifically as follows: Linear interpolation: Linear interpolation is performed on samples with the same label. Two spectral data sets that share the same label are randomly selected for each operation. Interpolation is implemented by generating two random numbers ranging from 0 to 1 and ensuring that their sum is equal to 1. These two random numbers are then multiplied by the corresponding spectral data and added to obtain new interpolated samples. Noise injection: Add a Gaussian signal to the original spectrum to simulate random noise; Spectral shift: The shift direction is randomly selected, including positive or negative, and the shift amount is uniformly sampled within the range of ±0.2nm. The shifted spectrum is resampled using the cubic spline interpolation algorithm. Spectrum clipping: Randomly select wavelength positions and set the corresponding spectral signal intensity value to zero at the selected wavelength position to simulate spectral missing or abnormal conditions in the measurement.

6. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In the above S4, the specific architecture of the DebNet model includes 4 convolutional layers, 4 pooling layers, 1 LSTM layer, 1 self-attention module, 2 fully connected layers, and the final output layer uses the Softmax function for multi-classification; Among them, the LeakyReLU activation function is used after the convolution layer to prevent the gradient vanishing problem; the maximum pooling layer is set after the first three convolution layers, and the global average pooling layer is set after the last convolution layer; the LSTM layer contains 100 hidden units to capture the time series characteristics of the spectral data and generate a global time-dependent representation; the self-attention module is used to optimize feature selection and enhance the ability to focus on key features; the two fully connected layers have 128 and 64 neurons respectively to further extract features, and regularization techniques are used to prevent overfitting; A batch normalization layer is set after each convolutional layer and fully connected layer to improve training speed and stability; the self-attention module includes a fully connected layer 1, a ReLU activation layer, a fully connected layer 2, and a Softmax layer, which are set in sequence.

7. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 6 is characterized in that: The process of obtaining the output of the DebNet model is as follows: After the enhanced spectral data is input into the convolution layer, the LeakyReLU activation function is used, which is expressed as: ; In the formula, represents the LeakyReLU activation function; x represents the input signal; a is a coefficient between 0 and 1 that controls the output slope when x is negative; The calculation formula of the one-dimensional structure convolution kernel used in the convolution layer is: ; In the formula, Represents the value of the feature map output after convolution at position i; is the value of the input signal at position i+m; is the weight parameter of the convolution kernel at position m; b is the bias term used to adjust the convolution output; M is the size of the convolution kernel; The maximum pooling layer is set after the first three convolutional layers, and the maximum pooling method is used for sampling, which is expressed as: ; In the formula, It represents the value of the jth feature map at position l after the hth convolution layer and the maximum pooling layer; l is the size of the convolution kernel; max represents the maximum pooling operation, which selects the maximum value from the given input; and Represents the two adjacent values ​​at position j in the feature map output by the hth convolutional layer; A global average pooling layer is set after the last convolutional layer, and its calculation formula is: ; Where y represents the output value of the global average pooling layer; H and W represent the height and width of the feature map respectively; Represents the value of the position (p, q) in the feature map, p represents the index on the height dimension of the feature map, and q represents the index on the width dimension of the feature map; After multiple layers of convolution and maximum pooling operations, the extracted sample features are processed by the global average pooling layer to reduce the dimension of the feature map and convert it into a vector representation of a fixed size. The feature vector is then flattened into a one-dimensional vector and input into the LSTM layer. The features output from the LSTM layer are further input into the self-attention module. The fully connected layer 1 of the self-attention module maps the features output by the LSTM to a low-dimensional attention space. The number of neurons is set to 100. Then, nonlinear mapping is introduced through the ReLU activation function to improve the flexibility of feature selection. Feature weights are generated from the low-dimensional attention space through the fully connected layer 2. The number of neurons is also 100. Finally, the generated weights are normalized by the Softmax function through the Softmax layer so that the sum of the weights is 1. The importance of different features is weighted. The generated attention weights are multiplied element by element with the features output by the LSTM to highlight key features and suppress irrelevant information. The features optimized by the self-attention module are input into two fully connected layers with 128 and 64 neurons respectively, and the Dropout ratios are 0.5 and 0.3 respectively to prevent overfitting; The Softmax activation function is used in the final output layer to realize the probability prediction of pesticide categories, and the number of neurons in the output layer is 4.

8. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 7 is characterized in that: In the DebNet model, a multi-scale temporal feature extraction unit is provided between the LSTM layer and the self-attention module. The unit is composed of a parallel temporal convolutional network TCN branch and a bidirectional gated recurrent unit BiGRU branch, wherein: The TCN branch adopts a dilated causal convolution structure, which includes three convolutional layers with dilation rates of 1, 2, and 4, and the convolution kernel size of each layer is 3, which is used to capture the multi-scale local temporal dependencies of the spectral sequence; The BiGRU branch contains 50 forward GRU units and 50 reverse GRU units, which are used to extract bidirectional long-range temporal features; The output features of TCN and BiGRU are dynamically integrated through the gated fusion mechanism. The fusion formula is: ; In the formula, Represents the feature output after fusion; Represents the Sigmoid function; is a learnable parameter matrix used to adjust the weights of input features; Represents the output features of TCN; Represents the output features of BiGRU; Indicates concatenating the output features of TCN and BiGRU; Represents element-wise multiplication.

9. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In the above S5, the training process is: in the forward propagation, the convolution, pooling, LSTM, self-attention module and fully connected layer are calculated in sequence, and finally the predicted value is generated, and the loss function is calculated using the true value; in the back propagation, the weights of each layer of the DebNet model are updated by calculating the gradient; The training continues until the loss function value converges to the minimum, and the final training result is output; the dynamic hard example mining loss function is used in back propagation, which is expressed as: ; In the formula, is the dynamic hard example mining loss value; N is the total number of samples in the batch, and n represents one of the samples; K is the total number of pesticide categories, and k represents one of the categories; is the probability that the nth sample is predicted to be the kth class; Indicates the historical batch The exponential moving average of is the focusing factor; is the true label predicted by the nth sample; ln represents the natural logarithm.

10. The DebNet intelligent classification algorithm based on fluorescence pesticide residue detection according to claim 1 is characterized in that: In S5, during training, the loss function is reduced using the Adam optimization algorithm, and the cosine annealing scheduler is integrated to dynamically adjust the learning rate, where the Adam optimization algorithm parameters are set to: the exponential decay rate of the first-order moment estimate The exponential decay rate of the second-order moment estimate is 0.9 is 0.999, the numerical stability parameter For 10 -7 ; Learning rate is 0.001; Cosine annealing learning rate The calculation formula is: ; In the formula, the number of cycles t is set to 100; the minimum learning rate is set to ; The maximum learning rate is set to ; To speed up the convergence of the model, the training samples were divided into multiple batches, and the number of batch samples was set to 128. The training samples, i.e., the enhanced spectral data, were randomly divided into three parts: 70% of the spectral data was used as the training set; 10% of the spectral data is used as a validation set to adjust the neuron weight parameters during the back-propagation training process; 20% of the spectral data is used as a test set to test the performance of the trained network model; After the training is completed, four evaluation indicators, namely precision, recall, F1 score and accuracy, are used to measure the classification performance of the DebNet model.

Citation Information

Patent Citations

  • Vegetable oil pesticide residue detection method based on three-dimensional fluorescence spectroscopic technology

    CN110702656A

  • Neural network training method based on improved loss function

    CN112603324A

  • Pesticide residue type identification method based on fluorescence spectra

    CN113916860A

  • Crop pesticide residue detection method and system based on convolutional neural network

    CN116884512A

  • Neural network-based temperature prediction method in optical network-on-chip

    CN117094218A

Cited By

  • GRU regression model and method for realizing intelligent detection of deltamethrin pesticide residues by using same

    CN120744869A

  • Intelligent monitoring method for strength of mine cemented filling body

    CN121636984A

  • Hyperspectral imaging and deep learning combined leek pesticide residue detection method

    CN122415628A