A leek pesticide residue detection method combining hyperspectral imaging and deep learning

CN122415628BActive Publication Date: 2026-08-11QINGDAO AGRI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0003]传统的韭菜农药残留检测方法主要依赖于色谱-质谱联用技术,这些方法虽然检测精度高,但存在样品前处理复杂、检测周期长、设备昂贵且需要专业人员操作等缺点,难以满足现场快速检测和大规模筛查的需求

Benefits of technology

一、本发明突破现有高光谱农残检测技术的应用局限,结合韭菜自身基质光谱干扰特性与多农残混合光谱重叠的核心难题,设计独创的解缠结分支网络结构,搭配自主研发的特征正交约束计算公式,实现韭菜基质背景特征与各类农药残留特征的完全拆分剥离,同步引入数据集分布均衡约束算法,规避传统随机划分方式造成的样本分层失衡问题,有效抑制内源物质对痕量农残光谱信号的掩盖作用,从特征提取根源降低多农药之间的特征交叉干扰,稳定提升混合农残同步分类识别与浓度反演的综合精度,检测检出范围全面覆盖国标限定的残留阈值,能够稳定实现多类别复合型农残的一体化精准检测。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122415628B_ABST
    Figure CN122415628B_ABST
Patent Text Reader

Abstract

This invention discloses a method for detecting pesticide residues in chives that combines hyperspectral imaging and deep learning, belonging to the field of non-destructive testing technology for agricultural products. By constructing a detangling neural network model, this invention innovatively designs a spectral feature encoding module, a detangling feature decomposition module, and a multi-task output head module, achieving complete separation and isolation of the chive matrix background features from various pesticide residue features. This effectively solves the core problems of spectral overlap and matrix interference from mixed pesticide residues. Simultaneously, this method combines a dataset distribution balancing constraint algorithm with a multi-task joint loss function, significantly improving the model's detection accuracy and generalization ability. It can achieve integrated and accurate detection of multiple types of compound pesticide residues while maintaining sample integrity and non-destructive testing, providing a new technical means for agricultural product quality and safety testing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of non-destructive testing technology for agricultural products, specifically a method for detecting pesticide residues in leeks that combines hyperspectral imaging and deep learning. Background Technology

[0002] With the rapid development of modern agriculture, pesticides play a vital role in increasing crop yields and controlling pests and diseases. However, the widespread use of pesticides has also led to pesticide residues in agricultural products, seriously threatening consumer health and safety. Leeks, as a common vegetable, are particularly vulnerable to pesticide residues. Due to their short growth cycle and susceptibility to pests and diseases, farmers frequently use pesticides during cultivation, resulting in a high variety and concentration of pesticide residues in leeks.

[0003] Traditional methods for detecting pesticide residues in leeks primarily rely on chromatography-mass spectrometry (GC-MS). While these methods offer high accuracy, they suffer from drawbacks such as complex sample pretreatment, long detection cycles, expensive equipment, and the need for specialized personnel, making them unsuitable for rapid on-site testing and large-scale screening. Furthermore, hyperspectral imaging, as a non-contact, non-destructive, and rapid detection method, has shown great potential in agricultural product quality testing. However, when dealing with complex matrix backgrounds such as those found in leeks and mixed contamination of multiple pesticide residues, traditional hyperspectral analysis methods often suffer from decreased accuracy due to severe overlap of spectral features and significant matrix interference, making it difficult to achieve simultaneous qualitative and quantitative detection of multiple pesticide residues. Specifically, traditional methods lack effective feature extraction and de-entanglement mechanisms when processing hyperspectral data, failing to accurately distinguish between leek matrix background features and pesticide residue characteristics. This results in trace pesticide residue signals being masked by strong background signals, thus affecting detection sensitivity and accuracy.

[0004] In view of the shortcomings of traditional methods for detecting pesticide residues in chives, this invention proposes a method for detecting pesticide residues in chives that combines hyperspectral imaging and deep learning, which is of particular importance. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a method for detecting pesticide residues in chives that combines hyperspectral imaging and deep learning. This method innovatively designs a spectral feature encoding module, a detangling feature decomposition module, and a multi-task output head module by constructing a detangling neural network model. This achieves complete separation and isolation of the chive matrix background features from various pesticide residue features, effectively solving the core problems of spectral overlap and matrix interference caused by mixed pesticide residues. Simultaneously, this method combines a dataset distribution balance constraint algorithm with a multi-task joint loss function, significantly improving the model's detection accuracy and generalization ability. It can achieve integrated and accurate detection of multiple types of compound pesticide residues while maintaining sample integrity, providing a new technical means for agricultural product quality and safety testing.

[0006] To solve the above-mentioned technical problems, this invention provides the following technical solution: a method for detecting pesticide residues in chives combining hyperspectral imaging and deep learning, the specific steps of which are as follows: S1. Leek Sample Preparation and Hyperspectral Imaging Data Acquisition: Prepare leek samples covering single pesticide residues and mixed pesticide residue contamination, along with blank control samples, and acquire hyperspectral image data of all samples using a hyperspectral imaging system; S2. Hyperspectral data preprocessing and true value calibration: The acquired hyperspectral image data is preprocessed to remove noise and background interference, and the average spectral data of the effective area of ​​the leek leaves is extracted; the types and concentrations of pesticide residues in each leek sample are determined by national standard physicochemical testing methods and used as label data for model training. S3. Construction of Hybrid Pesticide Residue Detection Dataset: The preprocessed spectral data is paired with the corresponding pesticide residue type and concentration label data one by one, and the training set, validation set and test set are divided according to the preset ratio to construct a dedicated dataset for training and validation of the untangled neural network. S4. Construction of a Disentangled Neural Network Model for Mixed Pesticide Residues: A disentangled neural network for the simultaneous qualitative and quantitative detection of mixed pesticide residues in leeks is constructed. The network sequentially includes a spectral feature encoding module, a disentangled feature decomposition module, and a multi-task output head module. The disentangled feature decomposition module is used to decompose the depth features of the mixed spectrum into mutually independent leek matrix background feature subspaces and individual pesticide residue-specific feature subspaces, thereby eliminating the cross-interference of spectral features of mixed pesticide residues and the background interference of the leek endogenous matrix. S5. Model Training and Validation: The untangled neural network is trained under supervision using the constructed dataset. The network parameters are iteratively optimized through a multi-task joint loss function. The model hyperparameters are adjusted through the validation set. The generalization and detection accuracy of the model are verified through the test set, resulting in a trained pesticide residue detection model for leeks. S6. Detection of Leek Samples to be Tested: Collect hyperspectral image data of leek samples to be tested, perform the same preprocessing operation as in S2, and input the trained leek pesticide residue detection model, including but not limited to on-site rapid detection and laboratory high-precision detection modes, and simultaneously output qualitative identification results of pesticide residue types and quantitative inversion results of concentration of each pesticide residue in the sample to be tested.

[0007] Furthermore, in S1, the leek samples were prepared from freshly harvested leeks from the same batch. After removing yellow, rotten, or mechanically damaged leaves, the leeks were rinsed three times with ultrapure water and then air-dried naturally in a constant temperature and humidity clean environment for 24 hours to ensure that the leaf surface was free of free moisture and exogenous impurities. The sample preparation covered six pesticides from three major categories that are frequently detected in leek cultivation: organophosphates, pyrethroids, and carbamates. Single pesticide residue pollution groups, binary mixed pollution groups, ternary mixed pollution groups, quaternary mixed pollution groups, and a blank control group were set up. Each group had eight concentration gradients. For each concentration gradient, at least 25 parallel samples were set up, and at least 120 parallel samples were set up for the blank control group. The hyperspectral imaging system adopted a linear array pushbroom imaging spectrometer with a spectral acquisition range of 400–1000 nm visible and near-infrared band and a spectral resolution of 3 nm. The acquisition process was completed in a sealed dark box, which was equipped with two sets of symmetrically distributed halogen linear light sources. The samples were placed on a high-precision electrically controlled translation stage. Before each acquisition, the system dark current calibration and white board calibration were performed to ensure the consistency and stability of the acquired data.

[0008] Furthermore, in S2, the hyperspectral data preprocessing is performed in a fixed sequence. First, black-and-white correction is performed on the acquired raw hyperspectral image to eliminate systematic errors caused by dark current noise in the imaging system and uneven illumination of the light source. Then, Savitzky-Golay smoothing and denoising is performed on the corrected image to eliminate random noise interference while preserving spectral characteristic peaks and inflection point information. Subsequently, multivariate scattering correction is performed on the smoothed spectral data to eliminate spectral baseline drift and scattering errors caused by uneven leaf thickness and different surface roughness of leek leaves. Finally, normalization is used to map the spectral reflectance values ​​to the 0-1 range, reducing the overall reflectance difference band between different samples. To mitigate interference, a fixed threshold segmentation method was employed to extract the region of interest. First, background areas were removed from the image using grayscale threshold segmentation. Then, texture features were used to filter out necrotic leaf areas, insect-hole areas, occluded areas, and edge-distorted areas, retaining only the healthy, effective areas of the leek leaves. The spectral average of all pixels within the effective area was calculated to obtain the one-dimensional average spectral vector corresponding to a single sample. True value calibration was performed using gas chromatography-mass spectrometry, strictly adhering to the GB23200.8-2016 standard for pretreatment and detection. Three parallel measurements were conducted, and the average value was used as the true value for the pesticide residue type and concentration of the sample. Simultaneously, blank controls and matrix matching calibration were performed to ensure the accuracy of the true data.

[0009] Furthermore, in S3, the dataset construction employs a stratified balanced sampling method to partition the dataset. First, all samples are stratified and grouped according to three dimensions: pesticide residue type, mixed pesticide residue quantity, concentration, and gradient. This ensures that samples within each stratum have consistent label features. Then, samples within each stratum are assigned to the training set, validation set, and test set respectively in a fixed ratio of 7:2:1. During the partitioning process, the sample distribution balance constraint formula controls the difference in sample distribution between different datasets. The formula is: ,in The coefficient of variation in the dataset distribution. This represents the total number of hierarchical groups. For the first The proportion of samples within each stratified group in the training set. For the first The proportion of samples within each stratified group in the test set, subject to mandatory constraints during the partitioning process. The value is no greater than 0.02 to ensure that the sample distribution of the training set, validation set and test set is completely consistent, and to avoid the problems of model overfitting and insufficient generalization caused by data distribution deviation. After the division is completed, the spectral data and label data of all datasets are matched and verified one by one. Invalid samples with missing labels and abnormal spectral data are removed. Finally, a special dataset adapted for training untangled neural networks is constructed. All samples in the dataset are equipped with complete spectral data, qualitative labels of pesticide residue types and quantitative labels of pesticide residue concentrations.

[0010] Furthermore, the spectral feature encoding module in S4 adopts a three-layer cascaded one-dimensional convolutional neural network structure. The input is the preprocessed one-dimensional spectral vector. The first convolutional layer sets a corresponding number of convolutional kernels, a stride of 1, and equal-length padding to ensure that the feature vector length remains consistent before and after the convolution operation. After the first convolutional layer, a batch normalization layer, a ReLU activation layer, and a max-pooling layer are added. The batch normalization layer is used to eliminate the training instability caused by feature distribution shift. The ReLU activation layer is used to introduce nonlinear transformation to improve feature fitting ability. The max-pooling layer completes the downsampling operation of the feature dimension. The second convolutional layer adjusts the convolution. The number of kernels is kept the same as the first layer, while the other convolutional parameters remain the same. The parameters of the subsequent batch normalization layers, activation layers, and pooling layers are exactly the same as the first layer. The number of kernels in the third convolutional layer is adjusted again, while the other convolutional parameters remain the same as the first two layers. Subsequent layers only have batch normalization layers and ReLU activation layers, without pooling layers, to avoid excessive loss of core features. The output of the encoding module is a fixed-dimensional mixed spectral depth encoded feature, realizing the step-by-step extraction from shallow spectral waveform features to deep abstract chemical features. At the same time, through equal-length padding and hierarchical convolutional kernel design, the weak spectral features corresponding to trace pesticide residues are preserved, avoiding the loss of effective information during the feature extraction process.

[0011] Furthermore, the untangled feature decomposition module in S4 sets up N+1 parallel feature untangling branches, where N is the preset total number of detectable pesticide residue types. One branch is dedicated to the leek substrate background, and the remaining N branches correspond one-to-one with the N target pesticide residues. Each untangling branch adopts a two-layer cascaded fully connected layer structure, and each branch is equipped with L1 sparse regularization constraints to force each branch to learn only the specific features of the corresponding target, suppressing the activation of non-target features. Orthogonal constraint layers are set between branches, and the feature subspaces output by different branches are forced to be orthogonal through the untangled orthogonal constraint loss function, as shown in the formula: ,in To resolve the loss due to entanglement orthogonal constraints, The feature matrix is ​​obtained by concatenating the feature vectors output from all untangled branches. Characteristic matrix The transpose of the matrix, It is the identity matrix. Using the Frobenius norm, this formula achieves complete decoupling of different feature subspaces by constraining the inner product of feature vectors of different branches to approach 0, completely eliminating feature cross-interference between mixed pesticide residues, and completely separating the background features of the leek matrix from the pesticide residue features, avoiding the strong background signal from masking the weak features of trace pesticide residues. The feature vector output by each branch is the exclusive independent feature subspace of the corresponding target.

[0012] Furthermore, the multi-task output head module in S4 includes two parallel branches: a qualitative classification head and a quantitative regression head. These two branches share the pesticide residue-specific feature subspace output by the detangling feature decomposition module. The qualitative classification head first concatenates the feature vectors output from the N pesticide residue feature subspaces to obtain a fused feature vector of the corresponding dimension. Then, the fused feature vector is input into a two-layer cascaded fully connected layer, followed by a Softmax activation layer. This layer outputs the probability value corresponding to each pesticide residue pollution type. The type corresponding to the highest probability is taken as the qualitative identification result, thus achieving the identification of all pesticides in mixed pesticide residues. For simultaneous identification of pesticide species, the quantitative regression head is configured with N parallel independent regression branches. Each regression branch is directly connected to the dedicated feature subspace of the corresponding pesticide residue. Each regression branch adopts a two-layer cascaded fully connected layer structure, directly outputting the concentration inversion value of the corresponding pesticide residue, realizing independent and accurate quantification of each pesticide residue, and avoiding quantitative interference between multiple pesticide residues. A feature linkage mechanism is set between the qualitative classification head and the quantitative regression head. Only when the qualitative classification head identifies the presence of the corresponding pesticide residue will the concentration value output by the quantitative regression branch of that pesticide residue be included in the final result, avoiding false positive quantitative output when there is no pesticide residue.

[0013] Furthermore, in S5, model training employs a multi-task joint supervised learning approach, iteratively optimizing all learnable parameters of the network through a multi-task joint loss function, as shown in the formula: ,in For the total loss of multiple tasks, This is the multi-class cross-entropy loss, used to optimize the recognition accuracy of the qualitative classification head. The mean squared error loss is used to optimize the inversion accuracy of the quantitative regression head. To resolve the entanglement orthogonal constraint loss, this is used to optimize the decoupling effect of the feature subspace. L1 sparse loss is used to optimize the feature-specific learning capability of each untangled branch. , , , The weighting coefficients for the four loss terms are determined using orthogonal experimentation, with the overall detection accuracy on the test set as the evaluation metric. The final weighting coefficients are determined after multiple sets of gradient experiments. The value is 0.3. The value is 0.4. The value is 0.2. The value is set to 0.1. By balancing the four optimization objectives of qualitative classification, quantitative regression, feature detangling, and sparse constraints through weight allocation, the model is ensured to have high classification accuracy, high quantitative accuracy, and strong feature decoupling ability. The Adam optimizer is used in the training process, and the corresponding initial learning rate, weight decay coefficient, batch size, and maximum number of iterations are set. A dynamic learning rate decay mechanism is set during training, and an early stopping mechanism is set. When the total loss of the validation set does not decrease for 15 consecutive rounds, the training is automatically terminated, and the model weights with the lowest loss on the validation set are saved as the final pesticide residue detection model for leeks.

[0014] Furthermore, model validation in S5 is divided into three stages: accuracy validation, generalization validation, and anti-interference validation. Accuracy validation is completed using test set samples, and the qualitative identification accuracy, precision, recall, and F1 score of the model are statistically analyzed. Simultaneously, the coefficient of determination, root mean square error, mean absolute error, and detection limit for quantitative inversion of each pesticide residue concentration are calculated to ensure that the model's detection accuracy meets the requirements of national food safety standards. Generalization validation is completed using leek samples from different batches, production areas, and varieties. Three different leek varieties from different production areas are selected, and pesticide residue contamination samples consistent with the training set are prepared. These samples are not used for model training but only for generalization testing. The model's performance across different batches and production areas is statistically analyzed. The detection accuracy variation in scene samples ensures that the model maintains stable detection performance in leek samples from different sources. Anti-interference verification is completed by simulating the complex environment of on-site detection. Random noise, baseline drift, and illumination fluctuation interference of different intensities are added to the spectral data. The changes in detection accuracy of the model under different interference intensities are statistically analyzed to ensure that the model has strong anti-interference ability in complex on-site environments. After verification, the model weights are quantized and compressed to reduce the number of model parameters without sacrificing detection accuracy, thereby improving the model's inference speed and ensuring that the detection time for a single sample does not exceed 5 seconds, which is suitable for the needs of rapid batch detection on-site.

[0015] Furthermore, the detection of the leek sample in S6 is divided into two modes: on-site rapid detection and laboratory high-precision detection. The on-site rapid detection mode uses a portable hyperspectral imaging device for data acquisition. Before acquisition, the device is calibrated on-site, including dark current calibration, white plate calibration, and standard substance calibration, to ensure consistency between the acquired data and the training set data. During acquisition, three complete and healthy leaves of the leek to be tested are selected, and hyperspectral images are acquired for each leaf. The images of each leaf undergo the same preprocessing operations as in S2, and the corresponding average spectral vector is extracted. These are then input into the trained detection model, and the average of the three detection results is taken as the final detection result. The laboratory high-precision detection mode uses... The desktop hyperspectral imaging system completes data acquisition with parameters identical to those of S1. Five intact, healthy leaves of the chives to be tested are selected, and hyperspectral image acquisition and preprocessing are performed separately. Pixel-level spectral data within the effective area of ​​each leaf are extracted and input into the trained detection model to obtain the spatial distribution visualization results and global average concentration results of pesticide residues on the leaf surface. Qualitative identification results of pesticide residue types and quantitative results of the concentration of each pesticide residue are output simultaneously. Both detection modes are equipped with an automatic result judgment function, which compares the detected pesticide residue concentration values ​​with the corresponding maximum residue limit values ​​in the GB2763-2021 standard, automatically outputs the judgment result of pass or fail, and generates a complete test report.

[0016] Compared with existing technologies, this method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning has the following advantages: I. This invention breaks through the application limitations of existing hyperspectral pesticide residue detection technologies. Combining the core challenges of the spectral interference characteristics of the leek matrix and the overlapping of multiple pesticide residue spectra, it designs an original untangled branch network structure and, with the help of a self-developed feature orthogonal constraint calculation formula, achieves complete separation and isolation of leek matrix background features and various pesticide residue features. Simultaneously, it introduces a dataset distribution balance constraint algorithm to avoid the sample stratification imbalance problem caused by traditional random partitioning methods, effectively suppressing the masking effect of endogenous substances on trace pesticide residue spectral signals, reducing feature cross-interference between multiple pesticides from the root of feature extraction, and steadily improving the comprehensive accuracy of simultaneous classification and identification and concentration inversion of mixed pesticide residues. The detection range fully covers the residue thresholds limited by national standards, and can stably achieve integrated and accurate detection of multiple types of compound pesticide residues.

[0017] II. This invention constructs a hierarchical convolutional coding structure adapted to leek spectral data, fully preserving the key spectral features of weak trace pesticide residues. It employs an original multi-task joint loss function to complete model iterative optimization, and relies on orthogonal experiments to complete the quantification and calibration of multiple weight parameters. It balances the collaborative optimization requirements of feature untangling, sparse constraints, classification tasks, and regression tasks. The model has undergone standardized verification under multiple scenarios and interference conditions, and has excellent generalization ability across production areas and batches. The entire detection process does not require the intervention of chemical reagents, keeping the sample intact and undamaged. At the same time, it distinguishes between two execution modes: rapid on-site detection and precise laboratory detection, simplifying the detection operation process and shortening the detection time per sample. It can be widely adapted to the batch detection needs of different application scenarios such as leek planting bases, fresh food wholesale markets, and food safety regulatory departments.

[0018] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0020] Figure 1 A flowchart of a method for detecting pesticide residues in chives that combines hyperspectral imaging and deep learning; Figure 2 A flowchart illustrating the construction of a detangled neural network model for a method combining hyperspectral imaging and deep learning to detect pesticide residues in chives; Figure 3 This is a flowchart illustrating the dual-mode detection process for pesticide residues in chives, a method combining hyperspectral imaging and deep learning. Detailed Implementation

[0021] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0022] Example: This example addresses the real-time screening scenario for chives in urban vegetable markets. It utilizes a chive pesticide residue detection method combining hyperspectral imaging and deep learning, perfectly meeting the needs for rapid, accurate, and convenient on-site testing. The specific implementation steps are as follows: S1. Leek Sample Preparation and Hyperspectral Imaging Data Acquisition: Fresh leeks from the same batch were selected to prepare test samples. The samples covered six frequently detected pesticides in leek cultivation, belonging to three major categories: organophosphates, pyrethroids, and carbamates. Single pesticide residue groups, binary mixed residue groups, ternary mixed residue groups, quaternary mixed residue groups, and a blank control group were sequentially set up. Eight concentration gradients were set up in each group to cover the common residue concentration ranges in actual detection. At least 25 parallel samples were set up for each concentration gradient to ensure statistical significance. The blank control group was set up... No fewer than 120 parallel samples were used to exclude matrix and environmental interference. Hyperspectral imaging was performed using a linear array pushbroom imaging spectrometer. The acquisition process was carried out in a sealed dark box to isolate external stray light interference. Two sets of symmetrically distributed halogen linear light sources were configured inside the dark box to ensure uniform illumination of the leek samples. The samples were placed on a high-precision electronically controlled translation stage to ensure accurate acquisition position and stable imaging. Before each acquisition, the system dark current calibration and whiteboard calibration were performed to eliminate equipment noise and system errors, thereby obtaining high-quality and highly stable raw hyperspectral image data.

[0023] S2. Hyperspectral Data Preprocessing and Truth Value Calibration: The original hyperspectral images undergo a complete preprocessing process in a fixed sequence. First, black-and-white correction is performed to eliminate equipment system errors. Then, Savitzky-Golay smoothing and denoising are performed to eliminate random noise interference while preserving spectral characteristic peaks and inflection point information. Subsequently, multivariate scattering correction is performed to offset spectral deviations caused by scattering from the sample surface. Next, normalization is used to map the spectral reflectance values ​​to the 0-1 range to unify the data scale and improve model computational efficiency. Finally, a fixed threshold segmentation method is used to extract the region of interest. Background areas are first removed, followed by areas with leaf necrosis, insect holes, occlusion, and edge distortion. Only healthy and valid areas are retained, and the spectral average of all pixels in that area is calculated to obtain a precise one-dimensional average spectral vector representing the sample characteristics. Truth value calibration is performed using gas chromatography-mass spectrometry (GC-MS) from national standard physicochemical testing methods. Three parallel measurements are taken, and the average value is used as the true value for pesticide residue types and concentrations. Simultaneously, blank controls and matrix matching calibration are performed to ensure the accuracy and authority of the label data, providing a reliable annotation basis for model training.

[0024] S3. Construction of Mixed Pesticide Residue Detection Dataset: Preprocessed spectral data are precisely paired with corresponding pesticide residue type and concentration label data. A stratified balanced sampling method is used to partition the dataset. First, stratification is performed based on three dimensions: pesticide residue pollution type, quantity of mixed pesticide residues, and concentration gradient. Then, samples within each stratified group are assigned to the training, validation, and test sets in a fixed ratio of 7:2:1. During the partitioning process, the difference in sample distribution between different datasets is strictly controlled using a dataset distribution difference coefficient constraint formula. The formula is: ,in The coefficient of variation in the dataset distribution. This represents the total number of hierarchical groups. For the first The proportion of samples within each stratified group in the training set. For the first The proportion of samples within each stratified group in the test set is used to avoid overfitting or insufficient generalization caused by sample distribution bias, thus constructing a dedicated dataset suitable for training and validating untangled neural networks.

[0025] S4. Construction of a Disentangled Neural Network Model for Mixed Pesticide Residues: A disentangled neural network was built for the simultaneous qualitative and quantitative detection of mixed pesticide residues in leeks. The network sequentially includes a spectral feature encoding module, a disentangled feature decomposition module, and a multi-task output head module. The spectral feature encoding module adopts a three-layer cascaded one-dimensional convolutional neural network structure, sequentially paired with a batch normalization layer, a ReLU activation layer, and a max pooling layer. The batch normalization layer is used to eliminate the training instability caused by feature distribution shift, the ReLU activation layer is used to introduce nonlinear transformation to improve feature fitting ability, and the max pooling layer is used to complete feature dimension downsampling. The extraction process proceeds step-by-step from shallow spectral waveform features to deep abstract chemical features, while preserving the weak spectral features corresponding to trace pesticide residues. The untangled feature decomposition module employs seven parallel feature untangling branches: one branch is dedicated to the leek matrix background, and six branches correspond one-to-one with the six target pesticide residues. Each branch uses a two-layer cascaded fully connected layer structure with L1 sparse regularization constraints, forcing branches to learn only the target-specific features and suppressing the activation of non-target features. Orthogonal constraint layers are set between branches, employing untangled orthogonal constraint loss to ensure that the feature subspaces output by different branches are independent. The formula is as follows: ,in To resolve the loss due to entanglement orthogonal constraints, The feature matrix is ​​obtained by concatenating the feature vectors output from all untangled branches. Characteristic matrix The transpose of the matrix, It is the identity matrix. Using the Frobenius norm, it effectively eliminates the cross-interference of mixed pesticide residue spectral features and the background interference of the endogenous matrix in chives. The multi-task output head module has two parallel branches: a qualitative classification head and a quantitative regression head. The two branches share the pesticide residue-specific feature subspace output by the detangling feature decomposition module. After the qualitative classification head splices the pesticide residue features, it outputs the pesticide residue type probability through a 2-layer cascaded fully connected layer and a Softmax activation layer, realizing the simultaneous identification of mixed pesticide residue types. The quantitative regression head has 6 independent regression branches and uses a 2-layer cascaded fully connected layer to output the concentration inversion value. A feature linkage mechanism is set between the two branches, and the quantitative result of the pesticide residue is only adopted when the corresponding pesticide residue is qualitatively identified, avoiding invalid quantitative output and improving detection accuracy.

[0026] S5. Model Training and Validation: The model is trained using a multi-task joint supervised learning approach. All learnable parameters of the network are iteratively optimized using a multi-task joint loss function, as shown in the formula: ,in For the total loss of multiple tasks, For multi-class cross-entropy loss, For mean square error loss, To resolve the loss due to entanglement orthogonal constraints, For L1 sparse loss, , , , The weight coefficients for the four loss terms are defined as follows: This loss function integrates multi-class cross-entropy loss, mean square error loss, untangled orthogonal constraint loss, and L1 sparsity loss, and matches the corresponding weight coefficients to synergistically optimize the qualitative identification and quantitative inversion effects. The model hyperparameters are adjusted through the validation set to optimize model performance, and the model's generalization and detection accuracy are verified through the test set. The qualitative identification accuracy, precision, recall, and F1 score are statistically analyzed, and the determination coefficient, root mean square error, mean absolute error, and detection limit for each pesticide residue concentration are calculated to comprehensively evaluate the model's basic detection capabilities. The generalization verification is completed using samples of leek varieties from three different production areas to test the model's cross-scenario detection stability. The anti-interference verification is completed by simulating complex on-site interferences such as random noise, baseline drift, and light fluctuations to test the model's stability in practical applications. After verification, the model weights are quantized and compressed to reduce the number of model parameters without sacrificing detection accuracy, adapting to the computing power conditions of portable detection devices.

[0027] S6. Rapid On-Site Detection of Leek Samples: On-site rapid detection is employed for sample testing. Portable hyperspectral imaging equipment is used for data acquisition. Before acquisition, the equipment undergoes on-site calibration, including dark current calibration, white board calibration, and standard substance calibration, to eliminate environmental and equipment errors. During acquisition, three intact, healthy leaves of the leeks are selected for hyperspectral image acquisition. The corresponding average spectral vector is extracted and input into the trained detection model. The average of the three test results is taken as the final result, reducing single-test error and improving result reliability. The detection system includes an automatic result judgment function, comparing the detected pesticide residue concentration value with the corresponding maximum residue limit value in GB2763-2021 standard, automatically outputting a pass / fail judgment result, and generating a complete on-site test report. This meets the practical application needs of rapid screening of pesticide residues in leeks in vegetable markets and immediate results.

[0028] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A method for detecting pesticide residues in leeks by combining hyperspectral imaging with deep learning, characterized in that, The specific steps of this method are as follows: S1. Leek Sample Preparation and Hyperspectral Imaging Data Acquisition: Prepare leek samples covering single pesticide residues and mixed pesticide residue contamination, along with blank control samples, and acquire hyperspectral image data of all samples using a hyperspectral imaging system; S2. Hyperspectral data preprocessing and true value calibration: The acquired hyperspectral image data is preprocessed to remove noise and background interference, and the average spectral data of the effective area of ​​the leek leaves is extracted; the types and concentrations of pesticide residues in each leek sample are determined by national standard physicochemical testing methods and used as label data for model training. S3. Construction of Hybrid Pesticide Residue Detection Dataset: The preprocessed spectral data is paired with the corresponding pesticide residue type and concentration label data one by one, and the training set, validation set and test set are divided according to the preset ratio to construct a dedicated dataset for training and validation of the untangled neural network. S4. Construction of a Disentangled Neural Network Model for Mixed Pesticide Residues: A disentangled neural network for the simultaneous qualitative and quantitative detection of mixed pesticide residues in leeks is constructed. The network sequentially includes a spectral feature encoding module, a disentangled feature decomposition module, and a multi-task output head module. The disentangled feature decomposition module is used to decompose the depth features of the mixed spectrum into mutually independent leek matrix background feature subspaces and individual pesticide residue-specific feature subspaces, thereby eliminating the cross-interference of spectral features of mixed pesticide residues and the background interference of the leek endogenous matrix. S5. Model Training and Validation: The untangled neural network is trained under supervision using the constructed dataset. The network parameters are iteratively optimized through a multi-task joint loss function. The model hyperparameters are adjusted through the validation set. The generalization and detection accuracy of the model are verified through the test set, resulting in a trained pesticide residue detection model for leeks. S6. Detection of Leek Samples to be Tested: Collect hyperspectral image data of leek samples to be tested, perform the same preprocessing operation as in S2, and input the trained leek pesticide residue detection model, including but not limited to on-site rapid detection and laboratory high-precision detection modes, and simultaneously output qualitative identification results of pesticide residue types and quantitative inversion results of concentration of each pesticide residue in the sample to be tested.

2. The method according to claim 1, wherein, In S1, the leek sample preparation selected fresh leeks harvested from the same batch. The sample preparation covered 6 pesticides from 3 major categories that were frequently detected in leek cultivation, namely organophosphates, pyrethroids, and carbamates. Single pesticide residue pollution group, binary mixed pollution group, ternary mixed pollution group, quaternary mixed pollution group, and blank control group were set up. Each group had 8 concentration gradients, and each concentration gradient had no less than 25 parallel samples. The blank control group had no less than 120 parallel samples. The hyperspectral imaging system adopted a linear array pushbroom imaging spectrometer. The acquisition process was completed in a sealed dark box. The dark box was equipped with two sets of symmetrically distributed halogen linear light sources. The samples were placed on a high-precision electrically controlled translation stage. Before each acquisition, the system dark current calibration and white board calibration were completed.

3. The method according to claim 1, wherein, In S2, the hyperspectral data preprocessing is performed in a fixed sequence. First, black and white correction is performed on the acquired raw hyperspectral image. Then, Savitzky-Golay smoothing and denoising is performed on the corrected image to eliminate random noise interference while preserving spectral characteristic peaks and inflection point information. Subsequently, multivariate scattering correction is performed on the smoothed spectral data. Then, normalization is used to map the spectral reflectance values ​​to the 0 to 1 range. Finally, a fixed threshold segmentation method is used to extract the region of interest. First, the background area in the image is removed by grayscale threshold segmentation. Then, the leaf necrosis area, insect hole area, occlusion area and edge distortion area are removed by texture feature screening. Only the healthy and effective area of ​​the leek leaf is retained. The spectral average value of all pixels in the effective area is calculated to obtain the one-dimensional average spectral vector corresponding to a single sample. The true value is determined by gas chromatography-mass spectrometry. Three parallel measurements are taken and the average value is taken as the true value of pesticide residue type and concentration for the sample. Blank control and matrix matching calibration are performed simultaneously in the detection process.

4. The method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The dataset construction in S3 employs a stratified balanced sampling method to partition the dataset. First, all samples are stratified and grouped according to three dimensions: pesticide residue type, mixed pesticide residue quantity, concentration gradient, and so on. Then, samples within each stratified group are assigned to the training set, validation set, and test set respectively, using a fixed ratio of 7:2:

1. During the partitioning process, the sample distribution balance constraint formula controls the difference in sample distribution between different datasets. The formula is: ,in Here, represents the dataset distribution dissimilarity coefficient, and represents the total number of stratified groups. For the first The proportion of samples within each stratified group in the training set. For the first The percentage of samples within each stratified group in the test set.

5. The method for detecting pesticide residues in leeks by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The spectral feature encoding module in S4 adopts a three-layer cascaded one-dimensional convolutional neural network structure. The input is the preprocessed one-dimensional spectral vector. The first convolutional layer sets the corresponding number of convolutional kernels, the stride is set to 1, and the padding method is set to equal-length padding. After the first convolutional layer, a batch normalization layer, a ReLU activation layer, and a max pooling layer are provided. The batch normalization layer is used to eliminate the training instability caused by feature distribution shift. The ReLU activation layer is used to introduce nonlinear transformation to improve feature fitting ability. The max pooling layer performs downsampling of the feature dimension. The second convolutional layer adjusts the convolutional vector. The number of kernels is kept the same as the first layer, and the parameters of the subsequent batch normalization layer, activation layer, and pooling layer are exactly the same as the first layer. The number of kernels in the third convolutional layer is adjusted again, and the parameters of the remaining convolutions are kept the same as the first two layers. Subsequent layers only have batch normalization layers and ReLU activation layers, and no pooling layers are set. The output of the encoding module is a fixed-dimensional mixed spectral depth encoding feature, which realizes the stepwise extraction from shallow spectral waveform features to deep abstract chemical features. At the same time, through equal-length padding and hierarchical convolutional kernel number design, the weak spectral features corresponding to trace pesticide residues are preserved.

6. The method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The S4 untangling feature decomposition module sets up N+1 parallel feature untangling branches, where N is the preset total number of detectable pesticide residue types. One branch is dedicated to the leek substrate background, and the remaining N branches correspond one-to-one with the N target pesticide residues. Each untangling branch adopts a two-layer cascaded fully connected layer structure, and each branch is equipped with L1 sparse regularization constraints to force each branch to learn only the specific features of the corresponding target and suppress the activation of non-target features. Orthogonal constraint layers are set between branches, and the feature subspaces output by different branches are forced to be orthogonal through the untangling orthogonal constraint loss function, as shown in the formula: ,in To resolve the loss due to entanglement orthogonal constraints, The feature matrix is ​​obtained by concatenating the feature vectors output from all untangled branches. Characteristic matrix The transpose of the matrix, It is the identity matrix. It is the Frobenius norm.

7. The method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The multi-task output head module in S4 includes two parallel branches: a qualitative classification head and a quantitative regression head. These two branches share the pesticide residue-specific feature subspace output by the detangling feature decomposition module. The qualitative classification head first concatenates the feature vectors output from the N pesticide residue feature subspaces to obtain a fused feature vector of the corresponding dimension. Then, the fused feature vector is input into a two-layer cascaded fully connected layer, followed by a Softmax activation layer. This outputs the probability value corresponding to each pesticide residue pollution type. The type corresponding to the highest probability is taken as the qualitative identification result, achieving simultaneous identification of all pesticide types in mixed pesticide residues. The quantitative regression head sets up N parallel independent regression branches. Each regression branch is directly connected to the corresponding pesticide residue's specific feature subspace. Each regression branch uses a two-layer cascaded fully connected layer structure, directly outputting the concentration inversion value of the corresponding pesticide residue. A feature linkage mechanism is set between the qualitative classification head and the quantitative regression head. Only when the qualitative classification head identifies the existence of a corresponding pesticide residue will the concentration value output by the corresponding quantitative regression branch be included in the final result.

8. The method for detecting pesticide residues in leeks by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, In S5, the model training adopts a multi-task joint supervised learning approach, iteratively optimizing all learnable parameters of the network through a multi-task joint loss function, as shown in the formula: ,in For the total loss of multiple tasks, For multi-class cross-entropy loss, For mean square error loss, To resolve the loss due to entanglement orthogonal constraints, For L1 sparse loss, , , , These are the weighting coefficients corresponding to the four loss terms.

9. The method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The model validation in S5 is divided into three stages: accuracy validation, generalization validation, and anti-interference validation. Accuracy validation is completed using test set samples, and the qualitative recognition accuracy, precision, recall, and F1 score of the model are statistically analyzed. At the same time, the coefficient of determination, root mean square error, mean absolute error, and detection limit for quantitative inversion of each pesticide residue concentration are calculated. Generalization validation is completed using leek samples from different batches, production areas, and varieties. Three different leek varieties from different production areas are selected, and pesticide residue contamination samples consistent with the training set are prepared. These samples are not used for model training but only for generalization testing. The changes in the detection accuracy of the model in cross-scenario samples are statistically analyzed. Anti-interference validation is completed by simulating the complex environment of on-site detection. Random noise, baseline drift, and illumination fluctuation interference of different intensities are added to the spectral data, and the changes in the detection accuracy of the model under different interference intensities are statistically analyzed. After validation, the model weights are quantized and compressed to compress the number of model parameters without sacrificing detection accuracy.

10. The method for detecting pesticide residues in chives by combining hyperspectral imaging and deep learning according to claim 1, characterized in that, The S6 method for detecting the leek samples to be tested is divided into two modes: on-site rapid detection and laboratory high-precision detection. The on-site rapid detection mode uses a portable hyperspectral imaging device for data acquisition. Before acquisition, the device undergoes on-site calibration, including dark current calibration, white plate calibration, and standard substance calibration. During acquisition, three intact, healthy leaves of the leek to be tested are selected, and hyperspectral images are acquired for each leaf. The corresponding average spectral vectors are extracted and input into the trained detection model. The average of the three detection results is taken as the final detection result. The laboratory high-precision detection mode uses a desktop hyperspectral imaging system for data acquisition. Five intact, healthy leaves of the chives to be tested were used to acquire and preprocess hyperspectral images. Pixel-level spectral data of the effective area of ​​each leaf were extracted and input into the trained detection model to obtain the spatial distribution visualization results and the global average concentration results of pesticide residues on the leaf surface. The qualitative identification results of pesticide residue types and the quantitative results of the concentration of each pesticide residue were output simultaneously. Both detection modes are equipped with an automatic result judgment function, which compares the detected pesticide residue concentration values ​​with the corresponding maximum residue limit values ​​in GB2763-2021 standard, automatically outputs the judgment result of qualified or unqualified, and generates a complete test report.

Citation Information

Patent Citations

  • Interferometric phase unwrapping method based on double-branch decoding residual network

    CN119667681A

  • Coarse cereal aflatoxin detection method based on combination of hyperspectral imaging and deep learning

    CN121805167A