Pesticide residue rapid screening method based on optimized electronic nose array

By optimizing the sensor arrangement and data processing methods of the electronic nose array, and combining the fusion model of support vector machine and random forest, the problems of low efficiency and poor accuracy of existing pesticide residue detection technologies have been solved, and rapid and accurate pesticide residue detection has been achieved.

CN121834713AActive Publication Date: 2026-04-10CHENGDU VOCATIONAL COLLEGE OF AGRI SCI & TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-10
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing pesticide residue detection technologies suffer from problems such as long detection cycles, high equipment costs, large sample losses, low detection accuracy, and poor data processing and model generalization capabilities, making it difficult to meet the needs of rapid, accurate, and convenient detection in agricultural production bases, wholesale markets, and other scenarios.

Method used

An optimized electronic nose array was adopted, and gas sensors with different sensitive materials were selected. The sensor array arrangement was optimized using an impedance matching algorithm. Combined with graded crushing, headspace extraction and dual-channel data acquisition, adaptive polynomial fitting and improved wavelet transform were used for data preprocessing. Finally, a fusion model of support vector machine and random forest was used to determine pesticide residues.

Benefits of technology

It enables real-time screening of pesticide residues, improves detection efficiency and accuracy, reduces false negative and false positive rates, is adaptable to agricultural products with different moisture contents, and ensures the reliability of test results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834713A_ABST
    Figure CN121834713A_ABST
Patent Text Reader

Abstract

The invention relates to a rapid pesticide residue screening method based on an optimized electronic nose array. The rapid pesticide residue screening method comprises the following steps: selecting a plurality of gas sensors made of different sensitive materials to form a basic sensor group, and performing array arrangement optimization; grading the to-be-detected agricultural product sample; synchronously collecting resistance change signals and response peak value occurrence time of the sensor in a specified time period; performing baseline correction on the dynamic response data, and then removing environmental interference noise; screening out specified key feature parameters; dividing the screened key feature parameters into a training set, a verification set and a final feature set according to a specified proportion, and performing training optimization verification on a fusion model of the support vector machine and the random forest; inputting the final feature set into the optimized screening model, and outputting a positive / negative judgment result and confidence; and respective processing is carried out according to the division of the confidence. Through sensor array optimization, pretreatment, multi-dimensional data preprocessing and fusion model construction, rapid and accurate screening of pesticide residues is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of pesticide residue detection, and particularly relates to a pesticide residue rapid screening method based on an optimized electronic nose array. BACKGROUND

[0002] Pesticides, as an important means to ensure the yield of agricultural products, are widely used in agricultural production. However, the problem of pesticide residues caused by excessive use has become a key hidden danger threatening food safety and human health. At present, pesticide residue detection technologies for agricultural products mainly include laboratory detection methods and on-site rapid detection methods. Both methods have obvious technical shortcomings and cannot meet the "fast, accurate and convenient" detection needs of agricultural production bases, wholesale markets and other scenes. The specific problems are as follows: The existing laboratory detection technology is represented by gas chromatography-mass spectrometry (GC-MS) and high performance liquid chromatography (HPLC). Although it has high detection accuracy (the minimum detection limit can reach 0.001 mg / kg), it has the following defects: Long detection period: The complex pretreatment process of sample crushing, solvent extraction, purification and concentration is required, and the single detection time usually takes 4-8 hours, which cannot realize real-time screening of agricultural products in the circulation link; High equipment cost: The unit price of GC-MS and HPLC equipment is as high as hundreds of thousands of yuan, and professional operators are needed for maintenance and operation, which is difficult to popularize in basic detection sites; Large sample loss: A large amount of organic solvents (such as acetonitrile and acetone) are consumed in the detection process, which not only pollutes the environment but also requires a sufficient amount of samples (usually ≥50g), which is not suitable for small batch sample detection. In order to solve the shortcomings of laboratory detection methods, on-site rapid detection technologies (such as colloidal gold immunochromatography, enzyme inhibition method and traditional electronic nose method) are gradually applied, but there are still core technical defects: Low detection accuracy: The colloidal gold immunochromatography and enzyme inhibition method have high cross-reactivity for organophosphorus and pyrethroid pesticides, and the detection accuracy of low concentration residues (<0.01 mg / kg) is often less than 80%, which is easy to cause "false negative" misjudgment; Electronic nose array performance is insufficient: The traditional electronic nose mainly uses a single type of sensitive material (such as metal oxide semiconductor), and the response threshold difference between sensors is irregular (the difference between some adjacent sensors is less than 5%), and the overall response deviation coefficient of the array is often greater than 8%, which causes the characteristic signal to overlap and cannot effectively distinguish different types and different concentrations of pesticide residues; Poor data processing and model generalization: Existing electronic nose detection technologies mostly use a single machine learning model, and do not perform targeted preprocessing on the response signal, such as baseline drift correction and environmental noise removal. The model is not suitable for complex matrix agricultural products such as high-moisture leafy vegetables and high-volatile fruits, and the detection accuracy fluctuates by more than 15%. Lack of intelligence in detection process: The existing rapid detection method does not consider the water content difference of agricultural products during sample pretreatment, and does not have a confidence determination and secondary detection mechanism, making it difficult to guarantee the reliability of the detection results. SUMMARY

[0003] The purpose of the present application is to provide a pesticide residue rapid screening method based on an optimized electronic nose array to solve the technical problems existing in the prior art.

[0004] To solve the above technical problems, the technical solution adopted by the present application is as follows: A pesticide residue rapid screening method based on an optimized electronic nose array, comprising the following steps: S1: Selecting a plurality of different sensitive material gas sensors to form a basic sensor group, and optimizing the array arrangement of the sensors through an impedance matching algorithm; S2: Grading the sample of the agricultural product to be detected, selecting a crushing method according to the water content of the sample, sieving after crushing, and placing a specified amount of the sieved sample in a sealed extraction bottle and adding an extraction liquid in a specified proportion for extraction; S3: Collecting dynamic response data: passing the gas phase in the upper layer of the extraction bottle into the optimized electronic nose array at a specified rate, using a double-channel data collection mode, and synchronously collecting the resistance change signal and response peak appearance time of the sensor within a specified time period; S4: Preprocessing the dynamic response data: first, baseline correction, and then removing environmental interference noise; combining time series characteristics and signal amplitude characteristics for double extraction, and screening out specified key feature parameters; S5: Dividing the screened key feature parameters into a training set, a validation set, and a final feature set in a specified proportion, inputting the training set and the validation set into a support vector machine and a random forest fusion model for model training and optimization verification; S6: Inputting the final feature set into the optimized screening model, outputting positive / negative determination results and confidence, when the confidence is greater than or equal to a first threshold value, directly outputting the determination results, when the confidence is between a second threshold value and the first threshold value, automatically triggering a secondary detection program to repeat steps S3 and S4, taking the average value of the two feature parameters to re-input the model, and when the confidence is less than the second threshold value, prompting that the sample pretreatment is abnormal and re-taking the sample for screening.

[0005] Preferably, the sensitive material in step S1 includes a rare earth element doped metal oxide semiconductor material, a functional conductive polymer material, and a modified graphene material, wherein the rare earth element includes lanthanum La and cerium Ce, the metal oxide semiconductor material includes SnO2, ZnO, and WO3, the functional conductive polymer material is a polypyrrole-carbon nanotube composite material, and the modified graphene material is a graphene-titanium dioxide-silver nanoparticle composite material.

[0006] Preferably, in step S1, the sensors are arrayed and optimized by an impedance matching algorithm, so that the response threshold difference of adjacent sensors is controlled within a specified range, and the overall response deviation coefficient of the array is ≤ a specified value. The specific process is as follows: S11: Establish a sensor impedance-response threshold database; S12: Iterative arrangement based on the impedance matching algorithm.

[0007] Preferably, the specific process of step S11 is as follows: S111: Sensor impedance and response threshold test: impedance test of multiple selected sensitive material sensors in a standard pesticide gas environment, record the impedance values of each sensor at different frequencies; record the response threshold of each sensor simultaneously, and obtain a single sensor response threshold database; S112: Impedance matching target parameter setting, including adjacent sensor response threshold difference, array overall response deviation coefficient, and impedance matching degree.

[0008] Preferably, the specific process of step S12 is as follows: S121: Sensor clustering and grouping based on impedance values: using the K-means clustering algorithm, taking the impedance value of a single sensor at 1 kHz as the clustering feature, dividing 10-12 sensors into 3-4 clustering groups, and ensuring that the similarity of the impedance values of the sensors in the same clustering group is ≥ 90%; S122: In-group sorting and preliminary arrangement based on response threshold: for the sensors in each clustering group, sort them from small to large according to the response threshold, calculate the response threshold difference ΔT of adjacent sensors, if ΔT is not within the range of 8%-12%, then adjust by replacing the sensor: select a sensor with impedance value similarity ≥ 90% and response threshold difference filling value from other clustering groups to replace the sensor that does not meet the requirements in the current group; S123: Iterative adjustment and overall deviation coefficient verification: the sorted sensors in each clustering group are arranged into a preliminary array according to the "inter-group alternation" principle, and the overall response deviation coefficient Cᵥ of the array is calculated. If Cᵥ>3%, then through local fine-tuning optimization: for the sensor with the largest deviation of response threshold from the average value, replace it with the sensor with the response threshold closer to the average value within the same cluster group and the adjacent difference still satisfying 8%-12%, repeat the calculation of Cᵥ until Cᵥ≤3%.

[0009] Preferably, the specific process of baseline correction in step S4 using the adaptive polynomial fitting algorithm is as follows: S41: The dynamic response data obtained in step S3 is denoted as R(t) and stored in the form of a two-dimensional array of time t-resistance change rate R; By calculating the first and second derivatives of the original curve, the degree of nonlinearity of the drift is determined: a first derivative change rate <5% is weak nonlinearity, and a low-order fitting is performed; a change rate ≥5% is strong nonlinearity, and a high-order fitting is performed; S42: Adaptive polynomial order determination and fitting calculation: using the piecewise error evaluation-order iterative optimization method, the optimal order among 3-5 order polynomials is selected: S43: Baseline correction: original data correction and result optimization: Preferably, the specific process of step S42 is as follows: S421: The original response curve is divided into three sections according to the detection stage, including the baseline stable section, the response rising section, and the response stable section, and baseline candidate points are extracted from the three sections: 31 points are taken from the baseline stable section, 10 smooth points are taken from each of the response rising section and the response stable section, and a total of 51 baseline candidate points are obtained; S422: 3rd, 4th, and 5th order polynomials are used to fit the 51 baseline candidate points, and the root mean square error RMSE and the maximum absolute error MAE of each order fitting are calculated; S423: Adaptive order selection: set the error threshold: RMSE≤0.1% and MAE≤0.05% as the optimal fitting standard, and select the order according to the following rules: If the 3rd order fitting meets the error threshold, and the second derivative of the fitting curve has a change rate <3%, then select the 3rd order; If the 3rd order fitting RMSE>0.1% or MAE>0.05%, but the 4th order fitting meets the error threshold, and the third derivative has a change rate <2%, then select the 4th order; If neither the 3rd order fitting nor the 4th order fitting meets the error threshold, or the original curve has strong nonlinearity drift, then select the 5th order.

[0010] Preferably, the specific process of step S43 is as follows: S431: Baseline deduction of original data: The resistance change rate R(t) of each time point in the original response curve is subtracted by the fitting baseline value B(t) at the corresponding time to obtain the corrected response value Rcorr (t), the formula is: R corr (t) = R(t) - B(t); S432: correction result verification: the mean of the corrected data of the baseline stable segment ≤0.1%, the standard deviation ≤0.05%, and the response peak deviation ≤2%; if the corrected data does not meet the requirements, automatically re-execute the order determination until it meets the standard.

[0011] Preferably, the specific process of screening the key characteristic parameters in step S4 is as follows: S44: based on the response signal denoised by improved wavelet transform, 12 initial candidate characteristic parameters are extracted; 3 types of samples are selected, each type of sample has 30 repetitions, and a total of 90 samples; 12 candidate characteristics are taken as columns, and 90 samples are taken as rows to construct a feature matrix X; the sample categories are taken as labels to construct a category vector Y; S45: principal component analysis PCA is performed on the feature matrix X , the variance contribution rate of the principal components is calculated, the first 5 principal components with cumulative variance contribution rate ≥85% are selected to form the dimension-reduced feature matrix X pca ; S46: through iterative calculation, latent variables X pca are extracted from t 1, u 1, t 1, u 1, 1, X pca and Y on t 1, u 1, the regression residual matrix E1, F1, if the residual variance ratio is >10%, then E1, F1 are taken as new X, Y, and the extraction of t 2, u 2 is repeated until the residual variance ratio is ≤10%, and finally 3 latent variables are extracted; characteristic importance evaluation: that is, the latent variable importance projection value VIP value is calculated.

[0012] Preferably, the fusion model of step S5 includes an input layer, a random forest sub-model RF, a support vector machine sub-model SVM, a fusion decision layer, and an output layer; The input layer selects 6 key characteristic parameters to construct a feature matrix X; The random forest sub-model reduces the risk of overfitting through ensemble learning of multiple decision trees; The number of decision trees: 100-150; Feature sampling strategy: randomly select 3-4 features each time to build the decision tree; Node splitting rule: adopt the principle of minimum Gini coefficient; Support vector machine sub-model: convert the linearly inseparable feature data to high-dimensional space through kernel function mapping, and construct the optimal classification hyperplane; Kernel function: select radial basis kernel function to adapt to the nonlinear distribution of pesticide residues; Fusion decision layer for weighted fusion: Weight determination logic: based on the accuracy of the sub-model on the validation set to allocate weights; Output layer output screening result: positive or negative; Confidence: probability after fusion P final ×100%.

[0013] The beneficial effects of the present application include: 1. The detection efficiency is significantly improved, which meets the real-time screening demand: according to the moisture content grading selection of pulverization method, and through the program temperature headspace extraction to shorten the extraction time, the total consumption time of pretreatment is less than or equal to 30 minutes; detection and data processing are rapid: adopt double-channel data acquisition mode, realize data rapid pretreatment through adaptive polynomial fitting and improved wavelet transform, reduce the inference time of the fusion model, the overall screening process consumes less time, and the efficiency is obviously improved compared with the existing laboratory detection method and the traditional electronic nose method.

[0014] 2. The detection accuracy is greatly improved, and the recognition ability of low concentration residue is strong: through the impedance matching algorithm, the difference between the response thresholds of adjacent sensors is controlled at 8%-12%, and the overall response deviation coefficient of the array is less than or equal to 3%, which effectively avoids the overlapping of sensor signals, and improves the signal discrimination degree compared with the traditional electronic nose array; adaptive polynomial fitting can reduce the baseline drift correction error, and improved wavelet transform can remove environmental noise at a high rate, providing high-quality data for feature extraction; the SVM-random forest fusion model combines the advantages of high-dimensional small sample classification of SVM and the anti-overfitting advantage of RF, and through weighted voting fusion, the recognition accuracy of low concentration pesticide residues is high, the overall pesticide residue recognition accuracy is high, the accuracy is improved compared with single SVM model, and the false negative and false positive misjudgment rate is effectively reduced.

[0015] 3. Through the grading pulverization and headspace extraction parameters, different moisture content, different matrix of leafy vegetables, melons and fruits and other agricultural products can be adapted, and the problem of detection distortion of traditional methods for high-moisture samples is solved.

[0016] 4. Realize hierarchical processing according to the confidence output by the fusion model: the confidence is greater than or equal to the first threshold value 90%, directly output the result, 70%-90% automatically trigger secondary detection, and less than 70% prompt to re-sample, through the "secondary verification" and "abnormal early warning", ensure the reliability of the detection result, compared with the existing method without confidence determination, the result credibility is improved. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 It is a flowchart of the pesticide residue rapid screening method based on the optimized electronic nose array of the application.

[0018] Figure 2 It is a flowchart of the baseline correction using the adaptive polynomial fitting algorithm of the application.

[0019] Figure 3 It is an architecture diagram of the fusion model of the application. DETAILED DESCRIPTION

[0020] The application will be further described below in combination with the accompanying drawings Figures 1-3 The application will be further described below in combination with the accompanying drawings Example 1 Referring to the accompanying drawings Figure 1 A pesticide residue rapid screening method based on an optimized electronic nose array, comprising the following steps: S1: Select a plurality of different sensitive material gas sensors to form a basic sensor group, and optimize the array arrangement of the sensors through an impedance matching algorithm; meanwhile, a porous nano coating is covered on the surface of the sensor, the pore size of the coating is 50-100 nm, which is used to enhance the adsorption selectivity to the volatile components of pesticides, and the proportion of the metal oxide semiconductor material is 45%-55%; S2: Pretreatment of the sample to be detected: The sample to be detected (vegetables / fruits) is subjected to hierarchical processing, and the crushing method is selected according to the water content of the sample, low-temperature crushing is used when the water content is greater than 85%, normal-temperature crushing is used when the water content is less than or equal to 85%, the crushed sample is sieved through a 50-mesh standard sieve, and 2-5 g of the undersize sample is placed in a sealed extraction bottle, and deionized water containing 0.1% antioxidant is added as an extraction liquid at a ratio of 1:8 (g / mL).

[0021] Programmed temperature headspace extraction is used, first constant temperature pretreatment at 30°C for 5 min, then heating to 50°C at a rate of 5°C / min, constant temperature extraction for 20 min, and at the same time, a blank control group (only extraction liquid) and three concentration gradient standard pesticide residue sample groups (0.005 mg / kg, 0.01 mg / kg, 0.05 mg / kg) are set, to ensure that the coefficient of variation of the pesticide component concentration in the upper gas phase of the extraction liquid after pretreatment is less than or equal to 5%.

[0022] S3: Acquisition of dynamic response data: The upper gas phase of the extraction bottle is introduced into the optimized electronic nose array at a constant rate of 80 mL / min. The detection environment is maintained at a temperature of 25±1℃ and a humidity of 40±2%. A dual-channel data acquisition mode is adopted to simultaneously acquire the resistance change signal of the sensor within 0-300s (sampling frequency 1Hz) and the time of occurrence of the response peak. When the response value of any sensor reaches a steady state (resistance change rate ≤0.5% for 10 consecutive seconds), the data acquisition termination program is automatically triggered to ensure the integrity and timeliness of the acquired data.

[0023] S4: Preprocess the dynamic response data: First, use an adaptive multinomial fitting algorithm (automatically adjust the fitting order to 3-5 according to the trend of the response curve) to perform baseline correction, and then use an improved wavelet transform algorithm (selecting the db4 wavelet basis function and decomposing 3 layers) to remove environmental interference noise.

[0024] By combining time series features and signal amplitude features for dual extraction, and using principal component analysis (PCA) for dimensionality reduction, partial least squares discriminant analysis (PLS-DA) was used to screen out six key feature parameters: response peak, response time, recovery time, steady-state response value, and slope at the peak. This ensures that the cumulative variance contribution rate of the screened feature parameters is ≥90%.

[0025] S5: Divide the selected key feature parameters into training set, validation set and final feature set in a 7:2:1 ratio. Input the training set and validation set into the fusion model of support vector machine (SVM) and random forest for training, validation and optimization. The SVM uses radial basis kernel function and the random forest is set to have 100-150 decision trees.

[0026] S6: Rapid screening and result determination of pesticide residues: The final feature set of the agricultural product sample to be tested is input into the optimized screening model. The model automatically outputs the positive / negative judgment result and the confidence level, with the confidence level ranging from 0 to 100%.

[0027] When the confidence level is ≥90%, the judgment result is output directly; when the confidence level is between 70% and 90%, the secondary detection procedure is automatically triggered, repeating steps S3 and S4, and the average value of the two feature parameters is re-inputted into the model; when the confidence level is <70%, an abnormality in sample pretreatment is indicated and resampling is performed for screening.

[0028] In this embodiment, the sensitive material in step S1 includes a metal oxide semiconductor material doped with rare earth elements, a functionalized conductive polymer material, and a modified graphene material. The rare earth elements include La and Ce, the metal oxide semiconductor material includes SnO2, ZnO, and WO3, the functionalized conductive polymer material is a polypyrrole-carbon nanotube composite material, and the modified graphene material is a graphene-titanium dioxide-silver nanoparticle composite material.

[0029] Example 2 Based on Example 1, in step S1, the sensor array arrangement is optimized using an impedance matching algorithm to control the difference in the sensitive material response threshold between adjacent sensors within 8%-12%, and the overall array response deviation coefficient ≤3%. The specific process is as follows: S11: Establish a sensor impedance-response threshold database; The specific process of step S11 is as follows: S111: Sensor Impedance and Response Threshold Test: Impedance tests were conducted on various selected sensitive material sensors under a standard pesticide gas environment with a concentration of 0.01 mg / kg. A precision impedance analyzer was used to test the frequency range of 100Hz-10kHz, and the impedance values ​​of each sensor at different frequencies (Z1, Z2...Z...) were recorded. n ); The response thresholds of each sensor are recorded synchronously, i.e., the pesticide gas concentration when the sensor resistance change rate reaches 10%. A single-sensor response threshold database (T1, T2...T...) is obtained through gradient concentration experiments for calibration. n ); S112: Impedance matching target parameter setting: Set the algorithm optimization goal: The difference in response threshold between adjacent sensors: ΔT = |T i -T i+1 |∈[8%,12%]; Measured as a percentage difference in response thresholds, such as T i =0.008mg / kg, T i+1 =0.009mg / kg, then ΔT=12.5%; Overall array response deviation coefficient: Cᵥ = (Standard deviation / mean value of response thresholds of all sensors in the array) × 100% ≤ 3%; Impedance matching degree: K = (Similarity of impedance values ​​between adjacent sensors / Preset similarity threshold) × 100% ≥ 90%; The similarity threshold was determined through preliminary experiments to ensure minimal signal interference between sensors.

[0030] S12: Iterative arrangement based on impedance matching algorithm.

[0031] The specific process of step S12 is as follows: S121: Sensor clustering and grouping based on impedance values: Using the K-means clustering algorithm, the impedance value of a single sensor at a frequency of 1kHz is used as the clustering feature (1kHz is the sensitive frequency of the response of volatile components of pesticides, which was determined through previous frequency scanning experiments). 10-12 sensors are divided into 3-4 clusters, with 3-4 sensors in each group, to ensure that the similarity of the sensor impedance values ​​within the same cluster is ≥90%, which meets the requirement of impedance matching degree K.

[0032] For example, cluster group 1 contains sensors with impedance values ​​between 1000-1050Ω, such as La-doped SnO2 and graphene-titanium dioxide-silver nanoparticle composite sensors; cluster group 2 contains sensors with impedance values ​​between 800-850Ω, such as polypyrrole-carbon nanotube composites and Ce-doped ZnO sensors, to avoid signal crosstalk caused by excessive impedance differences. S122: In-group sorting and preliminary arrangement based on response threshold: For the sensors within each cluster, sort them in ascending order of response threshold, T1 < T2 < ... < T m m is the number of sensors in the group; Calculate the response threshold difference ΔT between adjacent sensors. If ΔT is not within the range of 8%-12%, adjust by replacing the sensor: select a sensor with impedance similarity ≥90% and whose response threshold can fill the difference from other clusters and replace the sensor in the current group that does not meet the requirements. For example, if the ΔT of sensor A (T=0.005mg / kg) and sensor B (T=0.006mg / kg) in cluster group 1 is 20% (exceeding the upper limit), then sensor C (T=0.0054mg / kg, impedance value 820Ω, impedance similarity with sensor A 92%) is selected from cluster group 2 to replace sensor B, so that ΔT=8% (meets the requirements). S123: Iterative Adjustment and Verification of Overall Deviation Coefficient: After sorting the sensors in each cluster, they are arranged into a preliminary array according to the principle of "inter-group alternation", such as cluster group 1 sensor 1 → cluster group 2 sensor 1 → cluster group 3 sensor 1 → cluster group 1 sensor 2 → ..., and the overall response deviation coefficient Cᵥ of the array is calculated. If Cᵥ > 3%, then optimize through "local fine-tuning": replace the sensor whose response threshold deviates the most from the average value with a sensor whose response threshold is closer to the average value within the same cluster and whose adjacent difference still meets the 8%-12% requirement, and repeat the calculation of Cᵥ until Cᵥ ≤ 3%.

[0033] Iteration termination condition: After three consecutive adjustments, the ΔT of adjacent sensors is within the range of 8%-12%, and Cᵥ≤3%. At this time, the final array arrangement order is output, such as the sensor number: S3→S7→S1→S9→S5→S2→S10→S4→S8→S6, which corresponds to the optimal arrangement of 10 sensors.

[0034] Example 3 Based on Example 1 or Example 2, see Figure 2 As shown, the specific process of baseline correction using the adaptive polynomial fitting algorithm in step S4 is as follows: S41: Record the dynamic response data obtained in step S3 as R(t), and store it in the format of a two-dimensional array of time t-resistance change rate R, such as R=0% when t=0s, R=8% when t=50s, R=15% when t=200s, and R=14.8% when t=300s.

[0035] Electronic nose sensors commonly exhibit drift in pesticide detection, which can be categorized into two types: Slow linear drift: The baseline rises / falls at a constant rate over time, such as the baseline drifting from 0% to 1% within 0-300s, with a slope ≤0.003% / s.

[0036] Fluctuating drift: The baseline exhibits local fluctuations, such as ±0.5% fluctuations within 100-150 seconds due to minor changes in ambient airflow during the mid-detection period.

[0037] The degree of nonlinearity of the drift is determined by calculating the first derivative (ΔR / Δt) and second derivative (Δ²R / Δt²) of the original curve. A rate of change of the first derivative of less than 5% indicates weak nonlinearity, and a low-order fitting is performed. A change rate ≥ 5% indicates strong nonlinearity, requiring high-order fitting. S42: Adaptive polynomial order determination and fitting calculation: The "piecewise error evaluation-order iterative optimization" method is used to automatically select the optimal order among polynomials of orders 3-5. The specific process of step S42 is as follows: S421: Curve Segmentation and Baseline Interval Extraction The original response curve is divided into three segments according to the "detection stage": Baseline stability segment (0-30s): The sensor is not in contact with pesticide gas, and the curve is in the initial stable state. Data in this interval is used as the baseline reference. Response rise phase (31-150s): After the sensor comes into contact with pesticide gas, the rate of change of resistance rises rapidly to its peak value; Response stability period (151-300s): After the resistance change rate reaches its peak, it remains in the steady state range (fluctuation ≤1%). Baseline candidate points were extracted from the three segments: all 31 points were taken from the stable baseline segment, and 10 stationary points were taken from each of the rising and stable response segments (these were selected using the sliding window method, with a window size of 5s and a standard deviation of ≤0.2% as stationary points), for a total of 51 baseline candidate points. S422: Multi-order polynomial fitting and error calculation: The 51 baseline candidate points were fitted using 3rd, 4th, and 5th order polynomials, respectively. The general form of the polynomials is as follows: B ( t )= a n t n +a n−1 t n−1 +...+a 1 t + a 0; in n =3, 4, 5 are the polynomial orders. a n ...a 0 The fitting coefficients are obtained using the least squares method. The root mean square error (RMSE) and maximum absolute error (MAE) of each order of fit are calculated using the following formulas: ; ; m =51 represents the number of baseline candidate points. R base ( t i () represents the rate of change of resistance at the baseline candidate points in the original curve. B ( t i ) is the fitted baseline at t i The calculated value at time; S423: Adaptive order selection: Set the error judgment thresholds as follows: RMSE ≤ 0.1% and MAE ≤ 0.05% are the optimal fitting criteria. Select the order according to the following rules: If a third-order fit satisfies the error threshold and the rate of change of the second derivative of the fitted curve is less than 3% (indicating that the curve is smooth and there is no overfitting), then a third-order fit is selected. If the RMSE of the 3rd-order fit is greater than 0.1% or the MAE is greater than 0.05%, but the 4th-order fit meets the error threshold and the rate of change of the third derivative is less than 2%, then the 4th-order fit is selected. If neither the 3rd nor the 4th order fitting meets the error threshold, or if the original curve exhibits strong nonlinear drift (the rate of change of the second derivative is ≥5%), then the 5th order fitting should be selected. Example: The original curve of a certain sensor exhibits fluctuating drift (MAE=0.08%), the RMSE of the 3rd-order fit is 0.12% (not satisfied), the RMSE of the 4th-order fit is 0.09%, MAE is 0.04% (meets the threshold), and the rate of change of the third derivative is 1.5% (no overfitting). Therefore, the 4th-order polynomial is automatically selected.

[0038] S43: Baseline Correction Execution: Raw Data Correction and Result Optimization The specific process of step S43 is as follows: S431: Baseline subtraction of raw data: The corrected response value is obtained by subtracting the corresponding baseline value B(t) from the resistance change rate R(t) at each time point in the original response curve. R corr (t), the formula is: R corr (t)=R(t)-B(t); For example, at t=100s, the original R(t)=12% and the fitted baseline B(t)=0.6%, then the corrected R_corr(t)=11.4%, eliminating the interference of baseline drift on the response signal.

[0039] S432: Verification of calibration results: The corrected curves are then verified to ensure that the correction effect meets the following requirements: The corrected data for the baseline stability period (0-30s) have a mean of ≤0.1% and a standard deviation of ≤0.05% (recovery to the ideal baseline level). Response peak deviation ≤2% (the difference between the corrected peak value and the theoretical peak value, which is calibrated using standard pesticide samples). If the corrected data does not meet the requirements, the order determination will be automatically re-executed (e.g., adjusting the 4th order to the 5th order) until it meets the standard.

[0040] Environmental interference noise was removed by improving the wavelet transform algorithm (using the db4 wavelet basis function and decomposing into 3 layers): Environmental interference noise types and feature identification: During electronic nose detection, environmental interference noise mainly comes from three sources: Airflow fluctuation noise: High-frequency noise caused by small fluctuations in the airflow rate (±5mL / min) when headspace extraction gas is introduced, with a frequency range of 10-20Hz and a signal amplitude ≤0.5% (measured by the rate of change of resistance). Electromagnetic interference noise: Electromagnetic noise generated by peripheral electronic equipment (such as temperature control unit and data acquisition unit) of the detection device, with a frequency concentrated at 50Hz (power frequency interference) and an amplitude ≤0.3%; Sensor intrinsic noise: Low-frequency fluctuations caused by thermal noise of the sensor's sensitive material, with a frequency < 1 Hz and an amplitude ≤ 0.2%; The baseline-corrected response signal was obtained by Fast Fourier Transform (FFT). R corr (t), derived from baseline correction results, are used to perform spectral analysis to determine the frequency distribution and amplitude range of the noise, providing a target basis for subsequent wavelet decomposition.

[0041] Signal preprocessing: Data normalization and sampling frequency matching: The baseline-corrected response signal is normalized to map the resistance change rate to the [0,1] interval, as shown in the formula: ; in R corr,max 、R corr,min These are the maximum and minimum values ​​of the corrected signal, respectively, to avoid the signal amplitude difference affecting the noise reduction effect; Three-level wavelet decomposition based on db4 wavelet basis: The Mallat algorithm is used to normalize the signal. R norm ( t Perform 3-level wavelet decomposition, the specific process is as follows: Level 1 decomposition: R norm ( t It is decomposed into the first-level approximation coefficient A1 and the first-level detail coefficient D1.

[0042] Among them, A1 corresponds to the low-frequency component of the signal (mainly the slow variation trend of the effective signal, frequency <0.25Hz), and D1 corresponds to the high-frequency noise (mainly airflow fluctuation noise, 10-20Hz). Second-level decomposition: The first-level approximation coefficient A1 is further decomposed to obtain the second-level approximation coefficient A2 and the second-level detail coefficient D2. A2 corresponds to the lower frequency effective signal trend (frequency < 0.125Hz), and D2 corresponds to the intermediate frequency noise (mainly electromagnetic interference noise, the low-frequency component of 50Hz power frequency interference after sampling frequency folding, 2-5Hz). The third-level decomposition: The second-level approximation coefficient A2 is decomposed again to obtain the third-level approximation coefficient A3 and the third-level detail coefficient D3. A3 corresponds to the core effective component of the signal (frequency < 0.0625Hz, such as the trend of the response rising and steady segment), and D3 corresponds to the low-frequency noise (mainly the sensor's own thermal noise, < 1Hz). The coefficient set obtained after decomposition is {A3, D3, D2, D1}, where A3 is the core carrier of the effective signal, and D1-D3 correspond to environmental interference noise of different frequencies.

[0043] Improved threshold function and noise figure suppression: Adaptive threshold calculation: For the detail coefficients (D1-D3) of different decomposition layers, the Stein unbiased risk estimation (SURE) is used to adaptively calculate the threshold, avoiding over-denoising or incomplete denoising caused by a fixed threshold. For D1 (high-frequency airflow noise): Threshold ; in The standard deviation of the D1 coefficient is calculated using the median absolute deviation method: ; N is the number of data points, N = 304; For D2 (intermediate frequency electromagnetic noise): threshold λ2 = 0.8λ1; the electromagnetic noise amplitude is lower than the airflow noise, so the threshold is lowered to enhance the suppression effect; For D3 (low-frequency thermal noise): threshold λ3 = 0.5λ1; the thermal noise amplitude is minimized, and the threshold is further reduced to avoid loss of effective signal.

[0044] Improved threshold function processing: An improved threshold function combining hard and soft thresholding is used to process the detail coefficients: When the absolute value of the detail factor : Determined as a valid signal component, the original value of the coefficient is retained (hard threshold characteristic, to avoid smoothing distortion of the valid signal); When the absolute value of the detail factor (Below the threshold): Determined as noise component, suppressed using a soft thresholding function: (The exponential decay function, compared to the linear decay of the traditional soft threshold, can suppress noise more smoothly and reduce signal oscillation.) After processing, a new set of detail coefficients {D1', D2', D3'} is obtained, in which noise components are significantly suppressed and effective signal components are preserved.

[0045] Inverse wavelet transform reconstructs the signal: The Mallat inverse algorithm is used to reconstruct the denoised response signal by combining the processed approximation coefficients A3 with the new detail coefficients D1', D2', and D3'. R denoise (t): Step 1: Inverse transform A3 and D3' to obtain A2'; Step 2: Inverse transform A2' and D2' to obtain A1'; Step 3: Inverse transform A1' and D1' to obtain R denoise (t); The reconstructed signal is then denormalized to restore it to the original resistance change rate range. The formula is as follows: R final (t)= R denoise (t)×( R corr,max −R corr,min )+ R corr,min .

[0046] In this embodiment, the specific process of screening key feature parameters through partial least squares discriminant analysis (PLS-DA) in step S4 is as follows: S44: Response signal after denoising with improved wavelet transform R final Based on (t), 12 initial candidate feature parameters are extracted, covering the three dimensions of signal amplitude, time, and slope, as shown in Table 1 below: Table 1 Initial Candidate Feature Parameter Table Sample set division: Three types of samples were selected: blank control group, low concentration pesticide residue group (0.005 mg / kg), and high concentration pesticide residue group (0.05 mg / kg). Each type of sample had 30 replicates, for a total of 90 samples. Feature matrix and category vector construction: Construct feature matrix X (90×12) with 12 candidate features as columns and 90 samples as rows; construct category vector Y (90×1) with sample categories as labels (blank group=0, low concentration group=1, high concentration group=2). S45: PCA dimensionality reduction preprocessing: For the characteristic matrix X Principal component analysis (PCA) was performed to calculate the variance contribution rate of the principal components (PCs). The top 5 principal components (PC1-PC5) with a cumulative variance contribution rate ≥ 85% were selected to form the dimensionality-reduced feature matrix. X pca(90×5) to reduce multicollinearity among features, such as the correlation between response time and peak occurrence time, to provide low-redundancy input data for the PLS-DA model and improve model stability.

[0047] S46: Extracting Latent Variables (LV): Through iterative calculation, from... X pca Extracting latent variables t 1 (First Latent Variable): Extract latent variables from Y. u 1. Make t 1 and u The covariance of 1 is the largest, i.e., Cov( t 1, u 1)→max.

[0048] Residual Analysis and Iteration: Calculation X pca With Y t 1, u If the residual variance percentage of the regression residual matrix E1 and F1 is greater than 10%, then E1 and F1 are used as new X and Y, and the extraction is repeated. t 2. u 2. Continue until the residual variance ratio is ≤10%, and finally extract 3 latent variables (LV1-LV3).

[0049] Feature importance assessment: This involves calculating the variable importance projection value (VIP). The VIP value reflects the contribution of each candidate feature to the model's classification ability. The calculation formula is as follows: ; in, K The total number of candidate features. m The number of latent variables, R 2 (Y; t k ) is the first k The explanatory power of each latent variable for Y w jk For the first j The feature in the first k Weights on each latent variable.

[0050] Rule: Features with a VIP value > 1 are considered to make a significant contribution to the model's classification (industry-standard threshold). The larger the VIP value, the stronger the feature's discriminative power.

[0051] Examples of VIP values ​​for the 12 candidate features are shown in Table 2 below: Table 2. Example of VIP values ​​for 12 candidate features Example 4 Based on Example 1, Example 2, or Example 3, see [link to example]. Figure 3 As shown, the fusion model in step S5 includes an input layer, a random forest sub-model (RF), a support vector machine sub-model (SVM), a fusion decision layer, and an output layer. The six key feature parameters selected by the input layer are: peak response P, response time T1, recovery time T2, steady-state response value S, and peak slope K. p The peak percentage correction value P / S' is used to construct the feature matrix X; The random forest sub-model reduces the risk of overfitting through ensemble learning of multiple decision trees and is good at handling nonlinear features and imbalanced data. Number of decision trees: 100-150; Feature sampling strategy: Randomly select 3-4 features (50%-67% of the total number of features) each time a decision tree is built to enhance the diversity between trees; Node splitting rule: The principle of minimizing the Gini coefficient (Gini=1-Σpᵢ², where pᵢ is the proportion of a certain type of sample within the node) is adopted to ensure the highest classification purity of the node; Support Vector Machine (SVM) sub-model: It transforms linearly inseparable feature data into a high-dimensional space through kernel function mapping, and constructs the optimal classification hyperplane. It is good at accurate classification of small sample and high-dimensional data.

[0052] Key parameter optimization: Kernel function: Radial basis function (RBF) is selected to adapt to the nonlinear distribution of pesticide residue characteristics (such as the nonlinear relationship between response peak and residue concentration). Weighted fusion of decision-making levels: Weight determination logic: Weights are assigned based on the accuracy of the sub-models on the validation set, using the following formula: ω R F = Acc R F / ( Acc R F + Acc S VM ); ω S VM = Acc S VM / ( Acc R F + Acc S VM ); in, Acc R F For the accuracy of the RF validation set, Acc S VM For the SVM validation set accuracy, then ω R F ≈0.497, ω S VM ≈0.503; When RF has high accuracy in identifying low-concentration residual samples, and SVM has high accuracy in identifying high-concentration residual samples, weighted fusion can achieve high accuracy in identifying samples across the entire concentration range.

[0053] The output layer outputs the screening results: positive (indicating the presence of pesticide residues) or negative (indicating the absence of pesticide residues). Confidence level: probability after fusion P final ×100%, such as P final =0.85 corresponds to an 85% confidence level.

Claims

1. A rapid pesticide residue screening method based on an optimized electronic nose array, characterized in that, Includes the following steps: S1: Select a variety of gas sensors with different sensitive materials to form a basic sensor group, and optimize the array arrangement of the sensors through an impedance matching algorithm; S2: Grading the agricultural product samples to be tested, selecting the pulverization method according to the sample moisture content, sieving after pulverization, taking a specified amount of the sieved sample and placing it in a sealed extraction bottle, adding the extraction solution according to the specified ratio for extraction. S3: Acquire dynamic response data: The upper gas phase of the extraction bottle is introduced into the optimized electronic nose array at a specified rate. A dual-channel data acquisition mode is used to simultaneously acquire the resistance change signal and the time of response peak occurrence of the sensor within a specified time period. S4: Preprocess the dynamic response data: first perform baseline correction, then remove environmental interference noise; combine time series features and signal amplitude features for dual extraction to select specified key feature parameters; S5: Divide the selected key feature parameters into training set, validation set and final feature set according to the specified ratio, and input the training set and validation set into the fusion model of support vector machine and random forest for model training, optimization and verification. S6: Input the final feature set into the optimized screening model and output the positive / negative judgment result and confidence level; when the confidence level is ≥ the first threshold, output the judgment result directly; when the confidence level is between the second threshold and the first threshold, automatically trigger the secondary detection program to repeat steps S3 and S4, take the average value of the two feature parameters and re-input it into the model; when the confidence level is < the second threshold, indicate that the sample preprocessing is abnormal and re-sample for screening.

2. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 1, characterized in that, The sensitive materials in step S1 include metal oxide semiconductor materials doped with rare earth elements, functionalized conductive polymer materials, and modified graphene materials. The rare earth elements include lanthanum and cerium, the metal oxide semiconductor materials include SnO2, ZnO, and WO3, the functionalized conductive polymer materials are polypyrrole-carbon nanotube composite materials, and the modified graphene materials are graphene-titanium dioxide-silver nanoparticle composite materials.

3. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 2, characterized in that, In step S1, a basic sensor group is composed of gas sensors with different sensitive materials. The specific process of optimizing the array arrangement of the sensors using an impedance matching algorithm is as follows: S11: Establish a sensor impedance-response threshold database; S12: Iterative arrangement based on impedance matching algorithm.

4. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 3, characterized in that, The specific process of step S11 is as follows: S111: Sensor Impedance and Response Threshold Test: Impedance tests were conducted on various selected sensitive material sensors under a standard pesticide gas environment, and the impedance values ​​of each sensor at different frequencies were recorded; the response thresholds of each sensor were recorded simultaneously to obtain a single sensor response threshold database. S112: Set the target parameters for impedance matching, including the difference in response thresholds between adjacent sensors, the overall response deviation coefficient of the array, and the impedance matching degree.

5. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 3, characterized in that, The specific process of step S12 is as follows: S121: Sensor clustering based on impedance values: Using the K-means clustering algorithm, the impedance value of a single sensor at a frequency of 1kHz is used as the clustering feature to divide 10-12 sensors into 3-4 clusters, ensuring that the similarity of sensor impedance values ​​within the same cluster is ≥90%; S122: Intra-group sorting and preliminary arrangement based on response threshold: For sensors in each cluster, sort them in ascending order of response threshold, calculate the response threshold difference ΔT between adjacent sensors, and if ΔT is not within the range of 8%-12%, adjust by sensor replacement: Select sensors from other clusters with impedance similarity ≥90% and response threshold that can fill the difference, and replace the sensors in the current group that do not meet the requirements. S123: Iterative adjustment and overall deviation coefficient verification: After sorting the sensors in each cluster group, a preliminary array is formed according to the principle of inter-group alternation, and the overall response deviation coefficient Cᵥ of the array is calculated. If Cᵥ > 3%, then optimize through local fine-tuning: for the sensor whose response threshold deviates the most from the average value, replace it with a sensor in the same cluster whose response threshold is closer to the average value and whose adjacent difference still meets 8%-12%, and repeat the calculation of Cᵥ until Cᵥ ≤ 3%.

6. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 1, characterized in that, The specific process of baseline correction using the adaptive polynomial fitting algorithm in step S4 is as follows: S41: Record the dynamic response data obtained in step S3 as R(t) and store it in the format of a two-dimensional array of time t - resistance change rate R; determine the degree of nonlinearity of the drift by calculating the first and second derivatives of the original curve: the first derivative change rate < 5% is weak nonlinearity, and low-order fitting is performed; the change rate ≥ 5% is strong nonlinearity, and high-order fitting is performed. S42: Perform adaptive polynomial order determination and fitting calculation: Use the piecewise error evaluation-order iterative optimization method to select the optimal order among polynomials of order 3-5. S43: Perform baseline correction: Correct the original data and optimize the results.

7. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 6, characterized in that, The specific process of step S42 is as follows: S421: Divide the original response curve into three segments according to the detection stage, including the baseline stable segment, the response rising segment, and the response stable segment. Extract baseline candidate points from the three segments respectively: take all 31 points for the baseline stable segment, and take 10 stable points for each of the response rising segment and the stable segment, for a total of 51 baseline candidate points. S422: The 51 baseline candidate points were fitted using 3rd, 4th, and 5th order polynomials, respectively, and the root mean square error (RMSE) and maximum absolute error (MAE) of each order of fitting were calculated. S423: Perform adaptive order selection: Set error judgment thresholds: RMSE ≤ 0.1% and MAE ≤ 0.05% are the optimal fitting criteria, and select the order according to the following rules: If a third-order fit satisfies the error threshold and the rate of change of the second derivative of the fitted curve is less than 3%, then a third-order fit is selected. If the RMSE of the 3rd-order fit is greater than 0.1% or the MAE is greater than 0.05%, but the 4th-order fit meets the error threshold and the rate of change of the third derivative is less than 2%, then the 4th-order fit is selected. If neither the 3rd nor the 4th order fitting meets the error threshold, or if the original curve exhibits strong nonlinear drift, then the 5th order fitting should be selected.

8. The method for rapid screening of pesticide residues based on an optimized electronic nose array according to claim 6, characterized in that, The specific process of step S43 is as follows: S431: Baseline subtraction of raw data: The corrected response value is obtained by subtracting the corresponding baseline value B(t) from the resistance change rate R(t) at each time point in the original response curve. R corr (t), the formula is: R corr (t)=R(t)-B(t); S432: Verification of calibration results: The mean of the calibrated data in the baseline stable segment is ≤0.1%, the standard deviation is ≤0.05%, and the peak deviation of the response is ≤2%. If the calibrated data does not meet the requirements, the order determination will be automatically re-executed until it meets the standard.

9. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 1, characterized in that, The specific process for selecting key feature parameters in step S4 is as follows: S44: Based on the response signal after denoising by improved wavelet transform, 12 initial candidate feature parameters are extracted; Three classes of samples were selected, with 30 replicates of each class, for a total of 90 samples; a feature matrix X was constructed with 12 candidate features as columns and 90 samples as rows. Construct a category vector Y using the sample category as the label; S45: For the characteristic matrix X Principal component analysis was performed, and the variance contribution rate of each principal component was calculated. The top 5 principal components with a cumulative variance contribution rate ≥ 85% were selected to form the dimensionality-reduced feature matrix. X pca ; S46: Through iterative calculation, from X pca Extracting latent variables t 1. Extract latent variables from Y u 1. Make t 1 and u The covariance is largest at 1; calculate X pca With Y t 1, u If the residual variance percentage of the regression residual matrix E1 and F1 is greater than 10%, then E1 and F1 are used as new X and Y, and the extraction is repeated. t 2. u 2. Continue until the residual variance percentage is ≤10%, and finally extract 3 latent variables; Perform feature importance assessment, i.e. calculate the latent variable importance projection value (VIP value).

10. The method for rapid pesticide residue screening based on an optimized electronic nose array according to claim 1, characterized in that, The fusion model in step S5 includes an input layer, a random forest sub-model (RF), a support vector machine sub-model (SVM), a fusion decision layer, and an output layer. The six key feature parameters filtered by the input layer are used to construct a feature matrix X; The random forest sub-model reduces the risk of overfitting through ensemble learning of multiple decision trees; Number of decision trees: 100-150; Feature sampling strategy: Randomly select 3-4 features each time a decision tree is built; Node splitting rule: adopts the principle of minimizing the Gini coefficient; Support Vector Machine sub-model: Transforms linearly inseparable feature data into a high-dimensional space through kernel function mapping to construct the optimal classification hyperplane; Kernel function: A radial basis function is selected to adapt to the nonlinear distribution of pesticide residue characteristics; Weighted fusion of decision-making levels: Weight determination logic: Weights are assigned based on the accuracy of the sub-models on the validation set; The output layer outputs the screening results: positive or negative. Confidence level: probability after fusion P final ×100%.

Citation Information

Patent Citations

  • Electronic nose gas sensor array optimization method based on dynamic characteristic importance

    CN109799269A

  • Method for detecting disinfectant concentration in cold chain environment based on electronic nose

    CN117169441A

  • Cut tobacco dryer predictive control method based on multi-model fusion

    CN121128952A

  • Tire wear particle environment fate prediction model construction method based on multi-dimensional data

    CN121351024A

  • Virtual impedance auto matching box

    KR101989518B1