Fresh tobacco leaf cross-year pesticide residue concentration detection method and detection system based on feature transfer learning

Through the feature transfer learning method, a cross-year transfer learning optimization model was constructed, which solved the problem of insufficient generalization ability of the tobacco leaf pesticide residue discrimination model under different environmental conditions and achieved the accuracy and robustness of cross-year fresh tobacco leaf pesticide residue concentration detection.

CN120833552APending Publication Date: 2025-10-24CHINA NAT TOBACCO CORP SHANDONG BRANCH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510923602.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

The existing tobacco leaf pesticide residue discrimination model has insufficient generalization ability under different environmental conditions, resulting in inaccurate detection results and making it difficult to meet the needs of field pesticide use supervision.

Method used

A method based on feature transfer learning was adopted. By obtaining hyperspectral image data of fresh tobacco leaf samples from the same producing area for two consecutive years, spectral correction and noise filtering preprocessing were performed. The GA-LSSVM model was used to train the pesticide residue concentration detection model. The feature distributions of the source domain and the target domain were aligned through the TCA and MEDA joint algorithm to generate a cross-year transfer learning optimization model.

Benefits of technology

The accuracy and robustness of pesticide residue concentration detection in fresh tobacco leaves across years were achieved, and the generalization ability of the model was improved. In particular, the accuracy rate reached 89.33%~83.33% in the five pesticide detection tasks, effectively overcoming the impact of year differences on model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120833552A_ABST
    Figure CN120833552A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of tobacco pesticide residue detection, and discloses a fresh tobacco cross-year pesticide residue concentration detection method based on feature transfer learning, and the method comprises the following steps: obtaining hyperspectral image data of fresh tobacco samples in two consecutive years in the same production area, and taking the hyperspectral image data as a source domain data set and a target domain data set; a feature importance evaluation algorithm is adopted to carry out dimensionality reduction on source domain full-spectrum data, screening features are input into a GA-LSSVM detection model, and a GA-LSSVM pesticide residue concentration detection model is obtained through training; a TCA and MEDA combined algorithm is adopted to align feature distribution of a source domain and a target domain, adapted features are input into the GA-LSSVM pesticide residue concentration detection model, and a cross-year transfer learning optimization model is generated; and inputting hyperspectral data of fresh tobacco leaves to be detected into the cross-year transfer learning optimization model, and outputting a pesticide residue concentration value. The invention provides a reliable transfer learning solution for cross-year detection of the pesticide residue concentration level of the fresh tobacco leaves.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of tobacco leaf pesticide residue detection, and relates to a fresh tobacco leaf cross-year pesticide residue concentration detection method and a detection system based on feature transfer learning. BACKGROUND

[0002] Chemical pesticides have irreplaceable effects in tobacco pest control due to their convenience, economy and high efficiency. However, due to the problems of product quality and safety caused by the illegal and excessive use of chemical pesticides, the use of pesticides in tobacco has become a key point in the management and control of product quality and safety. At present, for the detection of pesticide use in tobacco, traditional and conventional chemical analysis methods (such as LC-MS, GC-MS, HPLC, etc.) have the advantages of high detection accuracy and low detection limit, but are limited by the disadvantages of complex pretreatment, high threshold requirement and long detection time, and are difficult to meet the needs of field pesticide supervision in tobacco production areas. The rapid test strips developed based on enzyme inhibition method and colloidal gold immunoassay have limited pesticide detection types, are easily affected by the detection environment to produce false positives, and can only detect one pesticide at a time. The detection speed, efficiency and cost are still difficult to meet the needs of field pesticide supervision.

[0003] As a new optical detection method, hyperspectral imaging technology can provide three-dimensional data cubes containing spectral and spatial information of target samples by combining spectroscopy and imaging technology. In recent years, this technology has shown significant application value in the field of agricultural product quality and safety detection due to its non-destructive, non-polluting, no sample pretreatment and high detection efficiency. At present, this technology has made progress in pesticide residue detection of various agricultural products, fully verifying its feasibility and effectiveness in pesticide residue detection. However, the characteristics of tobacco are affected by various ecological environmental factors such as climate and light intensity, resulting in differences in spectral characteristics of tobacco grown in different years. This causes the pesticide residue discrimination model of tobacco based on hyperspectral technology to often face the problem of insufficient generalization ability, resulting in unstable operation effect under different environmental conditions and difficulty in ensuring the accuracy of detection. Traditional solutions usually require re-collection of a large amount of data and re-training of the model. However, this process requires a lot of time, effort and resources, making it extremely challenging to establish a pesticide residue discrimination model with wide applicability. Therefore, there is an urgent need to develop a method to enhance the generalization ability of the model to ensure its higher adaptability and robustness in practical applications. SUMMARY

[0004] The present application aims at the technical problem of inaccurate detection results caused by insufficient generalization ability of existing tobacco pesticide residue discrimination models, and provides a fresh tobacco cross-year pesticide residue concentration level detection method based on feature transfer learning, which can effectively overcome the influence of changes in tobacco characteristics caused by year differences on model performance, thereby realizing cross-year fresh tobacco pesticide residue concentration detection and providing a reliable transfer learning solution for cross-year fresh tobacco pesticide residue concentration detection.

[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions:

[0006] The present application provides a fresh tobacco cross-year pesticide residue concentration level detection method based on feature transfer learning, comprising the following steps:

[0007] S1: Obtain hyperspectral image data of fresh tobacco samples in the same production area for two consecutive years, take the earlier year data as the source domain data set, the newer year data as the target domain data set, and perform spectral correction and noise filtering preprocessing;

[0008] S2: Use a feature importance evaluation algorithm to reduce the dimensionality of the source domain full spectrum data, input the filtered features into a GA-LSSVM detection model, and train to obtain a GA-LSSVM pesticide residue concentration detection model;

[0009] S3: Align the feature distribution of the source domain and the target domain using a TCA and MEDA joint algorithm, input the adapted features into the GA-LSSVM pesticide residue concentration detection model, and generate a cross-year transfer learning optimization model;

[0010] S4: After the hyperspectral data of the fresh tobacco to be detected is preprocessed by spectral correction and noise filtering, input it into the cross-year transfer learning optimization model, and output the pesticide residue concentration value.

[0011] In the above technical solution of the present application, the band range of the fresh tobacco pesticide residue category hyperspectral data is 420-900nm.

[0012] In the above technical solution of the present application, the preprocessing adopts a first derivative (FD), second derivative (SD), mean variance transformation (MVT), or standard normal variate transformation (SNV).

[0013] In the technical scheme of the present application, the feature importance evaluation algorithm is a Competitive Adaptive Reweighted Sampling (CARS) algorithm or a Successive Projections Algorithm (SPA) algorithm.

[0014] The present application also provides a fresh tobacco leaf cross-year pesticide residue concentration level detection system based on feature transfer learning, comprising a data acquisition module, a spectrum visualization module, a data preprocessing module, and a pesticide residue concentration detection module.

[0015] The data acquisition module is configured to acquire hyperspectral data of a fresh tobacco leaf sample to be detected.

[0016] The spectrum visualization module is configured to calculate an average spectrum based on the hyperspectral data and display an average spectrum curve.

[0017] The data preprocessing module is configured to preprocess the hyperspectral data.

[0018] The pesticide residue concentration detection module is configured to determine a pesticide residue concentration value of the fresh tobacco leaf sample to be detected based on the cross-year transfer learning optimized model.

[0019] The present application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning.

[0020] The present application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning.

[0021] Compared with the prior art, the present application has the following advantages:

[0022] The constructed cross-year migration learning optimization model (TCA-MEDA-GA-LSSVM) effectively reduces the domain difference through feature space mapping, and MEDA further aligns the distributions of the source domain and the target domain in the manifold space, thereby achieving a more optimal migration effect. In five pesticide detection tasks, the optimal performance is achieved, and the accuracy rates are: acetamiprid 89.33%, fosamine 76%, thiophanate-methyl 80.67%, lambda-cyhalothrin 89.33%, and propamocarb 83.33%. The influence of the change of tobacco characteristics caused by year differences on the performance of the model can be effectively overcome, so as to realize the cross-year detection of fresh tobacco pesticide residue concentration and provide a reliable transfer learning solution for the cross-year detection of fresh tobacco pesticide residue concentration. BRIEF DESCRIPTION OF DRAWINGS

[0023] Figure 1 It is the average spectrum curve of fresh tobacco leaves of acetamiprid in 2022 under different residual concentrations, and low, normal and high in the figure represent concentrations of 0.04 mL / L, 0.17 mL / L and 0.34 mL / L, respectively.

[0024] Figure 2 It is the average spectrum curve of fresh tobacco leaves of fosamine in 2022 under different residual concentrations, and low, normal and high in the figure represent concentrations of 0.5 g / L, 2.5 g / L and 5.0 g / L, respectively.

[0025] Figure 3 It is the average spectrum curve of fresh tobacco leaves of thiophanate-methyl in 2022 under different residual concentrations, and low, normal and high in the figure represent concentrations of 0.17 g / L, 0.85 g / L and 1.70 g / L, respectively.

[0026] Figure 4 It is the average spectrum curve of fresh tobacco leaves of lambda-cyhalothrin in 2022 under different residual concentrations, and low, normal and high in the figure represent concentrations of 0.07 mL / L, 0.33 mL / L and 0.66 mL / L, respectively.

[0027] Figure 5 It is the average spectrum curve of fresh tobacco leaves of propamocarb in 2022 under different residual concentrations, and low, normal and high in the figure represent concentrations of 0.2 mL / L, 1.0 mL / L and 2.0 mL / L, respectively. DETAILED DESCRIPTION

[0028] The following examples are used to illustrate the present application, but are not used to limit the protection scope of the present application. If not specifically indicated, the technical means used in the examples is the conventional means known to those skilled in the art. The test methods in the following examples are conventional methods, unless otherwise specified.

[0029] Example 1

[0030] Fresh tobacco leaves were from Baoji tobacco leaf science and technology demonstration garden in Shaanxi province. The selection of pesticide standard samples for the experiment was based on the current situation of pesticide use in tobacco planting in this area, including five kinds of representative and commonly used pesticides: lambda-cyhalothrin, acetamiprid, propamocarb, metalaxyl-M and thiophanate-methyl. The experimental samples were selected from the middle part of the tobacco leaves without obvious defects, with good and basically consistent growth conditions. The precise quantity of pesticides was completed using a pipette gun and a high-precision electronic scale. After the pesticide configuration was completed, a special spraying device was used to uniformly spray the pesticides on the surface of the tobacco leaves. The samples were picked 24 hours after the spraying was completed and the hyperspectral data was collected. The data was collected once in 2022 and 2023, and the effectiveness of the model established by the data in 2022 was verified by the data in 2023. The number of samples collected in 2022 was 200 pieces of tobacco leaves for each type, and the number of samples collected in 2023 was 50 pieces of tobacco leaves for each type. A blank control was also set. The types of pesticides, concentrations of pesticides and sample numbers selected in the experiment are shown in Table 1.

[0031] Table 1 Information of pesticides used in the experiment and the matching concentration

[0032]

[0033] (1) Collecting hyperspectral data. The hyperspectral image of each sample was collected using a hyperspectral experimental instrument (Shenzhen Haipu Nanometer Optics Technology Co., Ltd.). To ensure the stability of the system and the accuracy of the data, the hyperspectral imaging system needs to be preheated for 10 minutes. Then, the system parameters are adjusted through the matching software VidiSpec to ensure that the image is clear and not distorted. The specific parameter settings are as follows: the power voltage is set to 9V, and the exposure time is adjusted to 30ms to ensure that the image is fully exposed and has no overexposure or underexposure phenomenon. After the parameter setting is completed, the light source and the dark box are turned off, and the black frame data is collected as a reference for background noise. Then, the light source is turned on and the dark box is turned off, and the standard white board with a reflectivity of 99% is used to collect white frame data to calibrate the reflectivity and spectral response of the system. After completing the collection of black frame and white frame, the tobacco leaf samples are laid flat in the dark box to ensure that the sample position is stable and there is no interference from external light sources. The hyperspectral image is collected. After the collection is completed, the black and white reference images are used to correct the original spectral image of the sample to obtain higher image quality.

[0034] The black and white correction formula is as follows:

[0035]

[0036] In the formula, I is the corrected pixel value (relative reflectivity); I0 is the original pixel value; I B is the collected black frame pixel value; I W is the collected white pixel value.

[0037] (2) Extract the region of interest and calculate the average spectrum. Using ENVI 5.6 software for region segmentation, the leaf vein and background are deducted, and the sample hyperspectral image containing only the leaf is extracted. The average spectrum of the hyperspectral image in the sample region of interest after processing is calculated using Matlab R2020a software. The hyperspectral image band range is 400-900 nm, with a step of 4 nm, a total of 126 bands. Because the 400-420 nm band has high noise, the band range is cut off, and the final band range is 420-900 nm, with a band number of 121.

[0038] Figure 1 The average spectrum curve of fresh tobacco leaves in 2022 under different residual concentrations of acetamiprid is shown in the figure, where low, normal, and high represent concentrations of 0.04 mL / L, 0.17 mL / L, and 0.34 mL / L, respectively. As can be seen from the figure, the difference between the three curves is most obvious in the 470-670 nm band range, with the lowest concentration curve having the highest reflectivity in this range, followed by the high concentration and normal concentration. In other band ranges, the three curves differ little and almost overlap, indicating that the spectral characteristic band of fresh tobacco leaves for acetamiprid pesticide residues should mainly be concentrated in the 470-670 nm band range.

[0039] Figure 2 The average spectrum curve of fresh tobacco leaves in 2022 under different residual concentrations of hymexazol is shown in the figure, where low, normal, and high represent concentrations of 0.5 g / L, 2.5 g / L, and 5.0 g / L, respectively. As can be seen from the figure, within the 470-900 nm band range, the normal residual level curve differs significantly from the high residual level curve and the low residual level curve. Specifically, in the 470-720 nm band range, the normal residual level curve is lower than the high residual level curve and the low residual level curve; while in the 720-900 nm band range, the normal residual level curve is higher than the two curves. The high residual level curve and the low residual level curve only show a slight difference in the 700-900 nm band range. Therefore, the spectral characteristic band of fresh tobacco leaves for hymexazol pesticide residues is mainly concentrated in the 700-900 nm band range.

[0040] Figure 3 The average spectrum curve of fresh tobacco leaves in 2022 under different residual concentrations of thiophanate-methyl is shown in the figure, where low, normal, and high represent concentrations of 0.17 g / L, 0.85 g / L, and 1.70 g / L, respectively. As can be seen from the figure, within the 470-700 nm band range, the normal residual level curve has the highest reflectivity, followed by the high residual level and the low residual level, and the difference between the three curves is relatively obvious. From the 700-900 nm band range, the reflectivity of the curves decreases in the order of low, high, and normal residual levels, and the difference between the three curves is more obvious.

[0041] Figure 4 The following are the average spectral curves of fresh tobacco leaves at different residue concentrations of lambda-cyhalothrin in 2022. Low, normal, and high represent concentrations of 0.07 mL / L, 0.33 mL / L, and 0.66 mL / L, respectively. As can be seen from the figure, within the 470-670 nm wavelength range, there are significant differences between normal residue levels and high and low residue levels, while the spectral curves for high and low residue levels almost overlap, with minimal differences. In contrast, within the 700-900 nm wavelength range, the spectral reflectance for high, medium, and low residue levels shows a clear sequential progression, from high to low, with significant differences between the three curves.

[0042] Figure 5 The following are the average spectral curves of fresh tobacco leaves at different propamocarb residue concentrations in 2022. Low, normal, and high represent concentrations of 0.2 mL / L, 1.0 mL / L, and 2.0 mL / L, respectively. As can be seen from the figure, within the 400-900 nm wavelength range, there are only slight differences between the three curves. The reflectance of the high, normal, and low level curves is arranged in descending order, indicating a relationship between propamocarb residue concentration and the spectral curve.

[0043] This example uses the 2022 samples as the source domain and divides the spectral data of the collected samples into five datasets: the acetamiprid residual concentration detection dataset (sample_d), the oxadiazine residual concentration detection dataset (sample_e), the methyl thiophanate residual concentration detection dataset (sample_j), the lambda-cyhalothrin residual concentration detection dataset (sample_l), and the propamocarb residual concentration detection dataset (sample_s). The 2023 samples are used as the target domain and divide the spectral data of the collected samples into five datasets: the acetamiprid residual concentration detection dataset (sample_d1), the lambda-cyhalothrin residual concentration detection dataset (sample_e1), the methyl thiophanate residual concentration detection dataset (sample_j1), the lambda-cyhalothrin residual concentration detection dataset (sample_l1), and the propamocarb residual concentration detection dataset (sample_s1).

[0044] Table 2 Number of original spectra in source and target domains

[0045]

[0046]

[0047] (2) Data preprocessing. In addition to the chemical information of the sample itself, the spectrum also contains other irrelevant information and noise, such as electrical noise, sample background and stray light, etc. These interferences will seriously affect the accuracy of hyperspectral data, and directly affect the correlation between the spectrum and the measured sample. In this embodiment, four kinds of spectral preprocessing are used for the constructed data set, which are first derivative (FD), second derivative (SD), mean variance (MVT) and standard normal variable transformation (SNV).

[0048] (4) Construction of GA-LSSVM model

[0049] In this embodiment, the genetic algorithm (Genetic Algorithm, GA) optimized least squares support vector machine (Least Squares Support Vector Machine, LSSVM) model, i.e. GA-LSSVM model, can improve the prediction performance and generalization ability of the model in the prediction task of the concentration level of fresh tobacco leaf pesticide residues.

[0050] For acetamiprid, the source domain data set sample_d processed by different preprocessing methods is input into the GA-LSSVM model, and the prediction results are shown in Table 3. It can be seen that when the preprocessing method is SNV, the accuracy of GA-LSSVM reaches 99.14%, the precision and recall are 99.39% and 98.85% respectively, and the F1 score is 0.9911, which shows the advantage of the model and the preprocessing method in the prediction task of acetamiprid residue concentration level.

[0051] For hymexazol, the source domain data set sample_e processed by different preprocessing methods is input into the GA-LSSVM model, and the prediction results are shown in Table 4. It can be seen that when the preprocessing method is FD, the accuracy, precision and recall of GA-LSSVM all reach 100%, and the F1 score also reaches 1.0000.

[0052] For thiophanate-methyl, the source domain data set sample_j processed by different preprocessing methods is input into the GA-LSSVM model, and the prediction results are shown in Table 5. It can be seen that when the preprocessing method is SNV, the accuracy, precision and recall of GA-LSSVM all reach 100%, and the F1 score also reaches 1.0000.

[0053] For lambda-cyhalothrin, the source domain data set sample_l processed by different preprocessing methods is input into the GA-LSSVM model, and the prediction results are shown in Table 6. It can be seen that when the preprocessing method is RAW, the accuracy, precision and recall of GA-LSSVM are all high, which can accurately predict the concentration level of lambda-cyhalothrin residue.

[0054] For propamocarb, the source domain dataset sample_s treated by different pretreatment methods was input into the GA-LSSVM model, and the prediction results are shown in Table 7. It can be seen that when the pretreatment method is SNV, the accuracy, precision and recall of GA-LSSVM are all high, which can accurately predict the concentration level of lambda-cyhalothrin residue.

[0055] Table 3 Prediction results of GA-LSSVM model on acetamiprid concentration level

[0056]

[0057] Table 4 Prediction results of GA-LSSVM model on propineb concentration level

[0058]

[0059] Table 5 Prediction results of GA-LSSVM model on thiophanate-methyl concentration level

[0060]

[0061]

[0062] Table 6 Prediction results of GA-LSSVM model on lambda-cyhalothrin concentration level

[0063]

[0064] Table 7 Prediction results of GA-LSSVM model on propamocarb concentration level

[0065]

[0066] (3) Data dimensionality reduction and training of GA-LSSVM pesticide residue concentration model

[0067] The spectral dimension of hyperspectral data is large, contains a lot of redundant information, especially there may be collinearity between adjacent bands, and high-dimensional data will also increase the complexity of the model and affect the performance of the model. Therefore, after spectral pretreatment, the data needs to be reduced, that is, to select some effective wavelengths from the full-band data as feature wavelengths. In this embodiment, competitive adaptive reweighted sampling algorithm (CARS) and successive projections algorithm (SPA) are used to reduce the dimension of the source domain full-spectrum data, eliminate redundant information, and compare the adaptability of GA-LSSVM model to different feature importance evaluation algorithms and different pretreatment methods.

[0068] The prediction results of GA-LSSVM pesticide residue concentration detection model on acetamiprid concentration level under different pretreatment conditions after SPA and CARS data dimension reduction are shown in Table 8. As can be seen from the table, when the pretreatment method is SNV and the feature selection method is CARS, the number of selected features is the most, but the model accuracy, precision, recall and F1 score are 98.28%, 98.81%, 97.70% and 0.9820 respectively, which is slightly lower than that of the full spectrum model.

[0069] The prediction results of GA-LSSVM pesticide residue concentration detection model on hymexazol concentration level under different pretreatment conditions after SPA and CARS data dimension reduction are shown in Table 9. As can be seen from the table, when the pretreatment method is FD and the feature selection method is CARS, the number of selected features is the most, and the performance is improved compared with the full spectrum model.

[0070] The prediction results of GA-LSSVM pesticide residue concentration detection model on hymexazol concentration level under different pretreatment conditions after SPA and CARS data dimension reduction are shown in Table 9. As can be seen from the table, when the pretreatment method is FD and the feature selection method is CARS, the number of selected features is the most, and the performance is improved compared with the full spectrum model.

[0071] The prediction results of GA-LSSVM pesticide residue concentration detection model on hymexazol concentration level under different pretreatment conditions after SPA and CARS data dimension reduction are shown in Table 9. As can be seen from the table, when the pretreatment method is FD and the feature selection method is CARS, the number of selected features is the most, and the performance is improved compared with the full spectrum model.

[0072] The prediction results of GA-LSSVM pesticide residue concentration detection model on hymexazol concentration level under different pretreatment conditions after SPA and CARS data dimension reduction are shown in Table 9. As can be seen from the table, when the pretreatment method is FD and the feature selection method is CARS, the number of selected features is the most, and the performance is improved compared with the full spectrum model.

[0073] Table 8 Prediction results of GA-LSSVM pesticide residue concentration detection model on acetamiprid concentration level

[0074]

[0075]

[0076] Table 9 Prediction results of GA-LSSVM pesticide residue concentration detection model on hymexazol concentration level

[0077]

[0078] Table 10 Prediction results of GA-LSSVM pesticide residue concentration detection model on thiophanate-methyl concentration level

[0079]

[0080] Table 11 Prediction results of GA-LSSVM pesticide residue concentration detection model on lambda-cyhalothrin concentration level

[0081]

[0082] Table 12 Prediction results of GA-LSSVM pesticide residue concentration detection model on probenazole concentration level

[0083]

[0084] (5) Constructing a cross-year transfer learning optimization model

[0085] Transfer learning (TL) is a machine learning method, and its core idea is that when there is a certain similarity between the source domain (D s ) and the target domain (D t ), the knowledge obtained by training the source domain can be transferred to the target domain, thereby reducing the large amount of labeling requirement for the target domain data, reducing the training cost, and improving the generalization ability of the model. Since the model fine-tuning can save training time and improve learning accuracy by realizing the transfer of deep network model, but this method also has the basic assumption that the training data or test data conforms to the same data distribution, which is not true in actual tasks, so the model fine-tuning cannot handle the case where the training data and test data have different distributions.

[0086] Therefore, in the embodiment, a feature-based transfer learning method is used to construct a cross-year fresh tobacco leaf pesticide residue concentration detection model, that is, a transfer component analysis (TCA)-manifold embedded distribution alignment (MEDA) joint strategy is used for transfer learning. First, the TCA algorithm is used to map the feature space of the source domain and the target domain data to eliminate the distribution difference between the domains. Then, the MEDA algorithm is used to align the feature distribution of the source domain and the target domain in the manifold space, thereby enhancing the domain invariance of the feature representation. Then, the adapted features are input into the GA-LSSVM pesticide residue concentration detection model to generate a cross-year transfer learning optimization model.

[0087] To verify the effectiveness of the method, five migration task experiments were conducted, the migration task setting information is shown in Table 13, and the prediction results are compared with the prediction results of the GA-LSSVM direct prediction and the MEDA migration learning method alone, as shown in Table 14. It can be seen that the TCA-MEDA-GA-LSSVM model (i.e. the cross-year migration learning optimization model) performs well in multiple migration tasks, especially in task 1, task 4 and task 5, with accuracy rates of 89.33%, 89.33% and 83.33% respectively, and precision rates of 89.32%, 91.49% and 85.86% respectively. However, in task 2, the model performance is lower, with accuracy and precision rates of 76.00% and 79.47% respectively, which may be due to the large domain difference in task 2, increasing the difficulty of migration.

[0088] Table 13 Migration task setting information

[0089] source domain number of samples target domain number of samples migration task number sample_d 600 sample_d1 150 1 sample_e 600 sample_e1 150 2 sample_j 600 sample_j1 150 3 sample_l 600 sample_l1 150 4 sample_s 600 sample_s1 150 5

[0090] Table 14 Prediction results of different migration learning methods in five migration tasks

[0091]

[0092]

[0093] As can be seen from Table 14, compared with the direct migration strategy, the model based on MEDA and TCA-MEDA joint migration strategy shows significant performance improvement in all migration tasks, with an average increase of more than 20% in accuracy, fully verifying the effectiveness of the migration learning algorithm in the cross-year fresh tobacco leaf pesticide residue concentration level detection task.

[0094] Embodiment two

[0095] A fresh tobacco leaf cross-year pesticide residue concentration level detection system based on feature migration learning, comprising: a data acquisition module, a spectrum visualization module, a data preprocessing module, and a pesticide residue concentration detection module.

[0096] The data acquisition module is used to acquire hyperspectral data of the fresh tobacco leaf sample to be detected;

[0097] The spectrum visualization module is used to calculate and display the average spectrum curve according to the hyperspectral data;

[0098] The data preprocessing module is used to preprocess the hyperspectral data, such as spectrum correction and noise filtering preprocessing;

[0099] The pesticide residue concentration detection module is used to determine the pesticide residue concentration value of the fresh tobacco leaf sample to be detected according to the cross-year migration learning optimization model.

[0100] The band range of the fresh tobacco leaf pesticide residue category hyperspectral data collected by the data acquisition module in the embodiment is 420-900nm.

[0101] The spectral visualization module in the embodiment can allow a user to dynamically view the hyperspectral image under different bands by sliding the band selection slider, support zooming of the image and multi-curve comparison analysis, and help the user intuitively understand the image difference of the sample under different bands.

[0102] The data preprocessing module in the embodiment supports preprocessing of the hyperspectral data by using a first derivative, a second derivative, mean variance, or standard normal variable transformation method.

[0103] The pesticide residue concentration detection module in the embodiment is embedded with a cross-year transfer learning optimization model, which can determine the pesticide residue concentration of the fresh tobacco leaf sample to be detected by taking the preprocessed fresh tobacco leaf hyperspectral data as input, and the detection result is displayed in real time on the "detection result display" page to provide intuitive feedback for the user.

[0104] Embodiment three

[0105] An electronic device includes a memory, a processor, a computer program stored on the memory and executable on the processor, a communication interface, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can invoke the logical instructions in the memory to implement the fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning when executed. In addition, the logical instructions in the memory can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer storable medium.

[0106] Based on such understanding, the technical solutions of the present application, in essence or the part that contributes to the prior art, or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes instructions to make a computer device execute all or part of the steps of the method described in various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), and various program code storage media.

[0107] A fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning includes the following steps:

[0108] S1: Obtain hyperspectral image data of fresh tobacco leaf samples in the same production area for two consecutive years, take the data of the earlier year as the source domain data set, the data of the newer year as the target domain data set, and perform spectral correction and noise filtering preprocessing;

[0109] S2: Dimensionality reduction is performed on the source domain full-spectrum data by using a feature importance evaluation algorithm, the screened features are input into a GA-LSSVM detection model, and a GA-LSSVM pesticide residue concentration detection model is trained and obtained;

[0110] S3: The feature distribution of the source domain and the target domain is aligned by using a TCA and MEDA joint algorithm, the adapted features are input into the GA-LSSVM pesticide residue concentration detection model, and a cross-year transfer learning optimization model is generated;

[0111] S4: After the hyperspectral data of the fresh tobacco leaf to be detected are preprocessed by spectral correction and noise filtering, the cross-year transfer learning optimization model is input, and the pesticide residue concentration value is output.

[0112] The above-described embodiments are only preferred embodiments of the present application, and are used to explain the present application, but do not limit the scope of the present application. For those skilled in the art, of course, other embodiments can be easily obtained by substitution or change based on the technical content disclosed in the present specification, and therefore, any changes and improvements made on the principle of the present application shall be included in the scope of the patent application of the present application.

Claims

1. A method for detecting the concentration level of fresh tobacco leaf pesticide residues across years based on feature transfer learning, characterized in that, The method comprises the following steps: S1: Obtain hyperspectral image data of fresh tobacco leaf samples in the same production area for two consecutive years, take the data of the earlier year as the source domain data set, take the data of the later year as the target domain data set, and perform spectral correction and noise filtering preprocessing; S2: Use a feature importance evaluation algorithm to reduce the dimension of the source domain full-spectrum data, input the screened features into a GA-LSSVM detection model, and train to obtain a GA-LSSVM pesticide residue concentration detection model; S3: Use a TCA and MEDA joint algorithm to align the feature distributions of the source domain and the target domain, input the adapted features into the GA-LSSVM pesticide residue concentration detection model, and generate a cross-year transfer learning optimization model; S4: After the hyperspectral data of the fresh tobacco leaf to be detected are preprocessed by spectral correction and noise filtering, input the data into the cross-year transfer learning optimization model, and output the pesticide residue concentration value.

2. The method for detecting the concentration of fresh tobacco leaf annual agricultural residues according to claim 1, characterized in that, The band range of the fresh tobacco leaf pesticide residue category hyperspectral data is 420-900 nm.

3. The method for detecting the concentration of fresh tobacco leaf annual agricultural residues according to claim 1, characterized in that, The preprocessing uses a median, second derivative, mean variance, or standard normal variable transformation.

4. The method for detecting the concentration of fresh tobacco leaf annual agricultural residues according to claim 1, characterized in that, The feature importance evaluation algorithm is a competitive adaptive reweighted sampling algorithm or a successive projections algorithm.

5. A system for detecting the concentration level of pesticide residues in fresh tobacco leaves across years based on feature transfer learning according to any one of claims 1-4, characterized in that, It comprises: a data acquisition module, a spectral visualization module, a data preprocessing module, and a pesticide residue concentration detection module; The data acquisition module is used to acquire hyperspectral data of fresh tobacco leaf samples to be detected; The spectral visualization module is used to calculate average spectra from the hyperspectral data and display the average spectrum curve; The data preprocessing module is used to preprocess the hyperspectral data; The pesticide residue concentration detection module is used to determine the pesticide residue concentration value of the fresh tobacco leaf sample to be detected according to the cross-year transfer learning optimization model.

6. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed to realize the fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning according to any one of claims 1-4.

7. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The computer program is executed by the processor to realize the fresh tobacco leaf cross-year pesticide residue concentration level detection method based on feature transfer learning according to any one of claims 1-4.