Tobacco virus disease remote sensing distinguishing method and system based on multi-source data

By combining leaf area index, chlorophyll content and spectral data, and using continuous wavelet transformation and random forest algorithm, a tobacco virus lesions distinction model was constructed, solving the problems of strong subjectivity of identification results and incomplete feature information in the existing technology, and achieving efficient disease distinction and monitoring.

CN120145221APending Publication Date: 2025-06-13HONGHEZHOU BRANCH OF YUNNAN TOBACCO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510140367.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When identifying the health status and disease types of tobacco plants, the prior art has problems such as strong subjectivity of identification results, incomplete feature information, and difficulty in efficiently distinguishing diseases in complex field environments.

Method used

By comprehensively utilizing the leaf area index, chlorophyll content and canopy spectrum data of tobacco plants, remote sensing features are extracted based on continuous wavelet transformation, and combined with random forest algorithm, a distinction model between healthy plants and mosaic and vermilion disease infecting plants was constructed.

Benefits of technology

Innovation in the construction of remote sensing feature of disease distinction has been achieved, significantly improving identification accuracy, reducing costs, and providing an efficient and economical solution for tobacco virus disease monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145221A_ABST
    Figure CN120145221A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of tobacco virus disease distinguishing, and discloses a tobacco virus disease remote sensing distinguishing method and system based on multi-source data, and the method comprises the following steps: carrying out chlorophyll content and leaf area index determination on a tobacco sample to obtain chlorophyll characteristics and leaf area characteristics; fusing the spectral wavelet features, the chlorophyll features and the leaf area features, and preprocessing the fused data to form a data set; calling a historical tobacco classification database, and constructing a tobacco virus disease distinguishing model; inputting the collected spectral wavelet features, chlorophyll features and leaf area features into a tobacco virus disease classification model for classification and identification; and evaluating data classified by the tobacco virus disease classification model, and adjusting the tobacco virus disease classification model according to an evaluation result. The system comprises a tobacco sample collecting module, a distinguishing model constructing module and a distinguishing evaluating module. According to the invention, accurate distinguishing and identification of tobacco virus diseases in a complex field environment are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tobacco virus disease differentiation, and particularly to a remote sensing differentiation method and system for tobacco virus diseases based on multi-source data. Background Art

[0002] Currently, when traditional methods identify the health status and disease types of tobacco plants, they rely on manual observation or single feature extraction, and there are limitations such as strong subjectivity of the identification results, incomplete feature information, and difficulty in efficiently differentiating diseases in complex field environments. The present invention comprehensively utilizes the leaf area index, chlorophyll content, and canopy spectral data of tobacco plants, extracts remote sensing features based on continuous wavelet transform, and combines with the random forest algorithm to construct a differentiation model for healthy plants and plants infected with mosaic disease and curly leaf disease. This method not only innovates in the construction of remote sensing features for disease differentiation, but also significantly improves the identification accuracy and reduces the cost, providing an efficient and economical solution for tobacco virus disease monitoring.

[0003] Tobacco virus diseases are important factors affecting the yield and quality of tobacco cultivation. Among them, mosaic disease and curly leaf disease are the two most common and severely harmful diseases. Accurately identifying and differentiating the disease types of tobacco plants is crucial for formulating effective prevention and control measures. However, traditional disease identification methods mainly rely on manual observation, which is not only inefficient and subjective, but also difficult to meet the requirements of modern agriculture for precision and high efficiency. With the rapid development of remote sensing technology, crop disease identification based on remote sensing data has become a research hotspot.

[0004] After a crop is infected by a pathogen, a series of physiological and biochemical characteristics begin to change, and then different spectral responses are generated. Therefore, the disease can be identified by analyzing the spectral reflectance of the crop. Common methods for detecting crop diseases based on hyperspectral spectral features include vegetation indices, selection of optimal bands of original spectra, continuous wavelet transform, etc. Wavelet transform is an emerging spectral analysis method. By decomposing spectral data at multiple scales, it is possible to capture fine spectral change information. For example, Shi et al. proposed a wavelet-based stripe rust spectral feature set (WFs) to reveal the related processes of wheat stripe rust occurrence [1]. Zhang et al. combined wavelet transform with partial least squares regression (PLSR) based on the hyperspectral information of diseased leaves to achieve the assessment of powdery mildew of winter wheat at the leaf level [2]. However, there is currently no research on identifying tobacco virus diseases using the wavelet transform method.

[0005] Compared with spectral features, the physical and chemical parameter features of crops can more directly reflect the physiological and chemical changes of crops after being infected by diseases, and have received increasing attention in crop disease detection. At present, some scholars have comprehensively used spectra and physical and chemical parameters to establish disease recognition and discrimination methods. For example, Wu et al. extracted multiple physical and chemical parameters such as chlorophyll content and LAI from the UAV hyperspectral images of jujube fruits during the swelling period, and established a health assessment model of jujube trees based on the physical and chemical parameter features [3]. Liu et al. achieved the severity assessment of apple mosaic disease at the leaf scale through the comprehensive analysis of hyperspectral data and chlorophyll content [4]. Cheng et al. combined spectral features such as vegetation indices and wavelet features with physical and chemical parameters such as chlorophyll and anthocyanin to achieve the early identification of powdery mildew of rubber trees [5]. However, the current research on the identification of tobacco diseases and pests using the combination of spectral and physical and chemical parameter features has not received enough attention and needs further in-depth study. Therefore, in this invention, taking the hyperspectral reflectance data and physical and chemical parameter data as data sources, and optimizing the determination of physical and chemical parameters according to the morphological characteristics of tobacco plants, an effective method for identifying tobacco virus diseases by combining spectral and physical and chemical parameter features is proposed.

[0006] In current research, many disease recognition technologies only use spectral reflectance values or vegetation indices (such as NDVI and EVI) to characterize disease features, while ignoring the roles of physiological and biochemical parameters such as leaf area index and chlorophyll content. For example, Wang et al. extracted five bands sensitive to tobacco virus disease at 631, 638, 696, 733, and 864 nm from the hyperspectral reflectance data of tobacco measured by an ASD spectrometer, and achieved the identification of tobacco virus disease [6, 7]. Yusuf and He identified the disease severity of tobacco black shank based on the plant senescence reflectance index (PSRI) extracted from the laboratory hyperspectral reflectance data [8]. However, single spectral features are difficult to accurately capture the multi-dimensional information of diseased plants, especially in the case of coexistence of multiple disease types, and their discrimination ability is greatly limited.

[0007] Currently, the existing technologies have problems such as ignoring the roles of physiological and biochemical parameters such as leaf area index and chlorophyll content, and single spectral features being difficult to accurately capture the multi-dimensional information of diseased plants. Especially in the case of coexistence of multiple disease types, their discrimination ability is greatly limited. To solve the above problems, this invention provides a remote sensing discrimination method and system for tobacco virus diseases based on multi-source data.

[0008] [1] Shi, Y.; Huang, W.; González-Moreno, P.; Luke, B.; Dong, Y.; Zheng, Q.; Ma, H.; Liu, L. Wavelet-Based Rust Spectral Feature Set (WRSFs): A Novel Spectral Feature Set Based on Continuous Wavelet Transformation for Tracking Progressive Host-Pathogen Interaction of Yellow Rust on Wheat. Remote Sens. 2018, 10, 525, doi:10.3390 / rs10040525.

[0009] [2] Zhang, J.; Lin, Y.; Wang, J.; Huang, W.; Chen, L.; Zhang, D. Spectroscopic Leaf Level Detection of Powdery Mildew for Winter Wheat Using Continuous Wavelet Analysis. J. Integr. Agric. 2012, 11, 1474 - 1484, doi:10.1016 / S2095-3119(12)60147-6.

[0010] [3] Wu, Y.; Zhao, Q.; Yin, X.; Wang, Y.; Tian, W. Multi-Parameter Health Assessment of Jujube Trees Based on Unmanned Aerial Vehicle Hyperspectral Remote Sensing. Agriculture 2023, 13, 1679, doi:10.3390 / agriculture13091679.

[0011] [4] Liu, Y.; Zhang, Y.; Jiang, D.; Zhang, Z.; Chang, Q. Quantitative Assessment of Apple Mosaic Disease Severity Based on Hyperspectral Images and Chlorophyll Content. Remote Sens. 2023, 15, 2202, doi:10.3390 / rs15082202.

[0012] [5] Cheng, X.; Huang, M.; Guo, A.; Huang, W.; Cai, Z.; Dong, Y.; Guo, J.; Hao, Z.; Huang, Y.; Ren, K.; et al. Early Detection of Rubber Tree Powdery Mildew by Combining Spectral and Physicochemical Parameter Features. Remote Sens. 2024, 16, 1634, doi:10.3390 / rs16091634.

[0013] [6] Wang, M.; Li, X.; Yao, Q.; Liu, Y. Extraction of Diseases and Insect Pests for Tobacco Based on Hyperspectral Remote Sensing. Geod List 2012.

[0014] [7] Wang, M.; Li, X.J.; Lu, Y.Y.; Guo, S.L. Tobacco Pest Monitoring Feasibility Analysis Based on RS. Adv. Mater. Res. 2011, 217 - 218, 1516 - 1519, doi:10.4028 / www.scientific.net / AMR.217 - 218.1516.

[0015] [8]Babangida Lawal Yusuf Application of Hyperspectral Imaging Sensorto Differentiate between the Moisture and Reflectance of Healthy and InfectedTobacco Leaves.Afr.J.Agric.RESEEARCH 2011,6,doi:10.5897 / AJAR11.1281. Summary of the Invention

[0016] The main object of the present invention is to provide a remote sensing discrimination method and system for tobacco virus diseases based on multi-source data, so as to solve the problems in the prior art that the roles of physiological and biochemical parameters such as leaf area index and chlorophyll content are ignored, and it is difficult for single spectral features to accurately capture the multi-dimensional information of diseased plants. Especially in the case of coexistence of multiple disease types, its discrimination ability is greatly limited.

[0017] To achieve the above object, the present invention provides the following technical solutions:

[0018] A remote sensing discrimination method for tobacco virus diseases based on multi-source data, the remote sensing discrimination method for tobacco virus diseases based on multi-source data includes:

[0019] Collect spectral data of tobacco samples, and use continuous wavelet transform to extract the spectral data characteristics of tobacco samples; classify the grades of tobacco virus diseases, measure the chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll characteristics and leaf area characteristics;

[0020] Fuse the spectral wavelet characteristics, chlorophyll characteristics and leaf area characteristics, preprocess the fused data to form a data set; divide the data set into a training set, a validation set and a test set; retrieve the historical tobacco classification database and construct a discrimination model for tobacco virus diseases;

[0021] Input the collected optical wavelet characteristics, chlorophyll characteristics and leaf area characteristics into the discrimination model for tobacco virus diseases for classification and discrimination; evaluate the data classified by the discrimination model for tobacco virus diseases, and adjust the discrimination model for tobacco virus diseases according to the evaluation results.

[0022] As a further improvement of the present invention, the process of using continuous wavelet transform to extract the spectral data characteristics of tobacco samples includes the following steps:

[0023] Collect multispectral data in the wavelength range of 350 to 2500 for tobacco samples, and use continuous wavelet transform to extract the spectral data characteristics of tobacco; analyze the original spectral curve and Gaussian function at different positions and scales to generate continuous wavelet power coefficients; extract weak information in different ice-white spectra;

[0024] Select the Mexican hat wavelet, which has absorption characteristics similar to vegetation indices, as the mother wavelet basis function, and only retain the wavelet power at decomposition scales of 2 n (n = 1, 2,..., 10), which are denoted as the 1st scale, 2nd scale,..., 10th scale respectively;

[0025] Based on the continuous wavelet transform, perform grade classification on tobacco virus diseases using wavelet features extracted based on the top 1% threshold of the determination coefficient between wavelet energy coefficients and tobacco virus disease grades. Measure the chlorophyll content and leaf area index of tobacco samples to obtain chlorophyll characteristics and leaf area characteristics.

[0026] As a further improvement of the present invention, the process of obtaining chlorophyll characteristics and leaf area characteristics includes the following steps:

[0027] Measure the chlorophyll content of tobacco samples. For each plant, collect two representative leaves, cut each leaf into upper, middle, and lower parts, measure two values in the middle of each part, and average the six values as the chlorophyll content. Take the average value as the chlorophyll characteristic. For the leaf area index of each sample, perform 3 sampling values, and take its average value and variance as the leaf area characteristic.

[0028] As a further improvement of the present invention, the process of constructing a tobacco virus disease discrimination model includes the following steps:

[0029] Fuse the spectral wavelet characteristics, chlorophyll characteristics, and leaf area characteristics, and update the ecological niche for network parameter allocation of the fused data. The ecological niche is determined by the competition strategy in the tobacco simulation ecosystem;

[0030] Perform preprocessing operations such as missing value filling and standardization on the fused data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database to construct a tobacco virus disease discrimination model;

[0031] Train the tobacco virus disease discrimination model with the training set. Among them, the network parameters of the tobacco virus disease discrimination model are allocated with an initial ecological niche; calculate the loss function corresponding to the network parameters of the tobacco virus disease discrimination model, and update the network parameters.

[0032] As a further improvement of the present invention, the process of updating the ecological niche for network parameter allocation of the fused data includes the following steps:

[0033] Input the fused tobacco training set into a preset autoencoder neural network. Through the encoder of the autoencoder neural network, map the tobacco training set to the initial low-dimensional feature space to obtain the initial low-dimensional features;

[0034] Calculate the initial low-dimensional features to obtain the corresponding loss function, and determine the feature importance of the low-dimensional features; Based on the feature importance, recursively optimize the feature weight vector of the autoencoder neural network; Based on the feature weight vector, determine the new low-dimensional features;

[0035] Perform feature reconstruction on the new low-dimensional features through the compiler until the autoencoder neural network meets the preset iteration conditions;

[0036] Among them, the low-dimensional feature mapping formula of the autoencoder neural network:

[0037] Z = f encoder (X; W e , b e ) = σ(W e ·X + b e )

[0038] In the formula, X represents the input tobacco training set data matrix, with a dimension of (n×d), where n is the number of samples and d is the feature dimension; W e represents the weight matrix of the encoder, with a dimension of (k×d), where k is the dimension of the low-dimensional feature space; b e represents the bias vector of the encoder, with a dimension of (k×l); σ(·) represents the activation function; Z represents the mapped low-dimensional feature matrix, with a dimension of (n×k);

[0039] Low-dimensional feature loss function calculation and feature importance evaluation formula:

[0040]

[0041] In the formula, X i represents the original high-dimensional feature vector of the i-th sample; Z i represents the low-dimensional feature vector of the i-th sample; f decoder (·) represents the decoder function, which is used to reconstruct the low-dimensional features into high-dimensional features; W d represents the weight matrix of the decoder, with a dimension of (d×k); b d represents the bias vector of the decoder, with a dimension of (d×l); λ represents the regularization coefficient, which is used to control the sparsity of the feature weight vector; w j represents the j-th feature weight;

[0042] Low-dimensional feature reconstruction and iterative optimization formula:

[0043]

[0044] In the formula, Z (t) represents the low-dimensional feature matrix at the t-th iteration; η represents the learning rate, which controls the iteration step size; represents the gradient of the loss function with respect to the low-dimensional feature; the low-dimensional feature is iteratively optimized by the gradient descent method until the preset iteration condition is satisfied.

[0045] As a further improvement of the present invention, the process of calculating the loss function corresponding to the network parameters of the tobacco virus disease discrimination model includes the following steps:

[0046] Calculate the loss function corresponding to the network parameters of the tobacco virus disease discrimination model, use the loss function as the adaptability score corresponding to the network parameters, and calculate the energy adjustment factor corresponding to the grid parameters; input the energy adjustment factor and the preset maximum energy to calculate the mutation energy corresponding to the grid parameters;

[0047] Based on the mutation energy and the preset migration and mutation adjustment coefficient, calculate the adaptability score to determine the dynamic mutation rate corresponding to the network parameters;

[0048] Migrate the initial niche based on the dynamic mutation rate, and update the grid parameters; input the tobacco sample set, sample the samples to obtain a sampling set containing m samples; train the t-th tobacco virus disease discrimination model with the sampling set D′; the category with the most votes cast by T tobacco virus disease discrimination models is the final category.

[0049] As a further improvement of the present invention, the process in which the category with the most votes cast by the decision tree model is the final category includes the following steps:

[0050] Initialize the structure of the tobacco virus disease discrimination model. At each decision node, at a certain decision node, select a feature to perform splitting by maximizing the local structure entropy gain;

[0051] Input the tobacco sample set, sample the tobacco samples to obtain a sampling set including m samples; perform node splitting according to the selected feature and its threshold, and allocate the tobacco sample data to the left and right child nodes. Split the tobacco sample data according to the selected feature and its optimal splitting point;

[0052] Perform leaf node calibration. At the leaf nodes of the decision tree, calibrate the node category according to the majority cleavage of the tobacco samples in the node; the category with the most votes cast by T tobacco virus disease discrimination models is the final category; after training, evaluate the classification accuracy of the tobacco virus disease discrimination model to obtain the classification effect.

[0053] As a further improvement of the present invention, the process of evaluating the classification accuracy of the tobacco virus disease discrimination model includes the following steps:

[0054] Construct a confusion matrix to analyze the performance of the test set in the tobacco virus disease discrimination model after training. Each cell in the confusion matrix contains the number of samples in the corresponding category; calculate the accuracy of the classification results for each category;

[0055] Calculate the classification efficiency of the classification results for each category according to the confusion matrix. According to the classification efficiency of the classification results for each category and the probability of the category appearing in the total samples, calculate the accuracy of the classification results for each category;

[0056] Determine the accuracy of the tobacco virus disease discrimination model according to the calculated accuracy of the classification results for each category; and calibrate it with the preset accuracy. If it is less than the preset accuracy, adjust the tobacco virus disease discrimination model.

[0057] As a further improvement of the present invention, the process of adjusting the tobacco virus disease classification model according to the evaluation results includes the following steps:

[0058] Input the collected optical wavelet features, chlorophyll features, and leaf area features into the tobacco virus disease classification model for classification and discrimination;

[0059] Evaluate the classified tobacco data, compare the evaluation results with the preset evaluation result threshold. If it is greater than the preset evaluation result threshold, increase the tobacco sample data and adjust the tobacco classification model;

[0060] Upload the final classification results to the tobacco pathology evaluation platform, obtain the tobacco planting management plan according to the classification results, and monitor the development of the disease in real time. If it is greater than the preset disease threshold, give an alarm.

[0061] To achieve the above object, the present invention also provides the following technical solution:

[0062] A remote sensing discrimination system for tobacco virus diseases based on multi-source data, which is applied to the remote sensing discrimination method for tobacco virus diseases based on multi-source data. The remote sensing discrimination system for tobacco virus diseases based on multi-source data includes:

[0063] A tobacco sample collection module, which is used to collect spectral data of tobacco samples, extract spectral data features of tobacco samples by using continuous wavelet transform; classify the grades of tobacco virus diseases, measure the chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll features and leaf area features;

[0064] A discrimination model construction module, which is used to fuse spectral wavelet features, chlorophyll features, and leaf area features, preprocess the fused data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database and construct a tobacco virus disease discrimination model;

[0065] An evaluation and discrimination module is used to input the collected optical wavelet features, chlorophyll features, and leaf area features into the tobacco virus disease classification model for classification and discrimination; evaluate the data classified by the tobacco virus disease classification model, and adjust the tobacco virus disease classification model according to the evaluation results.

[0066] The present invention proposes a method for measuring and extracting features of tobacco plants. By combining the morphological characteristics of tobacco plants, the measurement methods of leaf area index and chlorophyll content are optimized, so as to more accurately reflect the physiological and biochemical changes of diseased plants, providing a scientific basis for comprehensively analyzing the growth status and disease characteristics of tobacco plants; innovatively combining spectral features, physical and chemical parameters with the random forest algorithm, a tobacco virus disease discrimination model with high accuracy and robustness is constructed. This method makes full use of the sensitivity of spectral data and the comprehensiveness of physical and chemical parameters, combines the superior classification ability of the random forest, realizes the accurate discrimination and identification of tobacco virus diseases in complex field environments, and significantly improves the applicability and popularization value of the model. Description of the Drawings

[0067] Figure 1 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention;

[0068] Figure 2 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention to extract spectral data features of tobacco samples using continuous wavelet transform;

[0069] Figure 3 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention to obtain chlorophyll features and leaf area features;

[0070] Figure 4 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention to construct a tobacco virus disease discrimination model;

[0071] Figure 5 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention to update the niche for network parameter allocation of the fused data;

[0072] Figure 6 It is a schematic diagram of the step flow of an embodiment of the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention to calculate the loss function corresponding to the network parameters of the tobacco virus disease discrimination model;

[0073] Figure 7 It is a schematic diagram of the principle for realizing the first classification function in the remote sensing discrimination method for tobacco virus diseases based on multi-source data of the present invention;

[0074] Figure 8 This is the schematic diagram for realizing the second classification function in the remote sensing discrimination method of tobacco virus disease based on multi-source data of the present invention;

[0075] Figure 9 This is the schematic diagram of the step process where the category with the most votes cast by the decision tree model is the final category in an embodiment of the remote sensing discrimination method of tobacco virus disease based on multi-source data of the present invention;

[0076] Figure 10 This is the schematic diagram of the step process for evaluating the classification accuracy of the tobacco virus disease discrimination model in an embodiment of the remote sensing discrimination method of tobacco virus disease based on multi-source data of the present invention;

[0077] Figure 11 This is the schematic diagram of the step process for adjusting the tobacco virus disease classification model according to the evaluation results in an embodiment of the remote sensing discrimination method of tobacco virus disease based on multi-source data of the present invention;

[0078] Figure 12 This is the schematic diagram of the functional modules in an embodiment of the remote sensing discrimination system of tobacco virus disease based on multi-source data of the present invention;

[0079] Figure 13 This is the schematic diagram of the structure in an embodiment of the electronic device of the present invention;

[0080] Figure 14 This is the schematic diagram of the structure in an embodiment of the storage medium of the present invention. Detailed implementation manners

[0081] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0082] The terms "first", "second", and "third" in the present invention are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", and "third" may explicitly or implicitly include at least one of such features. In the description of the present invention, the meaning of "a plurality" is at least two, such as two, three, etc., unless otherwise specifically defined. All directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will change accordingly. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products, or devices.

[0083] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with the embodiments can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0084] As Figure 1 shown, this embodiment provides an embodiment of a remote sensing discrimination method for tobacco virus diseases based on multi-source data. In this embodiment, the remote sensing discrimination method for tobacco virus diseases based on multi-source data specifically includes the following steps:

[0085] Step S1: Collect spectral data of tobacco samples in the wavelength range from 350 to 2500, and extract spectral data features of tobacco samples using continuous wavelet transform; classify the grades of tobacco virus diseases, measure the chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll features and leaf area features;

[0086] Step S2: Integrate the spectral wavelet features, chlorophyll features, and leaf area features, preprocess the integrated data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database and construct a discrimination model for tobacco virus diseases;

[0087] Step S3: Input the collected optical wavelet features, chlorophyll features, and leaf area features into the tobacco virus disease classification model for classification and discrimination; evaluate the data classified by the tobacco virus disease classification model, and adjust the tobacco virus disease classification model according to the evaluation results.

[0088] Preferably, in this embodiment, based on the hyperspectral reflectance and physicochemical parameter data of tobacco, precise differentiation and identification of the main tobacco diseases and pests are carried out. The steps are divided into tobacco feature extraction and the construction of a disease and pest differentiation and identification model. Among them, tobacco feature extraction is further divided into three parts: spectral wavelet feature extraction, chlorophyll feature extraction, and leaf area index feature extraction; First, for the multispectral data of tobacco samples collected in the wavelength range of 350-2500, the spectral data features of tobacco samples are extracted by using continuous wavelet transform. Continuous wavelet transform is an important signal processing method that can localize both the frequency domain and the time domain at the same time, and refine functions or signals at different scales and positions. Based on continuous wavelet transform, correlation analysis is carried out on the original spectral curve and the Gaussian function at different positions and scales to generate a series of continuous wavelet energy coefficients, and these energy coefficients can extract weak information in the spectra of different diseases. In this embodiment, the Mexican hat wavelet, which is similar to the absorption characteristics of vegetation indices, is selected as the mother wavelet basis function. For the convenience of calculation and without affecting the accuracy of the continuous wavelet transform method, only the decomposition scale of 2 is retained n(n = 1, 2, …, 10) wavelet power is respectively denoted as the 1st scale, the 2nd scale, …, the 10th scale. Based on continuous wavelet transform, wavelet features are extracted based on the determination coefficient between wavelet energy coefficients and tobacco virus disease grades with the top 1% threshold. The threshold extraction method is to sort the determination coefficients at all scales from large to small and use the method of threshold limitation for screening. The wavelet energy coefficients calculated above are subjected to correlation analysis with the tobacco virus disease grades to generate a determination coefficient (R2). The R2 values of wavelet energy coefficients at different bands and different scales form a correlation scale map, which characterizes the sensitivity of each wavelet energy coefficient to tobacco virus disease. Elements with R2 in the top 1% are retained as wavelet feature regions. To reduce redundancy, within each wavelet feature region, the wavelet energy coefficient with the highest R2 is retained as the extracted wavelet feature. Then, chlorophyll content determination is carried out on the samples. The method is to collect two representative unfolded leaves from each plant, cut each leaf into three parts: upper, middle, and lower, measure two values in the middle of each part, and the average of the six values is the chlorophyll content of the leaf, and the mean value is taken as the chlorophyll feature. Finally, the leaf area index of each sample is sampled three times, and its mean value and variance are taken as the leaf area feature. By fusing spectral wavelet features, chlorophyll features, and leaf area features, the random forest method is used to train and classify the samples, and the accuracy rate reaches 88.9%. In view of the morphological characteristics of tobacco plants, the present invention designs a unique measurement method during the determination of physical and chemical parameters. By reasonably selecting key parameters such as leaf area index and chlorophyll content, it more accurately characterizes the physiological and biochemical changes of tobacco virus disease plants. This optimization not only improves the representativeness of physical and chemical parameter data but also provides more reliable feature inputs for subsequent modeling. The present invention innovatively combines spectral features with physical and chemical parameters to construct a remote sensing discrimination model for tobacco virus disease based on multi-source data. The model makes full use of the response ability of spectral data to disease-sensitive bands and the multi-dimensional characterization ability of physical and chemical parameters to the health status of tobacco, and solves the problem of singularity in feature extraction and model construction in traditional technologies. Through the fusion of multi-source features, the present invention significantly improves the accuracy and reliability of tobacco virus disease discrimination, reduces the interference of the complex field environment on the recognition effect, and provides a more universal and economical solution for the precise monitoring of tobacco virus disease.

[0089] Further, as Figure 2 shown, the process of step S1 for extracting spectral data features of tobacco samples using continuous wavelet transform specifically includes the following steps:

[0090] Step S11: Collect multispectral data of tobacco samples in the wavelength range of 350 to 2500, and use continuous wavelet transform to extract spectral data features of tobacco; analyze the original spectral curve and Gaussian function at different positions and scales to generate continuous wavelet energy coefficients; extract weak information in different ice-white spectra;

[0091] Among them, the output of the continuous wavelet transform is as follows:

[0092]

[0093] In the formula, f(λ) is the original spectrum, λ = 1, 2,..., m, where m is the number of bands, and W f (a, b) represents the wavelet energy coefficient, and ψ a,b (λ) is the adopted mother wavelet basis function; the general form is as follows:

[0094]

[0095] Among them, a is the scale factor representing the wavelet width, and b is the shift factor representing the wavelet position;

[0096] Step S12: Select the Mexican hat wavelet similar to the absorption characteristics of the vegetation index as the mother wavelet basis function, and only retain the wavelet power with the decomposition scale of 2 n (n = 1, 2,..., 10), which are respectively denoted as the 1st scale, the 2nd scale,..., the 10th scale;

[0097] Step S13: Based on the continuous wavelet transform, classify the tobacco virus disease by extracting the wavelet features extracted from the top 1% threshold of the determination coefficient between the wavelet energy coefficient and the tobacco virus disease level, and measure the chlorophyll content and leaf area index of the tobacco samples to obtain the chlorophyll characteristics and leaf area characteristics.

[0098] Preferably, in this embodiment, step S11 collects multi-spectral data of tobacco samples in the wavelength range of 350 to 2500 nm; uses continuous wavelet transform to extract spectral data features of tobacco; analyzes the original spectral curve and Gaussian function at different positions and scales to generate continuous wavelet power coefficients; and extracts weak information in different ice-white spectra. Through continuous wavelet transform, this embodiment can analyze the spectral characteristics of tobacco at different scales, capture information of different frequency components; extract key features in the tobacco spectrum, which helps subsequent disease level classification and chlorophyll content determination; and can effectively extract weak information in the spectrum, improving the signal-to-noise ratio of the data. In step S12, the Mexican hat wavelet, which is similar to the absorption characteristics of vegetation indices, is selected as the mother wavelet basis function; by choosing the Mexican hat wavelet in this embodiment, it can better match the absorption characteristics of vegetation indices, improving the correlation of features; retaining the wavelet power at specific scales can extract spectral features at different scales, enhancing the ability to identify diseases. In step S13, wavelet features are extracted based on the top 1% threshold of the determination coefficient between the wavelet energy coefficients extracted by continuous wavelet transform and the tobacco virus disease level; the chlorophyll content and leaf area index of tobacco samples are measured to obtain chlorophyll features and leaf area features; by extracting features highly correlated with the disease level in this embodiment, the tobacco virus disease can be accurately classified; combined with the chlorophyll content measurement, it can more comprehensively evaluate the health status of tobacco samples, providing a scientific basis for disease control.

[0099] Further, as Figure 3 shown, the process of obtaining chlorophyll features and leaf area features in step S13 specifically includes the following steps:

[0100] Step S131: Extract wavelet features based on the top 1% threshold of the determination coefficient between the wavelet energy coefficients extracted by continuous wavelet transform and the tobacco virus disease level; sort the determination coefficients at all scales from large to small, and use the method of threshold limitation for screening;

[0101] Step S132: Analyze the wavelet energy coefficients and the tobacco virus disease level to generate a determination coefficient. The determination coefficient values of the wavelet energy coefficients at different bands and different scales form a correlation scale map, obtaining the sensitivity of each wavelet energy coefficient to the tobacco virus disease;

[0102] Among them, the elements with a determination coefficient in the top 1% are marked as wavelet feature regions. In each wavelet feature region, the wavelet energy coefficient with the highest determination coefficient is retained as the extracted wavelet feature;

[0103] Step S133: Measure the chlorophyll content of the tobacco samples. For each plant, collect two representative leaves. Cut each leaf into three parts: upper, middle, and lower. Measure two values in the middle of each part. The average of the six values is the chlorophyll content. Take the mean value as the chlorophyll feature. For the leaf area index of each sample, take three samples and record the values. Take the mean and variance as the leaf area features.

[0104] Preferably, in step S131 of this embodiment, wavelet energy coefficients are extracted based on continuous wavelet transform; the wavelet energy coefficients are analyzed with the tobacco virus disease level, and the determination coefficient is calculated; the determination coefficients at all scales are sorted, and the top 1% of the high determination coefficient elements are selected by using a threshold-limiting method. This embodiment extracts features highly correlated with the tobacco virus disease level, namely the wavelet feature region; selects the wavelet energy coefficients most sensitive to the tobacco virus disease for subsequent analysis and modeling, improving the accuracy and reliability of the model. In step S132, the wavelet energy coefficients are analyzed with the tobacco virus disease level to generate the determination coefficient; a correlation scale map between the wavelet energy coefficients at different bands and scales and the disease level is constructed; the sensitivity of each wavelet energy coefficient to the tobacco virus disease is determined. This embodiment clarifies which wavelet energy coefficients are the most sensitive to the tobacco virus disease, so that these features can be preferentially considered for further analysis and application, providing an intuitive display of the sensitivity to the tobacco virus disease and helping to understand the importance of features at different bands and scales. In step S133, the chlorophyll content of the tobacco samples is measured. Each leaf is cut into three parts: upper, middle, and lower. Two values are measured for each part, and the average value is taken as the chlorophyll feature; for the leaf area index of each sample, three samples are taken and the values are recorded. Take the mean and variance as the leaf area features; it provides important indicators for the health status of tobacco, namely the chlorophyll content and the leaf area index; it can reflect the growth state and potential disease risks of tobacco, providing a scientific basis for tobacco planting and management.

[0105] Further, as Figure 4 shown, the process of constructing a tobacco virus disease discrimination model in step S2 specifically includes the following steps:

[0106] Step S21: Fuse the spectral wavelet features, chlorophyll features, and leaf area features, and update the ecological niche for network parameter allocation of the fused data. The ecological niche is determined by the competition strategy in the tobacco simulation ecosystem;

[0107] Step S22: Perform preprocessing operations such as missing value filling and standardization on the fused data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database and construct a tobacco virus disease discrimination model;

[0108] Step S23: Train the tobacco virus disease discrimination model with the training set. Among them, the network parameters of the tobacco virus disease discrimination model are assigned an initial niche; calculate the loss function corresponding to the network parameters of the tobacco virus disease discrimination model, and update the network parameters.

[0109] Preferably, in step S21 of this embodiment, the spectral wavelet features, chlorophyll features, and leaf area features are fused; the niche assigned to the network parameters is updated, and the niche is determined by simulating the competition strategy in the tobacco ecosystem; in this embodiment, by fusing multiple features (spectral wavelet, chlorophyll, leaf area), the comprehensive information volume of the data is improved, thereby enhancing the model's ability to identify tobacco virus diseases. At the same time, through the update of the niche, the competition strategy of tobacco in the ecosystem is simulated, enabling the model to better adapt to the disease identification requirements under different environmental conditions and improving the robustness and accuracy of the model. In step S22, preprocessing operations such as missing value filling and standardization are performed on the fused data; the data set is divided into a training set, a validation set, and a test set; the historical tobacco classification database is retrieved to construct the tobacco virus disease discrimination model. In this embodiment, through missing value filling and data standardization, the quality and consistency of the data are ensured, and the model deviation caused by data quality problems is reduced. The division of the data set helps the training and validation of the model, ensuring the stable performance of the model on different data subsets. In addition, using the historical tobacco classification database to construct the model can make full use of existing knowledge and experience, improving the accuracy and generalization ability of the model. In step S23, the training set is used to train the tobacco virus disease discrimination model and an initial niche is assigned; the loss function corresponding to the network parameters is calculated, and the network parameters are updated. In this embodiment, the model is trained with the training set and an initial niche is assigned, enabling the model to gradually optimize its parameters during the training process. Calculating the loss function and updating the network parameters helps to improve the performance of the model, enabling it to more accurately identify tobacco virus diseases.

[0110] Further, as Figure 5 shown, the process of updating the niche assigned to the network parameters for the fused data in step S21 specifically includes the following steps:

[0111] Step S211: Input the fused tobacco training set into a preset autoencoder neural network, and map the tobacco training set to the initial low-dimensional feature space with a low status through the encoder of the autoencoder neural network to obtain the initial low-dimensional features;

[0112] Step S212: Calculate the corresponding loss function for the initial low-dimensional features, determine the feature importance of the low-dimensional features; based on the feature importance, recursively optimize the feature weight vector of the autoencoder neural network; based on the feature weight vector, determine the new low-dimensional features;

[0113] Step S213: Reconstruct the new low-dimensional features through a compiler until the autoencoder neural network meets the preset iteration conditions.

[0114] Among them, the low-dimensional feature mapping formula of the autoencoder neural network in step S211:

[0115] Z = f encoder (X; W e , b e ) = σ(W e ·X + b e )

[0116] In the formula, X represents the input tobacco training set data matrix, with a dimension of (n×d), where n is the number of samples and d is the feature dimension; We e represents the weight matrix of the encoder, with a dimension of (k×d), where k is the dimension of the low-dimensional feature space; b e represents the bias vector of the encoder, with a dimension of (k×l); σ(·) represents the activation function (such as ReLU or Sigmoid); Z represents the mapped low-dimensional feature matrix, with a dimension of (n×k); map the high-dimensional tobacco spectral data to the low-dimensional feature space, extract key features, and reduce data redundancy.

[0117] The calculation formula of the low-dimensional feature loss function and the evaluation formula of feature importance in step S212:

[0118]

[0119] In the formula, X i represents the original high-dimensional feature vector of the i-th sample; Z i represents the low-dimensional feature vector of the i-th sample; f decoder (·) represents the decoder function, which is used to reconstruct the low-dimensional features into high-dimensional features; W d represents the weight matrix of the decoder, with a dimension of (d×k); b d represents the bias vector of the decoder, with a dimension of (d×l); λ represents the regularization coefficient, which is used to control the sparsity of the feature weight vector; w j represents the j-th feature weight; calculate the reconstruction error of the low-dimensional features, and evaluate the feature importance through the regularization term to optimize the feature weight vector.

[0120] The low-dimensional feature reconstruction and iterative optimization formula in step S213:

[0121]

[0122] In the formula, Z (t) represents the low-dimensional feature matrix at the t-th iteration; η represents the learning rate, which controls the iteration step size; Denote the gradient of the loss function with respect to the low-dimensional features; Iteratively optimize the low-dimensional features by the gradient descent method until the preset iteration conditions are met.

[0123] Preferably, in step S211 of this embodiment, the autoencoder consists of an encoder and a decoder. The encoder compresses the high-dimensional input data into a low-dimensional representation; The autoencoder is an unsupervised learning algorithm that can perform feature extraction and dimensionality reduction without label information; In this embodiment, by mapping the original data to a low-dimensional feature space, the complexity of the data is reduced while the main features of the data are retained; The key features of the tobacco data are initially extracted, laying a foundation for further analysis and processing. In step S212, calculate the reconstruction error or loss function to measure the difference between the low-dimensional features output by the encoder and the original data reconstructed by the decoder; Evaluate the importance of each feature through the loss function, and identify the key features that contribute more to the reconstruction error; Recursively optimize the weight vector of the autoencoder to minimize the reconstruction error and improve the model performance. In this embodiment, by optimizing the feature weight vector, the reconstruction accuracy of the autoencoder is improved, enabling the model to more effectively extract and represent the key information of the data; And by identifying important features, the generalization ability and robustness of the model are further improved. In step S213, repeatedly execute the encoding and decoding operations, continuously adjust the network parameters until the preset number of iterations is reached or the convergence conditions are met; Use the optimized weight vector to reconstruct the low-dimensional features back to a form close to the original data to verify the reconstruction ability of the model; Through iterative training, the autoencoder gradually approaches the optimal solution and finally reaches the preset iteration conditions, ensuring the stability and reliability of the model; The reconstructed data is close to the original data, indicating that the model has successfully extracted the core features of the data and can effectively restore the original information.

[0124] Furthermore, as Figure 6 shown, the process of calculating the loss function corresponding to the network parameters of the tobacco virus disease discrimination model in step S23 specifically includes the following steps:

[0125] Step S231: Calculate the loss function corresponding to the network parameters of the tobacco virus disease discrimination model, use the loss function as the adaptability score corresponding to the network parameters, and calculate the energy adjustment factor corresponding to the grid parameters; Input the energy adjustment factor and the preset maximum energy to calculate the mutation energy corresponding to the grid parameters;

[0126] Step S232: Based on the mutation energy and the preset migration and mutation adjustment coefficients, calculate the adaptability score to determine the dynamic mutation rate corresponding to the network parameters;

[0127] Step S233: Migrate the initial niche based on the dynamic mutation rate and update the grid parameters; input the tobacco sample set, sample the samples to obtain a sampling set containing m samples; train the t-th tobacco virus disease discrimination model with the sampling set D'; the one with the most votes from the T tobacco virus disease discrimination models is the final category;

[0128] Among them, the sample set is:

[0129] D = {(x 1 , y 1 ), (x 2 , y 2 ),...,(x n , y n )};

[0130] The sampling set containing m samples is:

[0131] D' = {(x i , y i ), i ∈ {1, 2,..., m}

[0132] Training the t'-th decision tree model with the sampling set D' is:

[0133] G t′ (x), t' ∈ {1, 2,..., T}.

[0134] Among them, step S231 represents the loss function and the adaptability score calculation formula:

[0135]

[0136] In the formula, y i represents the true label of the i-th sample; represents the predicted label of the i-th sample; θ j′ represents the j'-th network parameter; γ represents the regularization coefficient, used to control the sparsity of the network parameters; calculate the loss function of the tobacco virus disease discrimination model and evaluate the adaptability of the model;

[0137] Step S232 represents the dynamic mutation rate calculation formula:

[0138]

[0139] In the formula, μ t represents the dynamic mutation rate of the t-th iteration; μ base represents the base mutation rate; β represents the adjustment coefficient, controlling the attenuation speed of the mutation rate; represents the preset maximum loss value; dynamically adjust the mutation rate according to the loss function of the model and optimize the update strategy of the network parameters;

[0140] Step S233 Network Parameter Update and Sampling Formula:

[0141]

[0142] In the formula, represents the j'-th network parameter in the t-th iteration; Δθ j represents the update amount of the j'-th network parameter; ∈ represents the noise coefficient; represents Gaussian noise with a mean of 0 and a variance of σ 2 ; Diversity is introduced through mutation and noise to optimize the update process of network parameters.

[0143] Preferably, in step S231 of this embodiment, the loss function corresponding to the network parameters of the tobacco virus disease discrimination model is calculated, and the loss function is used as the adaptability score corresponding to the network parameters; the energy adjustment factor corresponding to the grid parameters is calculated; the energy adjustment factor and the preset maximum energy are input to calculate the mutation energy corresponding to the grid parameters. By calculating the loss function in this embodiment, the performance of the model can be quantified, thereby evaluating the adaptability of the network parameters; the energy adjustment factor can reflect the influence of the grid parameters on the model performance, which helps to optimize the structure and parameters of the model; the mutation energy is used to measure the change degree of the model under different parameter configurations, which helps to dynamically adjust the complexity and robustness of the model. In step S232, based on the mutation energy and the preset migration and mutation adjustment coefficients, the adaptability score is calculated to determine the dynamic mutation rate corresponding to the network parameters; the dynamic mutation rate can reflect the change trend of the model under different parameter configurations, which helps to optimize the training process of the model and improve the generalization ability of the model; by introducing the mutation energy and the migration adjustment coefficient in this embodiment, the adaptability of the network parameters can be more accurately evaluated, thereby optimizing the training strategy of the model. In step S233, the initial niche is migrated based on the dynamic mutation rate; the grid parameters are updated; the tobacco sample set is input, and the samples are sampled to obtain a sampling set containing m samples; the t-th tobacco virus disease discrimination model is trained with the sampling set D'; the category with the most votes cast by the T tobacco virus disease discrimination models is the final category; by migrating the initial niche in this example, a wider parameter space can be explored, improving the diversity and robustness of the model; updating the grid parameters helps to optimize the structure and parameter configuration of the model and improve the performance of the model; and by sampling the tobacco sample set, the computational complexity can be reduced and the training efficiency can be improved; through the voting mechanism of multiple models, the accuracy and reliability of classification can be improved (for the specific principle, refer to Appendix Figure 7 and Appendix Figure 8 ).

[0144] Furthermore, as Figure 9 shown, the process that the category with the most votes cast by the decision tree model in step S233 is the final category specifically includes the following steps:

[0145] Step S2331: Initialize the structure of the tobacco virus disease discrimination model. At each decision node, at a certain decision node, select features to perform splitting by maximizing the local structure entropy gain;

[0146] Step S2332: Input the tobacco sample set, sample the tobacco samples to obtain a sampling set including m samples; perform node splitting according to the selected features and their thresholds, and allocate the tobacco sample data to the left and right child nodes. Split the tobacco sample data according to the selected features and their optimal splitting points;

[0147] Step S2333: Perform leaf node calibration. At the leaf nodes of the decision tree, calibrate the node categories according to the majority lysis of the tobacco samples in the nodes; The one with the most votes cast by T tobacco virus disease discrimination models is the final category; After training, evaluate the classification accuracy of the tobacco virus disease discrimination model to obtain the classification effect.

[0148] Preferably, in step S2331 of this embodiment, the structure of the tobacco virus disease discrimination model is initialized; at each decision node, the feature that can most improve the classification accuracy is selected for splitting; at a certain decision node, splitting is performed by maximizing the local structure entropy gain; initializing the model structure lays the foundation for subsequent training and prediction; in this embodiment, by selecting the features that can improve the classification accuracy for splitting, it can be ensured that the model can effectively reduce uncertainty at each node, thereby improving the overall classification performance; the strategy of maximizing the local structure entropy gain helps to further optimize the classification effect at a specific node, making the model more accurate when dealing with complex data. In step S2332, the strategy of maximizing the local structure entropy gain helps to further optimize the classification effect at a specific node, making the model more accurate when dealing with complex data; according to the selected feature and its threshold, node splitting is performed, and the tobacco sample data is assigned to the left and right child nodes; according to the selected feature and its optimal splitting point, the tobacco sample data is split. In this example, the sample set is obtained through the sampling method, which can effectively reduce the calculation amount and improve the training efficiency; using the selected feature and its threshold for node splitting ensures the reasonable division of the data, making the samples in the subset as homogeneous as possible, thereby improving the classification accuracy; the selection of the optimal splitting point further improves the precision of the model, making the data of each child node more pure and contributing to the subsequent classification task. In step S2333, leaf node calibration is performed. At the leaf nodes of the decision tree, the node category is calibrated according to the majority lysis of the tobacco samples in the node; the one with the most votes cast by T tobacco virus disease discrimination models is the final category; after training is completed, the classification accuracy of the tobacco virus disease discrimination model is evaluated to obtain the classification effect. Leaf node calibration determines the category through majority voting, improving the robustness of the classification decision; the multi-model voting mechanism (ensemble learning) can effectively reduce the risk of overfitting and improve the generalization ability of the model on different data sets; evaluating the classification accuracy of the model ensures the effectiveness of the model in practical applications and provides reliable decision-making support for subsequent production applications.

[0149] Further, as Figure 10 shown, the process of evaluating the classification accuracy of the tobacco virus disease discrimination model in step S2333 specifically includes the following steps:

[0150] Step S23331: Construct a confusion matrix to analyze the performance of the test set in the tobacco virus disease discrimination model after training is completed. Among them, each cell in the confusion matrix contains the number of samples corresponding to the category; calculate the accuracy of the classification results of each category.

[0151] Step S23332: Calculate the classification efficiency of the classification results of each category according to the confusion matrix, and calculate the accuracy of the classification results of each category according to the classification efficiency of the classification results of each category and the probability of the category appearing in the total samples.

[0152] Step S23333: Determine the accuracy of the tobacco virus disease discrimination model based on the accuracy of the classification results of each category; and calibrate it with the preset accuracy. If it is less than the preset accuracy, adjust the tobacco virus disease discrimination model.

[0153] Among them, in Step S23331, confusion matrix and classification accuracy calculation:

[0154]

[0155] In the formula, TP c represents the number of true positive samples of category c; TN c represents the number of true negative samples of category c; FP c represents the number of false positive samples of category c; FN c represents the number of false negative samples of category c; Calculate the classification accuracy of each category to evaluate the classification performance of the model;

[0156] Step S23332 classification efficiency and accuracy calibration formula:

[0157]

[0158] In the formula, Precision c represents the precision rate of category c; Recall c represents the recall rate of category c; Calculate the classification efficiency of each category to further calibrate the accuracy of the model;

[0159] Step S23333 model accuracy evaluation formula:

[0160]

[0161] In the formula, C represents the total number of categories; Calculate the overall classification accuracy of the model and compare it with the preset accuracy to decide whether to adjust the model.

[0162] Preferably, in step S23331 of this embodiment, the confusion matrix is a tool for evaluating the performance of a classification model. By showing the relationship between the true labels and the model's predicted labels, it intuitively reflects the accuracy and effectiveness of the model. Through the confusion matrix in this embodiment, we can clearly understand the prediction accuracy of the model for different categories, identify which categories the model performs well in, and which categories have problems of confusion or misclassification, thus helping us understand the advantages and weaknesses of the model. In step S23332, the classification efficiency can be calculated through the diagonal elements (TP and TN) in the confusion matrix, which represent the number of correctly classified samples. Combining the probability of each category appearing in the total samples, the accuracy of each category is further calculated. To comprehensively evaluate the performance of the model for different categories, especially in the case of class imbalance, the traditional accuracy rate may not be the best evaluation index. By considering the class probabilities, the performance of the model in actual applications can be more accurately reflected. In step S23333, according to the classification result accuracy of each category calculated in step S23332, the accuracy of the entire tobacco virus disease classification model can be comprehensively evaluated. If the accuracy of the model is higher than the preset value, it indicates that the model performance is good and no adjustment is required; otherwise, the model needs to be optimized and adjusted. Through calibration with the preset accuracy in this embodiment, it can be ensured that the model meets the expected performance standards in actual applications. If the model does not reach the preset accuracy, the model performance can be improved by adjusting model parameters, increasing training data, or improving feature selection, etc.

[0163] Further, as Figure 11 shown, the process of adjusting the tobacco virus disease classification model according to the evaluation results in step S3 specifically includes the following steps:

[0164] Step S31: Input the collected optical wavelet features, chlorophyll features, and leaf area features into the tobacco virus disease classification model for classification and discrimination;

[0165] Step S32: Evaluate the classified tobacco data, compare the evaluation result with the preset evaluation result threshold. If it is greater than the preset evaluation result threshold, increase the tobacco sample data and adjust the tobacco classification model;

[0166] Step S33: Upload the final classification result to the tobacco pathology evaluation platform, obtain the tobacco planting management plan according to the classification result, and monitor the development of the disease in real time. If it is greater than the preset disease threshold, give an alarm.

[0167] Preferably, in step S31 of this embodiment, the collected optical wavelet features, chlorophyll features, and leaf area features are input into the tobacco virus disease classification model for classification and discrimination; by combining various plant physiological and spectral features, the accuracy and efficiency of tobacco virus disease recognition can be improved. These features can reflect the changes of plants in different disease states, thus helping the model to classify and discriminate more accurately. In step S32, the classified tobacco data is evaluated, and the evaluation result is compared with the preset evaluation result threshold. If it is greater than the preset evaluation result threshold, the tobacco sample data is increased, and the tobacco classification model is adjusted; in this embodiment, by dynamically adjusting the model parameters and increasing the sample data, the adaptability and robustness of the model are improved. When the evaluation result exceeds the preset threshold, it indicates that the current model may not be able to effectively identify certain diseases. Therefore, it is necessary to optimize its performance by increasing the sample data and adjusting the model, so as to improve the overall recognition accuracy. In step S33, the final classification result is uploaded to the tobacco pathology evaluation platform, and a tobacco planting management plan is obtained according to the classification result to monitor the disease development situation in real time. If it is greater than the preset disease threshold, a warning is issued; in this embodiment, by uploading the classification result to the evaluation platform in real time, dynamic monitoring and warning of tobacco virus disease can be realized. When the disease development exceeds the preset threshold, the system will automatically trigger the warning mechanism to help growers take preventive measures in time, thereby reducing the impact of the disease on tobacco yield and quality. This real-time monitoring and warning mechanism significantly improves the efficiency and response speed of disease management.

[0168] As Figure 12 shown, this embodiment also provides an embodiment of a tobacco virus disease remote sensing discrimination system based on multi-source data. In this embodiment, the tobacco virus disease remote sensing discrimination system based on multi-source data is applied to the tobacco virus disease remote sensing discrimination method based on multi-source data as described in the above embodiment. The tobacco virus disease remote sensing discrimination system based on multi-source data includes:

[0169] A tobacco sample collection module 1, which is used to collect spectral data of tobacco samples in the wavelength range of 350 to 2500, and extract the spectral data features of tobacco samples by using continuous wavelet transform; classify the grades of tobacco virus diseases, measure the chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll features and leaf area features;

[0170] A discrimination model construction module 2, which is used to fuse the spectral wavelet features, chlorophyll features, and leaf area features, preprocess the fused data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database and construct a tobacco virus disease discrimination model;

[0171] An evaluation and discrimination module 3 is used to input the collected optical wavelet features, chlorophyll features, and leaf area features into a tobacco virus disease classification model for classification and discrimination; evaluate the data classified by the tobacco virus disease classification model, and adjust the tobacco virus disease classification model according to the evaluation results.

[0172] Preferably, in this embodiment, the tobacco sample collection module is used for a wavelength range of 350 to 2500 nanometers, covering the visible light and near-infrared regions, which helps to capture the spectral information of tobacco samples; it is used to extract the spectral data features of tobacco samples. CWT is an effective signal processing method that can decompose signals and extract their time-frequency features, thereby helping to identify the spectral features of tobacco samples; the spectral features of tobacco samples are extracted, providing basic data for subsequent disease level classification and chlorophyll content determination; it enhances the ability to analyze the spectral information of tobacco samples, helping to more accurately identify and analyze tobacco samples. The discrimination model construction module 2 fuses the spectral wavelet features, chlorophyll features, and leaf area features to form a comprehensive data set. This multi-dimensional feature fusion can improve the classification accuracy and robustness of the model. Preprocess the fused data, including operations such as standardization and normalization, to ensure the quality and consistency of the data. Divide the data set into a training set, a validation set, and a test set for model training, validation, and final evaluation. This embodiment provides a comprehensive data set for constructing a tobacco virus disease classification model; data preprocessing improves data quality and reduces the impact of noise on the model, thereby enhancing the performance of the model; data set division ensures the effectiveness and generalization ability of model training. The evaluation and discrimination module inputs the collected optical wavelet features, chlorophyll features, and leaf area features into the tobacco virus disease classification model; evaluates according to the model classification results, and adjusts the model parameters according to the evaluation results to optimize the model performance; realizes the automatic classification and identification of tobacco virus diseases, improving the efficiency and accuracy of tobacco virus disease detection; the evaluation and adjustment process ensures the stability and reliability of the model in practical applications.

[0173] As Figure 13 shown, this embodiment provides an embodiment of an electronic device. In this embodiment, the electronic device 4 includes a processor 41 and a memory 42 coupled to the processor 41.

[0174] The memory 42 stores program instructions for implementing the layout method of the tobacco virus disease remote sensing discrimination method based on multi-source data in any of the above embodiments.

[0175] The processor 41 is used to execute the program instructions stored in the memory 42 to perform the layout of the tobacco virus disease remote sensing discrimination method based on multi-source data.

[0176] Among them, the processor 41 can also be referred to as a CPU (Central Processing Unit). The processor 41 may be an integrated circuit chip with signal processing capabilities. The processor 41 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0177] Furthermore, Figure 14 FIG. is a schematic structural diagram of a storage medium according to an embodiment of the present application. The storage medium 5 of the embodiment of the present application stores program instructions 51 capable of implementing all the above methods. Among them, the program instructions 51 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, or a tablet.

[0178] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, indirect couplings or communication connections of devices or units, and can be in electrical, mechanical, or other forms.

[0179] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. The above is only the embodiment of the present invention, and does not limit the patent scope of the present invention. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

[0180] The specific embodiments of the invention have been described in detail above. However, these are only examples, and the invention is not limited to the specific embodiments described above. For those skilled in the art, any equivalent modifications or substitutions to the invention are also within the scope of the invention. Therefore, all equivalent transformations, modifications, improvements, etc. made without departing from the spirit and principles of the invention should be covered by the scope of the invention.

Claims

1. A remote sensing differentiation method for tobacco virus diseases based on multi-source data, characterized in that: The tobacco virus disease remote sensing differentiation method based on multi-source data includes: Collect spectral data of tobacco samples, and use continuous wavelet transform to extract spectral data features of tobacco samples; classify tobacco virus diseases, measure chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll characteristics and leaf area characteristics; The spectral wavelet features, chlorophyll features and leaf area features are integrated, and the integrated data are preprocessed to form a data set; the data set is divided into a training set, a validation set and a test set; the historical tobacco classification database is retrieved to build a tobacco virus disease differentiation model; The collected optical wavelet features, chlorophyll features and leaf area features are input into the tobacco virus disease classification model for classification and identification; the data classified by the tobacco virus disease classification model is evaluated, and the tobacco virus disease classification model is adjusted according to the evaluation results.

2. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 1, characterized in that: The process of extracting the spectral data features of tobacco samples using continuous wavelet transform includes the following steps: Collect multi-spectral data of tobacco samples with a wavelength range of 350 to 2500, and use continuous wavelet transform to extract the spectral data characteristics of tobacco; analyze the spectrum after wavelet transform at different positions and scales to generate continuous wavelet capacity coefficients; extract weak information from the spectra of different diseases; The Mexican hat wavelet with similar absorption characteristics to the vegetation index is selected as the mother wavelet basis function, and only the decomposition scale of 2 is retained. n The wavelet powers of (n=1, 2, ..., 10) are recorded as the 1st scale, the 2nd scale, ..., the 10th scale respectively; Tobacco virus diseases are graded based on wavelet features extracted based on continuous wavelet transform and the top 1% threshold of the determination coefficient between the wavelet energy coefficient and the tobacco virus disease grade. The chlorophyll content and leaf area index of tobacco samples are measured to obtain chlorophyll characteristics and leaf area characteristics.

3. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 2, characterized in that: The process of obtaining chlorophyll characteristics and leaf area characteristics includes the following steps: The chlorophyll content of tobacco samples was measured. Two representative leaves were collected from each plant and each leaf was cut into three parts: upper, middle and lower. Two values ​​were measured in the middle of each part. The six values ​​were averaged as the chlorophyll content. The mean was taken as the chlorophyll characteristic. The leaf area index of each sample was sampled three times, and the mean and variance were taken as the leaf area characteristics.

4. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 1, characterized in that: The process of building a tobacco virus disease differentiation model includes the following steps: The spectral wavelet features, chlorophyll features and leaf area features are fused, and the fused data are used to update the ecological niche of network parameter allocation. The ecological niche is determined by the competition strategy in the tobacco simulation ecosystem. The fused data is filled with missing values ​​and preprocessed with standardization to form a data set; the data set is divided into a training set, a validation set, and a test set; the historical tobacco classification database is retrieved to build a tobacco virus disease differentiation model; The tobacco virus disease differentiation model is trained using a training set, wherein the network parameters of the tobacco virus disease differentiation model are assigned initial ecological niches; the loss function corresponding to the network parameters of the tobacco virus disease differentiation model is calculated, and the network parameters are updated.

5. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 4, characterized in that: The process of updating the niche of network parameter allocation using the fused data includes the following steps: The fused tobacco training set is input into a preset autoencoder neural network, and the tobacco training set is mapped to an initial low-dimensional feature space through the encoder of the autoencoder neural network to obtain an initial low-dimensional feature; Calculate the initial low-dimensional features to obtain the corresponding loss function and determine the feature importance of the low-dimensional features; based on the feature importance, recursively optimize the feature weight vector of the autoencoder neural network; based on the feature weight vector, determine the new low-dimensional features; The compiler reconstructs the new low-dimensional features until the autoencoder neural network meets the preset iteration conditions; Among them, the low-dimensional feature mapping formula of the autoencoder neural network is: Z=f encoder (;W e ,b e )=σ(W e ·X+b e ) Where X represents the input tobacco training set data matrix, with a dimension of (n×d), where n is the number of samples and d is the feature dimension; W e represents the weight matrix of the encoder, with a dimension of (k×d), where k is the dimension of the low-dimensional feature space; b e represents the bias vector of the encoder, with a dimension of (k×l); σ(·) represents the activation function; Z represents the low-dimensional feature matrix after mapping, with a dimension of (n×k); Low-dimensional feature loss function calculation and feature importance evaluation formula: In the formula, X i Represents the original high-dimensional feature vector of the i-th sample; Z i Represents the low-dimensional feature vector of the i-th sample; f decoder (·) represents the decoder function, which is used to reconstruct low-dimensional features into high-dimensional features; W d represents the weight matrix of the decoder, with dimension (d×k); b d represents the bias vector of the decoder, with a dimension of (d×l); λ represents the regularization coefficient, which is used to control the sparsity of the feature weight vector; w j represents the jth feature weight; Low-dimensional feature reconstruction and iterative optimization formula: In the formula, Z (t) represents the low-dimensional feature matrix of the tth iteration; η represents the learning rate, which controls the iteration step size; Represents the gradient of the loss function with respect to the low-dimensional features; the low-dimensional features are iteratively optimized through the gradient descent method until the preset iteration conditions are met.

6. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 4, characterized in that: The process of calculating the loss function corresponding to the network parameters of the tobacco virus disease discrimination model includes the following steps: Calculate the loss function corresponding to the network parameters of the tobacco virus disease differentiation model, use the loss function as the fitness score corresponding to the network parameters, calculate the energy adjustment factor corresponding to the grid parameters; input the energy adjustment factor and the preset maximum capacity into the calculation of the mutation energy corresponding to the grid parameters; Based on the mutation energy and the preset migration and mutation adjustment coefficients, the fitness score is calculated to determine the dynamic mutation rate corresponding to the network parameters; The initial ecological niche is migrated based on the dynamic mutation rate, and the network parameters are updated; the tobacco sample set is input, and the samples are sampled to obtain a sampling set containing m samples; the sampling set D′ is used to train the tth tobacco virus disease discrimination model; the one with the most votes from the T tobacco virus disease discrimination models is the final category.

7. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 6, characterized in that: The process of the decision tree model casting the most votes as the final category includes the following steps: Initialize the structure of the tobacco virus disease discrimination model. At each decision node, select features to split by maximizing the local structural entropy gain. Input a tobacco sample set, sample the tobacco samples, and obtain a sampling set including m samples; perform node splitting according to the selected features and their thresholds, assign tobacco sample data to left and right child nodes, and split the tobacco sample data according to the selected features and their optimal splitting points; Leaf node calibration is performed. At the leaf nodes of the decision tree, the node categories are calibrated according to the majority of tobacco samples in the nodes; the category with the most votes cast by the tobacco virus disease discrimination model is the final category; after the training is completed, the classification accuracy of the tobacco virus disease discrimination model is evaluated to obtain the classification effect.

8. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 7, characterized in that: The process of evaluating the classification accuracy of the tobacco virus disease discrimination model includes the following steps: Construct a confusion matrix to analyze the performance of the test set in the trained tobacco virus disease discrimination model, where each cell in the confusion matrix contains the number of samples of the corresponding category; calculate the accuracy of the classification results of each category; The classification efficiency of each category classification result is calculated according to the confusion matrix, and the accuracy of each category classification result is calculated according to the classification efficiency of each category classification result and the probability of the category appearing in the total samples; The accuracy of the tobacco virus disease differentiation model is determined based on the accuracy of the calculated classification results of each category; and the accuracy is calibrated with the preset accuracy. If it is less than the preset accuracy, the tobacco virus disease differentiation model is adjusted.

9. The method for distinguishing tobacco virus diseases by remote sensing based on multi-source data according to claim 1, characterized in that: The process of adjusting the tobacco virus disease classification model based on the evaluation results includes the following steps: The collected spectral wavelet features, chlorophyll features and leaf area features are input into the tobacco virus disease classification model for classification and identification; Evaluate the classified tobacco data, compare the evaluation result with the preset evaluation result threshold, and if it is greater than the preset evaluation result threshold, add tobacco sample data and adjust the tobacco classification model; The final classification results are uploaded to the tobacco pathology assessment platform, and the tobacco planting management plan is obtained based on the classification results. The disease development is monitored in real time, and an early warning is issued if it is greater than the preset disease threshold.

10. A tobacco virus disease remote sensing differentiation system based on multi-source data, which is applied to the tobacco virus disease remote sensing differentiation method based on multi-source data as claimed in any one of claims 1 to 9, characterized in that: The tobacco virus disease remote sensing differentiation system based on multi-source data includes: The tobacco sample collection module is used to collect spectral data of tobacco samples and extract the spectral data characteristics of tobacco samples using continuous wavelet transform; classify tobacco virus diseases, measure the chlorophyll content and leaf area index of tobacco samples, and obtain chlorophyll characteristics and leaf area characteristics; Construct a differentiation model module to fuse spectral wavelet features, chlorophyll features, and leaf area features, preprocess the fused data to form a data set; divide the data set into a training set, a validation set, and a test set; retrieve the historical tobacco classification database to construct a tobacco virus disease differentiation model; The evaluation and differentiation module is used to input the collected optical wavelet features, chlorophyll features and leaf area features into the tobacco virus disease classification model for classification and identification; evaluate the data classified by the tobacco virus disease classification model, and adjust the tobacco virus disease classification model according to the evaluation results.

Citation Information

Cited By

  • Tobacco virus classification model construction method based on tobacco hyperspectrum

    CN121637134A