Tobacco disease and insect pest identification method and system
Through the multi-spectral reflection data acquisition and dynamic convolutional algorithm combined with multimodal feature fusion, the problems of low accuracy and poor environmental adaptability of tobacco leaf disease and pest recognition in the existing technology are solved, and accurate identification and automated prevention and control of complex pests are achieved, and the level of intelligence in agricultural production is improved.
Patent Information
- Application Number
- CN202510341082.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing tobacco leaf pest recognition methods rely on a single image feature, making it difficult to cope with complex pest symptoms, and have poor adaptability to environmental changes, resulting in low recognition accuracy and unable to meet the needs of accurate monitoring and rapid response of modern agricultural production.
Multispectral reflection data acquisition and dynamic convolutional algorithm are used to extract local features, combine multimodal feature fusion and adaptive weight classifier for pest identification, generate pest type identifiers and adjust agricultural equipment operating parameters.
It improves the accuracy and robustness of pest identification, realizes accurate identification of complex pest symptoms, reduces the use of pesticides, and improves the intelligence level of agricultural production and tobacco leaf quality.
Smart Images

Figure CN120259685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural pest and disease monitoring and control, and more specifically, the present invention relates to a method and system for identifying tobacco leaf pests and diseases. Background Art
[0002] In the field of agricultural planting, especially in tobacco planting, the timely and accurate identification of pests and diseases is crucial for ensuring the healthy growth of crops and increasing yields. Traditional methods for identifying tobacco leaf pests and diseases mainly rely on manual observation and empirical judgment. This method is not only inefficient but also easily affected by subjective factors, resulting in inaccurate identification results. In recent years, with the development of image recognition technology and sensor technology, pest and disease identification methods based on image analysis have gradually emerged. These methods collect image information of tobacco leaves and use computer vision algorithms to analyze the images to achieve automatic identification of pests and diseases. However, most of the existing image recognition methods only rely on a single image feature, such as color or texture, and have limited effects on identifying complex pest and disease symptoms. Especially when facing mixed infections of multiple pests and diseases, the recognition accuracy is relatively low. In addition, the existing technology also has problems such as poor adaptability to environmental changes and difficulty in real-time dynamic adjustment of recognition parameters in practical applications, which limit its wide application in large-scale agricultural production.
[0003] In the process of implementing the embodiments of the present invention, the inventors found that there are at least the following problems or defects in the existing technology: Most of the existing pest and disease identification methods rely on a single image feature and are difficult to cope with complex pest and disease symptoms; they have poor adaptability to environmental changes and are difficult to adjust recognition parameters in real time dynamically, resulting in low recognition accuracy and unable to meet the requirements of modern agricultural production for precise pest and disease monitoring and rapid response. Summary of the Invention
[0004] The present invention provides a method and system for identifying tobacco leaf pests and diseases.
[0005] In the first aspect of the present invention, a method for identifying tobacco leaf pests and diseases is provided, including:
[0006] S1. Collect tobacco leaf images and obtain multispectral reflection data on the surface of the tobacco leaves through an image sensor;
[0007] S2. Preprocess the multispectral reflection data to generate a standardized image matrix;
[0008] S3. Extract local feature vectors of the standardized image matrix based on the dynamic convolution kernel algorithm;
[0009] S4. Through a multimodal feature fusion algorithm, match the local feature vectors with a preset pest and disease spectral feature library to generate a fusion feature tensor;
[0010] S5. Use an adaptive weight classifier to perform non - linear mapping on the fused feature tensor and output the pest and disease probability distribution;
[0011] S6. Execute logical decision - making according to the probability distribution to generate a pest and disease type identifier;
[0012] S7. Trigger a control instruction according to the type identifier to adjust the operating parameters of the associated agricultural equipment.
[0013] Further, the step S2 includes:
[0014] S21. Perform noise suppression on the multi - spectral reflection data and calculate the neighborhood gradient variance σ of each pixel point 2 , when σ 2 > the threshold T, update the pixel value using the following correction formula:
[0015]
[0016] where I(x,y) is the original pixel value, is the Gaussian weight kernel, and k is the window radius;
[0017] S22. Perform contrast enhancement on the denoised data and adjust the pixel dynamic range through a piece - wise linear transformation function;
[0018] S23. Normalize the enhanced data to the interval [0,1] to generate the standardized image matrix.
[0019] Further, the dynamic convolution kernel algorithm in the step S3 includes:
[0020] S31. Dynamically generate the convolution kernel size according to the gradient magnitude of the local region of the image
[0021]
[0022]
[0023]
[0024] S32. Construct a deformable convolution kernel parameter group θ = {δ x ,δ y ,α}, where the offsets δ x ,δ y are determined by the eigenvectors of the local curvature matrix, and the scaling factor
[0025] S33. Calculate the feature response map through the following formula:
[0026]
[0027] Among them, W(i,j) is a trainable weight matrix.
[0028] Further, the step S4 includes:
[0029] S41. Construct a multi-modal feature vector V = [v texture , v color , v spectral , where:
[0030] v texture is the texture feature extracted by the Gabor filter bank;
[0031] v color is the KL divergence feature of the HSV space color histogram;
[0032] v spectral is the reflectance ratio feature of the near-infrared band to the visible light band;
[0033] S42. Calculate the feature weights using the attention mechanism:
[0034]
[0035] where W m , b m are learnable parameters, and σ is the Sigmoid function;
[0036] S43. Generate the fused feature tensor through weighted summation M m is a preset modal matching matrix.
[0037] Further, the calculation process of the adaptive weight classifier in the step S5 includes:
[0038] S51. Construct a two-stream neural network, where the first branch processes spatial features and the second branch processes spectral features;
[0039] S52. Define the cross-adaptation function:
[0040]
[0041] where w1, w2 are dynamic weights, and γ is the coupling coefficient;
[0042] S53. Optimize the weight parameters through backpropagation to minimize the loss function where C is the number of pest and disease categories.
[0043] Further, the logical decision-making process in the step S6 includes:
[0044] S61. Calculate the confidence index
[0045]
[0046] When C < 0.3, start the multi - model voting mechanism;
[0047] S62. Use the fuzzy logic rule base for decision correction and define the membership function:
[0048]
[0049] S63. Perform decision tree traversal. When the confidence of the leaf node is lower than the threshold, trigger the manual review mark.
[0050] Furthermore, the control instruction generation in step S7 includes:
[0051] S71. Establish a mapping relationship table between pest and disease types and application parameters;
[0052] S72. Calculate the application rate through the following formula:
[0053]
[0054] where S is the proportion of the infected area, D is the disease severity level, t is the duration, and k1, k2, λ are adjustment coefficients;
[0055] S73. Generate a JSON control instruction package containing application coordinates, chemical types, and dosage parameters.
[0056] Furthermore, the calculation of the gradient amplitude in step S31 uses an improved Sobel operator:
[0057] Horizontal direction operator
[0058]
[0059] Vertical direction operator V = H T , gradient amplitude
[0060] Furthermore, the implementation of the attention mechanism in step S42 includes:
[0061] Use a multi - head attention structure. The dimension of each head is 64. The query matrix Q = W q ·T, the key matrix K = W k ·T, the value matrix V = W v ·T, and the attention weight is calculated as:
[0062]
[0063] where, d k is the dimension of the key vector.
[0064] In the second aspect of the present invention, a tobacco leaf pest and disease identification system is provided, including:
[0065] An image acquisition module, configured to acquire tobacco leaf images and obtain multispectral reflection data on the surface of tobacco leaves through an image sensor;
[0066] A processing module, configured to preprocess the multispectral reflection data to generate a standardized image matrix; extract local feature vectors of the standardized image matrix based on a dynamic convolution kernel algorithm; match the local feature vectors with a preset pest and disease spectral feature library through a multimodal feature fusion algorithm to generate a fusion feature tensor; perform a non-linear mapping on the fusion feature tensor using an adaptive weight classifier to output a pest and disease probability distribution; execute a logical decision based on the probability distribution to generate a pest and disease type identifier;
[0067] A control module, configured to trigger a control instruction according to the type identifier and adjust the operating parameters of associated agricultural equipment;
[0068] A communication module, configured to transmit the control instruction to an agricultural equipment execution terminal.
[0069] According to the above embodiments of the present invention, there are at least the following beneficial effects: First, by acquiring the multispectral images of tobacco leaves and combining the dynamic convolution kernel algorithm and multimodal feature fusion technology, accurate extraction and efficient matching of the characteristics of tobacco leaf pests and diseases can be achieved. This method can not only make full use of the advantages of different spectral information, but also effectively improve the accuracy and robustness of pest and disease identification. Especially when facing complex pest and disease symptoms and mixed infections of multiple pests and diseases, the identification effect can be significantly improved.
[0070] Second, the present invention adopts an adaptive weight classifier and a logical decision mechanism, which can dynamically adjust the decision-making process according to the probability distribution of pests and diseases and generate corresponding control instructions to adjust the operating parameters of agricultural equipment. This can not only achieve automatic monitoring and precise prevention and control of pests and diseases, but also reduce the usage amount of pesticides, lower the agricultural production cost, and at the same time improve the intelligent level of agricultural production and the quality of tobacco leaves. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become easily understandable. In the drawings, several embodiments of the present invention are shown in an exemplary rather than restrictive manner, wherein:
[0072] Figure 1 is a schematic flowchart of a tobacco leaf pest and disease identification method provided by an embodiment of the present invention;
[0073] Figure 2Schematic structural diagram of a tobacco leaf pest and disease identification system provided by an embodiment of the present invention;
[0074] Figure 3 Schematically shows a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0075] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and then implement the present invention, rather than limiting the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present invention more thorough and complete, and to be able to convey the scope of the present invention fully to those skilled in the art.
[0076] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, equipment, method or computer program product. Therefore, the present invention can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0077] It should be noted that the number of any element in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0078] Next, refer to Figure 1 , Figure 1 Schematic flowchart of a tobacco leaf pest and disease identification method provided by an embodiment of the present invention. As Figure 1 shown, a tobacco leaf pest and disease identification method 100 includes:
[0079] S1. Collect tobacco leaf images and obtain multi-spectral reflection data on the surface of the tobacco leaves through an image sensor;
[0080] S2. Preprocess the multi-spectral reflection data to generate a standardized image matrix;
[0081] S3. Extract local feature vectors of the standardized image matrix based on the dynamic convolution kernel algorithm;
[0082] S4. Through the multi-modal feature fusion algorithm, match the local feature vectors with a preset pest and disease spectral feature library to generate a fusion feature tensor;
[0083] S5. Use an adaptive weight classifier to perform non-linear mapping on the fusion feature tensor and output a pest and disease probability distribution;
[0084] S6. Perform logical decision-making according to the probability distribution to generate a pest and disease type identifier;
[0085] S7. Trigger a control instruction according to the type identifier to adjust the operating parameters of the associated agricultural equipment.
[0086] It should be noted that in the present invention, first, a tobacco leaf image needs to be collected, and multispectral reflection data on the surface of the tobacco leaf is obtained through an image sensor. Here, the "multispectral reflection data" refers to the intensity information of the reflected light on the surface of the tobacco leaf at multiple different wavelength bands, which usually include visible light wavelength bands (such as red, green, and blue light) and near-infrared wavelength bands, etc. Multispectral imaging technology can capture the characteristics of tobacco leaves under different spectra, thus providing richer information for subsequent pest and disease identification. This data collection method can effectively reflect the subtle changes on the surface of the tobacco leaf, including characteristics such as disease spots and pest damage, laying a foundation for subsequent image processing and analysis.
[0087] Specifically, an image sensor is a device that can convert optical signals into electrical signals. Common types include charge-coupled device (CCD) or complementary metal-oxide-semiconductor (CMOS) sensors. When collecting multispectral reflection data, the sensor needs to image the tobacco leaf at multiple preset wavelength bands. For example, the red wavelength band (about 600 - 700 nm), the green wavelength band (about 490 - 570 nm), the blue wavelength band (about 450 - 490 nm), and the near-infrared wavelength band (about 700 - 1000 nm) can be selected. The selection of these wavelength bands is based on the reflection characteristics of tobacco leaves under different spectra and the influence of pests and diseases on spectral reflection. For example, certain diseases will cause the reflectivity of tobacco leaves in the near-infrared wavelength band to decrease, while the reflectivity change in the visible light wavelength band is relatively small. These differences can be effectively distinguished through multispectral imaging. In addition, the collection device needs to have high resolution and high sensitivity to ensure that the acquired image data can accurately reflect the details on the surface of the tobacco leaf.
[0088] Preferably, when collecting the tobacco leaf image, multi-angle imaging technology can be adopted to take pictures of the tobacco leaf from different directions to obtain more comprehensive reflection data. At the same time, in order to improve the stability and reliability of the data, a calibration module can be introduced during the collection process to calibrate the spectral response of the sensor in real time to ensure the data consistency under different wavelength bands. In addition, automatic focusing and exposure control functions can also be adopted to meet the imaging requirements of tobacco leaves under different lighting conditions. After the data collection is completed, the collected multispectral images can be preliminarily screened to remove abnormal images caused by equipment failures or environmental interferences, thereby improving the efficiency and accuracy of subsequent processing.
[0089] In some embodiments, the step S2 includes:
[0090] S21. Perform noise suppression on the multispectral reflection data and calculate the neighborhood gradient variance σ of each pixel point 2 , when σ 2When it is greater than the threshold T, the pixel value is updated using the following correction formula:
[0091]
[0092] where I(x, y) is the original pixel value, is the Gaussian weight kernel, and k is the window radius;
[0093] S22. Perform contrast enhancement on the denoised data, and adjust the pixel dynamic range through a piecewise linear transformation function;
[0094] S23. Normalize the enhanced data to the interval [0, 1] to generate the standardized image matrix.
[0095] It should be noted that when preprocessing the collected multi-spectral reflection data, the main purposes are to remove noise, enhance the image contrast, and normalize it to generate a standardized image matrix. The noise suppression here refers to eliminating the random interference signals in the image through a specific algorithm, while the contrast enhancement is to adjust the pixel dynamic range of the image to make the details of the image clearer. The final normalization is to adjust the pixel values of the image data to a unified range (such as the interval [0, 1]) so that the subsequent processing steps can be carried out more efficiently. These preprocessing steps can effectively improve the image quality and provide a more accurate data basis for subsequent feature extraction and pest identification.
[0096] Specifically, in the noise suppression stage, calculate the neighborhood gradient variance σ of each pixel point 2 , when σ 2 exceeds the set threshold T, the pixel value is corrected using the Gaussian weight kernel G(x, y). The Gaussian weight kernel is a filter based on the Gaussian function, and its formula is
[0097]
[0098] where σ is the standard deviation, which is used to control the width of the Gaussian kernel. The window radius k determines the size of the neighborhood and is usually selected according to the resolution and noise level of the image. For example, for a high-resolution image, the window radius can be set to 3 or 5 pixels. In the contrast enhancement stage, the pixel dynamic range is adjusted through a piecewise linear transformation function, which can be adaptively adjusted according to the histogram distribution of the image to enhance the image contrast. Finally, the enhanced data is normalized to the interval [0, 1], and the specific method is to divide each pixel value by the maximum pixel value of the image.
[0099] Preferably, during the noise suppression process, an adaptive threshold T can be adopted, which can be dynamically adjusted according to the local statistical characteristics of the image to better adapt to images with different noise levels. For example, for image regions with more noise, the threshold T can be appropriately increased to more effectively remove the noise. During the contrast enhancement stage, histogram equalization technology can be introduced to further optimize the contrast of the image. Histogram equalization enhances the overall contrast of the image by adjusting the histogram distribution of the image to make the gray value distribution of the image more uniform. In addition, the normalization process can also be combined with the adjustment of the dynamic range of the image. For example, through linear or non-linear transformation, the pixel values of the image are mapped to a more appropriate interval to improve the efficiency and accuracy of subsequent processing.
[0100] In some embodiments, the dynamic convolution kernel algorithm in step S3 includes:
[0101] S31. Dynamically generate the convolution kernel size according to the gradient magnitude of the local region of the image
[0102]
[0103] of the convolution kernel
[0104]
[0105] S32. Construct a deformable convolution kernel parameter group θ = {δ x , δ y , α}, where the offset δ x , δ y is determined by the eigenvectors of the local curvature matrix, and the scaling factor
[0106] S33. Calculate the feature response map through the following formula:
[0107]
[0108] where W(i, j) is a trainable weight matrix.
[0109] It should be noted that the dynamic convolution kernel algorithm is an adaptive feature extraction method. Its core lies in dynamically adjusting the size and shape of the convolution kernel according to the gradient magnitude of the local region of the image, so as to better capture the key features in the image. Here, the gradient magnitude refers to the intensity of the change in pixel values in the image, reflecting the significance of image edges and textures. By dynamically generating the convolution kernel size and constructing a deformable convolution kernel parameter group, this algorithm can adapt to the feature changes in different image regions, improving the flexibility and accuracy of feature extraction. This method is particularly suitable for the identification of tobacco leaf diseases and pests because the symptoms of diseases and pests often have irregular shapes and edge features, and the dynamic convolution kernel algorithm can more effectively extract these features, providing richer information for subsequent classification of diseases and pests.
[0110] Specifically, in the first step of the dynamic convolution kernel algorithm, the gradient magnitude of the local region of the image is calculated
[0111]
[0112] where I is the grayscale value of the image. The calculation of the gradient magnitude can be achieved through an improved Sobel operator, which calculates the gradient in the horizontal and vertical directions respectively and can more accurately reflect the edge information. The convolution kernel size is dynamically generated according to the gradient magnitude
[0113]
[0114] where G max is the maximum gradient magnitude of the local region. This dynamic generation method can adjust the size of the convolution kernel according to the complexity of the local image, so as to better capture the detail features. In the constructed deformable convolution kernel parameter group Θ = {Δx, Δy, s}, the offsets Δx and Δy are determined by the eigenvectors of the local curvature matrix, and the scaling factor s is dynamically adjusted according to the magnitude of the local curvature, further enhancing the adaptability of the convolution kernel. Through these parameters, the feature response map can be calculated, and thus the key features in the image can be extracted.
[0115] Preferably, when calculating the gradient magnitude, a more advanced edge detection algorithm, such as the Canny edge detection algorithm, can be used to improve the accuracy and anti-noise performance of edge detection. In addition, the dynamic generation formula of the convolution kernel size can be adjusted according to the actual application requirements, for example, introducing a non-linear factor to enhance the sensitivity to high-gradient regions. When constructing the deformable convolution kernel parameter group, multi-scale analysis can be introduced. By calculating the offsets and scaling factors at different scales, the robustness of feature extraction can be further improved. For example, the wavelet transform can be combined to perform multi-scale decomposition on the image, and then the convolution kernel parameters can be calculated separately at each scale, so as to better capture the multi-scale features in the image. In addition, deep learning methods can also be introduced to automatically learn the optimal configuration of the convolution kernel parameters through training a neural network, further improving the performance of feature extraction.
[0116] In some embodiments, the step S4 includes:
[0117] S41. Construct a multi-modal feature vector V = [v texture , v color , v spectral , where:
[0118] v texture is the texture feature extracted by the Gabor filter bank;
[0119] v colorIt is the KL divergence feature of the HSV space color histogram;
[0120] v spectral It is the reflectance ratio feature between the near-infrared band and the visible light band;
[0121] S42. Calculate the feature weights using the attention mechanism:
[0122]
[0123] Among them, W m , b m are learnable parameters, and σ is the Sigmoid function;
[0124] S43. Generate the fused feature tensor through weighted summation M m is a preset modal matching matrix.
[0125] It should be noted that the multi-modal feature fusion algorithm is a process of integrating different types of feature vectors extracted from the standardized image matrix to generate a fused feature tensor. The multi-modal features here refer to various different types of features extracted from the image, such as texture features, color features, and spectral features, etc. These features respectively reflect the information of the image in different aspects. By fusing these features, the pest and disease characteristics of tobacco leaves can be described more comprehensively, thereby improving the accuracy and robustness of pest and disease identification. Specifically, this algorithm constructs multi-modal feature vectors and uses the attention mechanism to calculate the feature weights, and finally generates a fused feature tensor, providing richer information for subsequent classification and decision-making.
[0126] Specifically, in the multi-modal feature fusion algorithm, first construct the multi-modal feature vector F = [F texture , F color , F spectral . Among them, the texture feature F texture is extracted through the Gabor filter bank. The Gabor filter is a filter that can effectively capture the local texture information of the image. Its parameters include direction and frequency. Usually, multiple direction and frequency combinations are set to extract multi-scale texture features. The color feature F color is extracted through the KL divergence feature of the HSV space color histogram. The HSV space is a color model based on human vision, which can better reflect the intuitive characteristics of colors. The KL divergence is used to measure the difference in color distributions. The spectral feature F spectral is the reflectance ratio feature between the near-infrared band and the visible light band. This feature can reflect the difference in the reflectance characteristics of tobacco leaves under different spectra and is of great significance for the identification of pests and diseases. When calculating the feature weights, the attention mechanism is adopted, and through the learnable parameter α iand β i and the Sigmoid function σ to dynamically adjust the weights of the features of each modality, so as to highlight the important features and suppress the unimportant features. Finally, the fused feature tensor T is generated through weighted summation, where the modality matching matrix M i is used to adjust the matching relationship between the features of different modalities.
[0127] Preferably, when constructing the multi-modal feature vector, the parameter settings of the Gabor filter bank can be further optimized, such as increasing the number of directions and the frequency range, to capture texture features more comprehensively. For color features, normalization processing of the color space can be introduced to reduce the influence of lighting conditions on color features. In terms of spectral features, in addition to the reflectance ratio of the near-infrared and visible light bands, other spectral features can be introduced, such as the principal component analysis (PCA) features of multi-spectral images, to further enrich the spectral information. In the attention mechanism, a multi-head attention structure can be adopted, with the dimension of each head set to 64, and the attention weights are calculated through the query matrix Q, the key matrix K, and the value matrix V, so as to more effectively allocate weights to important features. In addition, an adaptive weight adjustment mechanism can be introduced to dynamically adjust the weights of the features of each modality according to the prior knowledge of the types of pests and diseases, so as to improve the pertinence and adaptability of the fused feature tensor.
[0128] In some embodiments, the calculation process of the adaptive weight classifier in step S5 includes:
[0129] S51. Construct a two-stream neural network, where the first branch processes spatial features and the second branch processes spectral features;
[0130] S52. Define a cross-adaptation function:
[0131]
[0132] where w1, w2 are dynamic weights and γ is a coupling coefficient;
[0133] S53. Optimize the weight parameters through backpropagation so that the loss function is minimized, and C is the number of pest and disease categories.
[0134] It should be noted that the adaptive weight classifier is a classification method that can dynamically adjust weights according to input features. Its purpose is to output the probability distribution of pests and diseases by performing a non-linear mapping on the fused feature tensor. The core of this method lies in constructing a two-stream neural network to process spatial features and spectral features respectively, and dynamically adjusting weights through a cross-adaptation function to improve the accuracy and adaptability of classification. By optimizing the weight parameters through backpropagation, the loss function is minimized, thereby achieving accurate identification of pests and diseases. This method is particularly suitable for the identification of tobacco leaf pests and diseases because the characteristics of tobacco leaf pests and diseases have complex variations in the spatial and spectral dimensions, and the adaptive weight classifier can better capture these variations and improve the identification effect.
[0135] Specifically, the implementation of the adaptive weight classifier includes the following key steps. First, construct a two-stream neural network, where the first branch is used to process spatial features and the second branch is used to process spectral features. Spatial features usually include information such as the texture and shape of the image, while spectral features reflect the reflection characteristics of tobacco leaves in different spectral bands. The design of the two-stream neural network allows for the simultaneous processing of these two types of features, thus more comprehensively describing the pest and disease situation of tobacco leaves.
[0136] Secondly, define the cross-adaptation function
[0137]
[0138] where x1 and x2 are the outputs of spatial features and spectral features respectively, W1 and w2 are dynamic weights, and γ is the coupling coefficient. This function adjusts the interaction between features through dynamic weights and the coupling coefficient, thereby improving the flexibility of classification. Finally, optimize the weight parameters through backpropagation to minimize the loss function where C is the number of pest and disease categories, y i is the true label, p i is the predicted probability, and λ is the regularization parameter. This optimization process can ensure that the classifier continuously adjusts the weights during training to better fit the data.
[0139] Preferably, when constructing the two-stream neural network, a Convolutional Neural Network (CNN) architecture can be adopted. Among them, the spatial feature branch can use multiple convolutional layers and pooling layers to extract local features of the image, while the spectral feature branch can use fully connected layers to process spectral data. For the cross-adaptation function, more interaction terms, such as high-order polynomial terms, can be introduced to more complexly describe the relationship between spatial features and spectral features. In addition, the Batch Normalization technique can be introduced to accelerate the training process and improve the stability of the model. During the optimization process, more advanced optimization algorithms, such as Adam or RMSprop, can be adopted. These algorithms can adaptively adjust the learning rate, thus converging to the optimal solution faster. In addition, the Early Stopping mechanism can also be introduced to stop training when the performance on the validation set no longer improves to prevent overfitting.
[0140] In some embodiments, the logical decision-making process in step S6 includes:
[0141] S61. Calculate the confidence index
[0142]
[0143] When C < 0.3, start the multi-model voting mechanism;
[0144] S62. Use the fuzzy logic rule base for decision correction, and define the membership function:
[0145]
[0146] S63. Perform decision tree traversal, and trigger manual review marking when the confidence of the leaf node is lower than the threshold.
[0147] It should be noted that the logical decision-making process is a key step in generating the pest and disease type identifier according to the probability distribution of pest and disease identification. This process evaluates the reliability of the identification result by calculating the confidence index, and starts the multi-model voting mechanism and the fuzzy logic rule base for decision correction when the confidence is low. Finally, the pest and disease type is determined through decision tree traversal, and manual review marking is triggered when necessary. This method can effectively improve the accuracy and robustness of decision-making and ensure the reliability of pest and disease identification results.
[0148] Specifically, in the logical decision-making process, first calculate the confidence index
[0149]
[0150] Among them, P is the probability distribution vector of pests and diseases, max(P) is the maximum probability value, mean(P) is the probability mean, and std(P) is the probability standard deviation. When the confidence level C < 0.3, the multi-model voting mechanism is activated to combine the prediction results of multiple pre-trained models to improve the reliability of decision-making. The fuzzy logic rule base processes uncertainty by defining the membership function μ(x). For example, when the probability value x falls within a certain interval, the membership function can assign different membership values to it, thereby realizing the quantification of uncertainty. The decision tree traversal gradually narrows down the range of pest and disease types through a series of predefined rule nodes, and finally determines the specific pest and disease types. If the confidence level of the leaf node is lower than the set threshold, an artificial review mark is triggered to ensure the accuracy of the final decision.
[0151] Preferably, when calculating the confidence index, the calculation formula of the confidence can be adjusted according to the prior knowledge of different pest and disease types. For example, category weights can be introduced to reflect the recognition difficulty of different pest and disease types. In the multi-model voting mechanism, a weighted voting method can be adopted, and different weights are assigned according to the performance of each model to improve the accuracy of the voting results. For the fuzzy logic rule base, more fuzzy rules and membership functions can be introduced to handle various uncertainty situations more carefully. For example, multiple fuzzy sets can be defined, such as low probability, medium probability, and high probability, and different membership functions can be designed for each set. During the decision tree traversal process, a deep learning model can be introduced as the decision basis for some nodes to improve the adaptability and accuracy of the decision tree. In addition, an online learning mechanism can be introduced to enable the system to dynamically adjust the decision rules according to new data, thereby further improving the robustness and adaptability of the system.
[0152] In some embodiments, the control instruction generation in step S7 includes:
[0153] S71. Establish a mapping relationship table between pest and disease types and application parameters;
[0154] S72. Calculate the application amount through the following formula:
[0155]
[0156] where S is the proportion of the infected area, D is the disease severity level, t is the duration, and k1, k2, and λ are adjustment coefficients;
[0157] S73. Generate a JSON control instruction package containing application coordinates, pesticide types, and dosage parameters.
[0158] It should be noted that control instruction generation is a process of triggering corresponding control instructions based on pest and disease type identifiers, which is used to adjust the operating parameters of associated agricultural equipment. The core of this process lies in establishing the mapping relationship between pest and disease types and application parameters, and generating specific control instructions by calculating the application amount. In this way, the system can automatically adjust the operating parameters of agricultural equipment according to the identified pest and disease types and severity levels, achieve precise pesticide application, and improve the efficiency and effectiveness of pest and disease control.
[0159] Specifically, during the control instruction generation process, first, a mapping relationship table between pest and disease types and application parameters is established. This mapping relationship table is a predefined database that stores the application parameters corresponding to different pest and disease types, including application concentration, application amount, application frequency, etc. Through the pest and disease type identifier, the system can look up the corresponding application parameters in the mapping relationship table. Secondly, the application amount is calculated based on the infection area ratio S, disease severity level D, and duration t. For example, the formula Dose = a1·S·D + a2·(1 - e -b·t ) can be used, where a1, a2, and b are adjustment coefficients used to adjust the relationship between the application amount and pest and disease parameters. Finally, a JSON control instruction package containing application coordinates, pesticide type, and dosage parameters is generated, and the control instructions are transmitted to the execution terminal of the agricultural equipment through the communication module to achieve automated pesticide application.
[0160] Preferably, when establishing the mapping relationship table between pest and disease types and application parameters, historical data and expert knowledge can be introduced to optimize the setting of application parameters. For example, according to past pest and disease control experience, adjust the application concentration and frequency of different pest and disease types. When calculating the application amount, adjustment coefficients of environmental factors such as temperature and humidity can be introduced to more accurately reflect the actual pesticide application requirements. In addition, machine learning algorithms can be used to dynamically optimize the application parameters and adjust the application amount according to real-time monitoring data to achieve more precise pest and disease control. When generating the control instruction package, a safety check mechanism can be introduced to ensure the accuracy and safety of the control instructions. For example, before sending the control instructions, the system can verify the instructions to ensure that they meet the preset safety standards.
[0161] In some embodiments, in step S31, the calculation of the gradient magnitude uses an improved Sobel operator:
[0162] Horizontal direction operator
[0163]
[0164] Vertical direction operator V = H T , Gradient magnitude
[0165] It should be noted that the improved Sobel operator is an algorithm for calculating the gradient magnitude of an image. It detects edges in the image by applying specific convolution kernels in the horizontal and vertical directions respectively. In the present invention, the improved Sobel operator in the horizontal direction is [-3, 0, 3; -10, 0, 10; -3, 0, 3], and the operator in the vertical direction is the transpose of the operator in the horizontal direction. Compared with the traditional Sobel operator, this improved operator can capture edge information in the image more effectively. Especially when processing multi-spectral images, it can better reflect the edge characteristics of tobacco leaf diseases and pests. By calculating the gradient magnitudes in the horizontal and vertical directions, the change situation of pixels in the image can be described more accurately, thus providing richer information for subsequent feature extraction.
[0166] Specifically, when calculating the gradient magnitude, the improved Sobel operator H in the horizontal direction and the operator V in the vertical direction are used to perform convolution operations on the image respectively. The operator H in the horizontal direction is [-3, 0, 3; -10, 0, 10; -3, 0, 3], and the operator V in the vertical direction is the transpose of H. By using these two operators to calculate the gradients of the image in the horizontal and vertical directions respectively, two gradient images G x and G y . Then, according to the formula
[0167]
[0168] the gradient magnitude of each pixel point is calculated, where * represents the convolution operation. This calculation method can effectively capture edge information in the image. Especially when processing multi-spectral images, it can better reflect the edge characteristics of tobacco leaf diseases and pests. The calculation results of the gradient magnitude can be used for subsequent feature extraction, such as the generation of dynamic convolution kernel sizes, etc.
[0169] Preferably, when using the improved Sobel operator, the weight parameters of the operator can be adjusted according to the resolution and noise level of the image. For example, for high-resolution images, the weight of the operator can be appropriately increased to more effectively capture subtle edge changes. In addition, multi-scale analysis can be introduced. By applying the improved Sobel operator at different scales, multi-scale gradient information is extracted, so as to better capture the edge characteristics in the image. For example, the image can be decomposed into multiple scales by combining wavelet transform, and then the gradient magnitudes are calculated separately at each scale. In addition, a non-linear filter, such as a median filter, can be introduced to post-process the calculated gradient magnitude image to further suppress the influence of noise. This multi-scale and non-linear processing method can improve the robustness and accuracy of gradient magnitude calculation, thus providing more reliable data support for subsequent feature extraction.
[0170] In some embodiments, the implementation of the attention mechanism in step S42 includes:
[0171] Using a multi - head attention structure, the dimension of each head is 64, and the query matrix Q = W q ·T, the key matrix K = W k ·T, the value matrix V = W v ·T, and the attention weights are calculated as:
[0172]
[0173] where d k is the dimension of the key vector.
[0174] It should be noted that the attention mechanism is a method for enhancing important parts in feature representation. By dynamically allocating weights to highlight key features, it improves the model's attention to important information. In the present invention, the attention mechanism is implemented through a multi - head attention structure, with the dimension of each head being 64. The query matrix, key matrix, and value matrix are respectively used to calculate the attention weights. This mechanism can effectively improve the efficiency and accuracy of feature fusion. Especially when dealing with multi - modal features, it can better capture the correlation between different features, thereby providing richer information for pest and disease identification.
[0175] Specifically, the core of the attention mechanism lies in calculating the attention weights through the query matrix Q, key matrix K, and value matrix V. Among them, the query matrix Q represents the features to be focused on, the key matrix K represents the context information of the features, and the value matrix V represents the actual feature values. The calculation formula for the attention weights is
[0176]
[0177] where d k is the dimension of the key vector. In the present invention, the dimension of each head is set to 64, which means that each head can process 64 - dimensional feature information. Through the multi - head attention structure, the model can capture the relationships between features from different perspectives, thereby improving the effect of feature fusion.
[0178] Preferably, when implementing the attention mechanism, the parameters of the multi-head attention structure can be adjusted according to actual application requirements. For example, the number of heads can be increased to further refine the capture of features, or the dimension of each head can be adjusted to adapt to feature data of different scales. In addition, residual connections and layer normalization techniques can be introduced to improve the training stability and performance of the model. Residual connections allow the model to pass the original input between different layers, thus avoiding the problem of vanishing gradients; layer normalization can normalize each feature, making the model more robust to the distribution of input data. In addition, an adaptive attention mechanism can be introduced to dynamically adjust the weights of each head by learning, so as to better adapt to different feature patterns. This adaptive mechanism can automatically adjust the allocation of attention weights according to the characteristics of the input data, further enhancing the flexibility and accuracy of feature fusion.
[0179] The above embodiments of the present invention have the following beneficial effects: First, by collecting multi-spectral images of tobacco leaves and preprocessing them, a standardized image matrix can be generated to ensure the consistency and reliability of the data. Based on the dynamic convolution kernel algorithm to extract local feature vectors, the subtle changes on the surface of tobacco leaves can be fully captured, improving the accuracy of feature extraction. The multi-modal feature fusion algorithm can effectively integrate different feature information to generate a fused feature tensor, thereby improving the accuracy and robustness of pest and disease identification. The adaptive weight classifier can perform non-linear mapping according to the probability distribution of pests and diseases and output more accurate pest and disease identification results.
[0180] In addition, through the logical decision-making mechanism, the present invention can generate pest and disease type identifiers according to the identification results and trigger corresponding control instructions to adjust the operating parameters of agricultural equipment. By adopting noise suppression, contrast enhancement and normalization processing, the quality of image data can be improved to ensure the accuracy of subsequent processing. The combination of the dynamic convolution kernel algorithm and the multi-modal feature fusion technology can significantly improve the effect of pest and disease identification. Through the multi-model voting mechanism and the fuzzy logic rule base, decision correction can be carried out in the case of low confidence to ensure the reliability of the identification results. Finally, the generated control instructions can achieve precise pesticide application, reduce the amount of pesticide used, lower the agricultural production cost, improve the quality of tobacco leaves and the intelligent level of agricultural production.
[0181] As Figure 2 shown, a tobacco leaf pest and disease identification system 200 of some embodiments, the system 200 includes:
[0182] An image acquisition module 201, configured to acquire tobacco leaf images and obtain multi-spectral reflection data on the surface of tobacco leaves through an image sensor;
[0183] The processing module 202 is configured to preprocess the multi-spectral reflection data to generate a normalized image matrix; extract local feature vectors of the normalized image matrix based on the dynamic convolution kernel algorithm; match the local feature vectors with a preset pest and disease spectral feature library through a multi-modal feature fusion algorithm to generate a fusion feature tensor; perform a non-linear mapping on the fusion feature tensor using an adaptive weight classifier to output a pest and disease probability distribution; and execute a logical decision based on the probability distribution to generate a pest and disease type identifier.
[0184] The control module 203 is configured to trigger a control instruction according to the type identifier and adjust the operating parameters of the associated agricultural equipment.
[0185] The communication module 204 is configured to transmit the control instruction to the agricultural equipment execution terminal. It can be understood that the modules described in the tobacco leaf pest and disease identification system 200 correspond to the respective steps in the tobacco leaf pest and disease identification method described in the reference. Figure 1 Therefore, the operations, features, and beneficial effects described above for the tobacco leaf pest and disease identification method also apply to the tobacco leaf pest and disease identification system 200 and the modules included therein, and will not be elaborated herein.
[0186] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present invention. The electronic devices in some embodiments of the present invention may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The terminal device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present invention.
[0187] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0188] Typically, the following devices can be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 can allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had. Figure 3 Each block shown in can represent one device or, as needed, multiple devices.
[0189] Further, the storage medium of the embodiments of the present application stores program instructions capable of implementing all the above methods. Among them, the program instructions can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. And the foregoing storage medium includes: various media that can store program codes such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a terminal device such as a computer, a server, a mobile phone, a tablet, etc.
[0190] The above description is only some preferred embodiments of the present invention and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present invention.
Claims
1. A method for identifying tobacco leaf diseases and pests, characterized in that, It includes the following steps: S1. Collect tobacco leaf images and obtain multi - spectral reflection data on the surface of tobacco leaves through an image sensor; S2. Pre - process the multi - spectral reflection data to generate a standardized image matrix; S3. Extract local feature vectors of the standardized image matrix based on the dynamic convolution kernel algorithm; S4. Through the multi - modal feature fusion algorithm, match the local feature vectors with a preset spectral feature library of pests and diseases to generate a fused feature tensor; S5. Use an adaptive weight classifier to perform non - linear mapping on the fused feature tensor and output the probability distribution of pests and diseases; S6. Execute logical decision - making based on the probability distribution to generate a pest and disease type identifier; S7. Trigger a control instruction according to the type identifier to adjust the operating parameters of associated agricultural equipment.
2. The method according to claim 1, characterized in that, The step S2 includes: S21. Perform noise suppression on the multi-spectral reflection data and calculate the neighborhood gradient variance σ of each pixel point 2 , when σ 2 > the threshold T, update the pixel value using the following correction formula: where I(x, y) is the original pixel value, is the Gaussian weight kernel, and k is the window radius; S22. Enhance the contrast of the denoised data and adjust the pixel dynamic range through a piece - wise linear transformation function; S23. Normalize the enhanced data to the interval [0, 1] to generate the standardized image matrix.
3. The method according to claim 1, wherein The dynamic convolution kernel algorithm in the step S3 includes: S31. Dynamically generate the convolution kernel size according to the gradient magnitude of the local region of the image; among them, the dynamic generation of the convolution kernel size is shown in the following formula: The dynamic generation of the convolution kernel size is shown in the following formula: S32. Construct a deformable convolution kernel parameter group θ = {δ x , δ y , α}, where the offsets δ x , δ y are determined by the eigenvectors of the local curvature matrix, and the scaling factor S33. Calculate the feature response map through the following formula: Among them, W(i, j) is a trainable weight matrix.
4. The method according to claim 1, wherein The step S4 includes: S41. Construct a multi-modal feature vector V = [v texture , v color , v spectral , where: v texture is the texture feature extracted by the Gabor filter bank; v color is the KL divergence feature of the HSV color space histogram; v spectral is the reflectance ratio feature between the near-infrared band and the visible light band; S42. Calculate the feature weights using the attention mechanism; Among them, W m , b m are learnable parameters, and σ is the Sigmoid function; S43. Generate the fused feature tensor through weighted summation M m is a preset modal matching matrix.
5. The method according to claim 1, wherein The calculation process of the adaptive weight classifier in the step S5 includes: S51. Construct a two - stream neural network, where the first branch processes spatial features and the second branch processes spectral features; S52. Define the cross - adaptation function: Among them, w1, w2 are dynamic weights, and γ is the coupling coefficient; S53. Optimize the weight parameters through backpropagation to minimize the loss function , where C is the number of pest and disease categories.
6. The method according to claim 1, wherein The logical decision - making process in the step S6 includes: S61. Calculate the confidence index: Start the multi - model voting mechanism when C < 0.3; S62. Use a fuzzy logic rule base for decision correction and define the membership function; S63. Execute decision - tree traversal and trigger manual review marking when the leaf - node confidence is lower than the threshold.
7. The method according to claim 1, characterized in that The generation of the control instruction in the step S7 includes: S71. Establish a mapping relationship table between pest and disease types and application parameters; S72. Calculate the application amount through the following formula: Among them, S is the proportion of the infected area, D is the disease severity level, t is the duration, and k1, k2, λ are adjustment coefficients; S73. Generate a JSON control instruction package containing application coordinates, pesticide types, and dosage parameters.
8. The method according to claim 3, wherein The calculation of the gradient magnitude in the step S31 uses an improved Sobel operator: Horizontal direction operator Vertical direction operator V = H T , gradient magnitude 9. The method according to claim 4, wherein The implementation of the attention mechanism in the step S42 includes: Using a multi-head attention structure, with the dimension of each head being 64, the query matrix Q = W q ·T, the key matrix K = W k ·T, the value matrix V = W v ·T, and the attention weights are calculated as: Among them, d k is the dimension of the key vector.
10. A tobacco leaf pest and disease identification system, characterized in that, Includes: An image acquisition module for collecting tobacco leaf images and obtaining multi - spectral reflection data on the surface of tobacco leaves through an image sensor; A processing module, configured to preprocess the multi-spectral reflection data to generate a normalized image matrix; extract local feature vectors of the normalized image matrix based on a dynamic convolution kernel algorithm; match the local feature vectors with a preset pest and disease spectral feature library through a multi-modal feature fusion algorithm to generate a fused feature tensor; perform a non-linear mapping on the fused feature tensor using an adaptive weight classifier to output a pest and disease probability distribution; execute a logical decision based on the probability distribution to generate a pest and disease type identifier; A control module, configured to trigger a control instruction according to the type identifier and adjust the operating parameters of associated agricultural equipment; A communication module, configured to transmit the control instruction to an agricultural equipment execution terminal.
Citation Information
Patent Citations
Intelligent identification method, medium and system for tobacco plant diseases and insect pests
CN117372881A
Tobacco disease and insect pest identification method, medium and system
CN118072251A
Digitized monitoring and early warning method and system for plant diseases and insect pests
CN118334582A
Method and device for judging health condition of tobacco
CN118429790A
Tobacco disease and insect pest pesticide matching method and system
CN119180334A
Cited By
Visual inspection system and method for tiny flaws of industrial products
CN120612327A