Crop disease detection method and system based on knowledge distillation
By adopting knowledge distillation-based methods and attention-guiding mechanisms in crop disease detection, a lightweight disease detection model is built, which solves the problems of excessive resource demand and insufficient discrimination caused by the complexity of the existing model, and achieves efficient and accurate disease detection, which is suitable for large-scale agricultural production.
Patent Information
- Application Number
- CN202510126233.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-06-20
AI Technical Summary
The existing general target detection model has high complexity and high demand for computing power, storage and energy, which exceeds the carrying range of existing agricultural equipment, limiting its deployment and application in the actual agricultural production environment. At the same time, the model lacks discrimination ability when facing crop diseases with similar characteristics, making it difficult to achieve accurate identification, and overfitting is prone to occur when the scale of the training data set is limited. In addition, the manual detection method relies on the subjective experience of farmers, resulting in poor consistency and reliability of the detection results, and is time-consuming and labor-consuming, making it difficult to be applicable to large-scale agricultural production.
Using knowledge distillation-based crop disease detection method, a lightweight disease detection model is constructed by extracting key knowledge of crop disease detection from large-scale teacher models and introducing attention-guiding mechanisms. The method includes image enhancement, data preprocessing, knowledge distillation model training and attention-guiding mechanism application, aiming to reduce the computational complexity and dependence of the model while improving the accuracy and robustness of the detection.
It effectively solves the problem of excessive resource demand caused by the complexity of existing models, reduces the number of parameters and calculations of the model, and improves the deployment efficiency in the agricultural production environment; at the same time, the model's ability to identify disease characteristics through attention guidance mechanism is enhanced, and the accuracy and robustness of detection is improved. In addition, the time and labor costs of data preparation and model training are reduced, making the method suitable for large-scale agricultural production.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_4
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and image processing in artificial intelligence, and more specifically, relates to a crop disease detection method and system based on knowledge distillation. Background Art
[0002] In modern agricultural economy, crops such as tomatoes, beans, and strawberries play a crucial role. They not only significantly increase farmers' economic income but also promote the diversified development of rural economy. However, these crops are vulnerable to diseases such as powdery mildew, leaf spot, and anthracnose during their growth cycle, which pose a serious threat to the yield and quality of the crops. Therefore, carrying out timely and accurate crop disease detection is of great significance in ensuring the healthy growth, stable yield, and improved quality of crops. It is not only an essential key link in agricultural production but also has a profound impact on promoting agricultural modernization, enhancing agricultural competitiveness, and achieving sustainable development.
[0003] Currently, crop disease detection mainly relies on two methods. One is manual detection, which is a traditional method that requires farmers to go to the fields in various harsh weather conditions, such as sweltering heat or severe cold, and rely on their own planting experience and naked-eye observation to identify diseases. The other is an automated detection method based on a general object detection model, which inputs crop images and uses a trained deep neural network model to automatically identify the specific location and type of diseases in the images, thereby achieving automated disease detection.
[0004] However, the above two existing crop disease detection methods still have some unavoidable defects:
[0005] (1) Due to the high complexity of the existing general object detection model, it has high requirements for computing power, storage, and energy, exceeding the capacity of existing agricultural equipment. Therefore, it limits its deployment and application in the actual agricultural production environment.
[0006] (2) The general object detection model uses convolution and pooling operations to extract common features, but when facing crop diseases with similar features, its discriminative ability is insufficient, making it difficult to accurately identify crop diseases.
[0007] (3) For the general object detection model, when the scale of the training dataset is limited, it is prone to overfitting, which will weaken its generalization ability for unknown data.
[0008] (4) The existing manual detection method overly relies on farmers' subjective experience, resulting in poor consistency and reliability of disease detection results. At the same time, this method consumes a large amount of time and manpower and is difficult to apply to large-scale agricultural production. Summary of the Invention
[0009] In view of the above defects or improvement requirements of the prior art, the present invention provides a method and system for crop disease detection based on knowledge distillation, aiming to solve the technical problems that the existing general object detection model has high complexity, large computational power, storage and energy requirements, exceeding the carrying capacity of existing agricultural equipment, thus limiting its deployment and application in the actual agricultural production environment, and when facing crop diseases with similar features, its discrimination ability is insufficient, making it difficult to achieve accurate identification of crop diseases, and when the scale of the training data set is limited, overfitting is likely to occur, thus weakening its generalization ability for unknown data, and the existing manual detection method overly relies on the subjective experience of farmers, resulting in poor consistency and reliability of disease detection results, and this method consumes a large amount of time and manpower and is difficult to be applied to large-scale agricultural production.
[0010] To achieve the above object, according to one aspect of the present invention, a method for crop disease detection based on knowledge distillation is provided, including the following steps:
[0011] (1) Obtain the crop image to be detected.
[0012] (2) Perform image enhancement processing on the crop image to be detected obtained in step (1) to obtain the enhanced image.
[0013] (3) Perform data preprocessing on the enhanced image obtained in step (2) to obtain the preprocessed image.
[0014] (4) Input the preprocessed image obtained in step (3) into a pre-trained crop disease detection model based on knowledge distillation to obtain the final detection result.
[0015] Preferably, any one or any combination of any of the following 8 methods is used for the data enhancement processing in step (2): HSV enhancement of the image, including hue, saturation and brightness enhancement, and the enhancement factors are set to 0.015, 0.7 and 0.4 respectively; image translation, and the translation factor is 0.2; image scaling, and the scaling factor is 0.9; image shearing, and the shearing factor is 0.5; image perspective, and the perspective factor is 0.001; left-right flipping of the image, and the flipping factor is 0.5; multi-image splicing, and the splicing factor is 1.0; image mixing, and the mixing factor is 0.15;
[0016] Preferably, the detection result obtained in step (4) exists in the form of a detection box, and each detection box marks the predicted disease position and disease category.
[0017] The crop disease detection model based on knowledge distillation in step (4) adopts a knowledge distillation model, which includes a teacher model NT and a student model N S 。
[0018] Preferably, the crop disease detection model based on knowledge distillation is trained through the following steps:
[0019] (5-1) Download the mixed dataset composed of the open-source strawberry dataset, tomato dataset, and pod dataset, and divide the dataset into a training set and a test set according to a ratio of 8:2.
[0020] (5-2) Perform image enhancement processing on the training set obtained in step (5-1) to obtain the enhanced training set.
[0021] (5-3) Perform data preprocessing on the enhanced training set obtained in step (5-2) to obtain the preprocessed training set.
[0022] (5-4) For each sample in the preprocessed training set obtained in step (5-3), input the sample into the teacher model N T for training, and after the training is completed, freeze the parameters of the teacher model N T .
[0023] (5-5) Set the counter cnt = 1, and initialize the total loss L cnt of the first iteration of the disease detection model as a random number between 0 and 1;
[0024] (5-6) Determine whether cnt is greater than the preset iteration threshold. If so, go to step (5-19); otherwise, go to step (5-7);
[0025] (5-7) For each sample in the preprocessed training set obtained in step (5-3), input the sample into the trained and parameter-frozen teacher model N T and the student model N S for forward propagation to obtain the intermediate layer features F T,cnt and F S,cnt .
[0026] (5-8) For each sample obtained in step (5-3), obtain the spatial attention of the intermediate layer feature F T,cnt corresponding to the sample obtained in step (5-7) and of the cnt-th iteration;
[0027] (5-9) For each sample obtained in step (5-3), obtain the spatial attention of the intermediate layer feature F S,cnt corresponding to the sample obtained in step (5-7) and of the cnt-th iteration;
[0028] (5 - 10) The intermediate - layer feature F obtained according to step (5 - 8) T,cnt 's spatial attention S(F T,cnt ) and the intermediate - layer feature F obtained in step (5 - 9) S,cnt 's spatial attention S(F S,cnt ) to obtain the spatial - attention weight W corresponding to this sample and for the cnt - th iteration sp,cnt (The purpose is to use the spatial attention of the teacher model to guide the student model to better learn spatial features).
[0029] (5 - 11) For each sample obtained in step (5 - 3), obtain the channel attention of the intermediate - layer feature F corresponding to this sample and for the cnt - th iteration obtained in step (5 - 7) T,cnt ;
[0030] (5 - 12) For each sample obtained in step (5 - 3), obtain the channel attention of the intermediate - layer feature F corresponding to this sample and for the cnt - th iteration obtained in step (5 - 7) S,cnt ;
[0031] (5 - 13) For each sample obtained in step (5 - 3), according to the channel attention C(F T,cnt ) of the intermediate - layer feature F obtained in step (5 - 11) and the channel attention C(F T,cnt ) of the intermediate - layer feature F obtained in step (5 - 12), obtain the channel - attention weight W corresponding to this sample and for the cnt - th iteration S,cnt in order to use the channel attention of the teacher model to guide the student model to better learn channel features. T,cnt ch,cnt
[0032] (5 - 14) For each sample obtained in step (5 - 3), calculate the total loss L of the disease - detection model cnt with respect to the gradient of the intermediate - layer feature F of the teacher model T,cnt .
[0033] (5 - 15) For each sample obtained in step (5 - 3), perform min - max normalization on the gradient obtained in step (5 - 14), that is, normalize the gradient value to between 0 and 1, and the resulting value is the pixel - point attention weight W T,cnt corresponding to the intermediate - layer feature F of this sample and for the cnt - th iteration pi,cnt .
[0034] (5-16) For each sample obtained in step (5-3), according to the spatial attention weight W of the cnt-th iteration corresponding to this sample obtained in step (5-10) sp,cnt , the channel attention weight W obtained in step (5-13) ch,cnt , and the pixel point attention weight W obtained in step (5-15) pi,cnt calculate the attention loss L of the cnt-th iteration corresponding to this sample AT,cnt and the loss L of the student model OR,cnt .
[0035] (5-17) For each sample obtained in step (5-3), according to the attention loss L of the cnt-th iteration corresponding to this sample calculated in step (5-16) AT,cnt and the loss L of the student model OR,cnt obtain the total loss L of the disease detection model of the cnt-th iteration corresponding to this sample cnt = αL AT,cnt + βL OR,cnt , where α and β represent loss weights, and the value ranges of both are between 0 and 1. Preferably, both are equal to 1.
[0036] (5-18) For each sample obtained in step (5-3), use the total loss L of the crop disease detection model of the cnt-th iteration obtained in step (5-17) cnt to perform gradient backpropagation on the student model N S and update the parameters of the student model N S through the stochastic gradient descent algorithm SGD so that the total loss L cnt is minimized, set the counter cnt = cnt + 1, and return to step (5-6).
[0037] (5-19) Obtain the optimal parameters of the student model N S to obtain a preliminarily trained crop disease detection model based on knowledge distillation.
[0038] (5-20) Use the test set obtained in step (5-1) to test the preliminarily trained crop disease detection model based on knowledge distillation in step (5-19) until the obtained classification accuracy reaches the optimum, so as to obtain a finally trained crop disease detection model.
[0039] Preferably, step (5-8) is specifically as follows. First, obtain the height H T of the intermediate layer feature F T,cnt of the teacher model N T,cnt , the width W T,cnt and C T,cntchannels; then, in the first branch, first apply a convolution operation to the intermediate layer feature F T,cnt with a kernel size of 1×9, and then apply another convolution operation to the output of the first convolution operation with a kernel size of 9×1, and finally output the feature map S1; subsequently, in the second branch, first apply a convolution operation to the intermediate layer feature F T,cnt with a kernel size of 9×1, and then apply another convolution operation to the output of the first convolution operation with a kernel size of 1×9, and finally output the feature map S2; then, perform an element-wise addition of the feature map S1 output by the first branch and the feature map S2 output by the second branch to obtain a fused feature map; finally, apply the Sigmoid activation function to the fused feature map to normalize the output value between 0 and 1 to obtain the spatial attention S(F T,cnt ) of the intermediate layer feature F T,cnt .
[0040] Specifically, step (5-9) is as follows: First, obtain the intermediate layer feature F S of the student model N S,cnt with height H S,cnt , width W S,cnt and C S,cnt channels, and then, using the same method as the teacher model, perform convolution operations on the intermediate layer feature F S,cnt in two parallel branches to obtain the feature maps S1 and S2 respectively. Then, perform an element-wise addition of the feature maps S1 and S2 output by the two branches to obtain a fused feature map. Finally, apply the Sigmoid activation function to the fused feature map to normalize the output value between 0 and 1 to obtain the spatial attention S(F S,cnt ) of the intermediate layer feature F S,cnt .
[0041] Preferably, step (5-10) is specifically: Calculate the absolute difference between the spatial attentions S(F T,cnt ) and S(F S,cnt ) of the intermediate layer features of the teacher model and the student model, divide it by the scaling factor τ (the value range of which is (0,1], preferably 0.5), and then apply the Softmax function to the result of the above division operation to obtain the spatial attention weight W sp,cnt :
[0042] W sp,cnt =H S,cnt ×W S,cnt ×Softmax(|S(F T,cnt ) - S(F S,cnt )| / τ)
[0043] Specifically, step (5-11) is as follows: First, obtain the teacher model NT The intermediate layer feature F T,cnt The height H T,cnt and the width W T,cnt and C T,cnt channels; then, perform global average pooling on the intermediate layer feature F T,cnt to obtain a feature map with dimensions 1×1×C T,cnt ; subsequently, input this feature map into the first fully connected layer (the number of neurons in this layer is C T,cnt / 4) to obtain a feature map with dimensions 1×1×(C T,cnt / 4); then, apply the ReLU activation function to the feature map output by the first fully connected layer, and input the output after ReLU activation into the second fully connected layer (the number of neurons in this layer is C T,cnt ) to obtain a feature map with dimensions 1×1×C T,cnt ; finally, apply the Sigmoid activation function to the feature map output by the second fully connected layer, normalize the output value between 0 and 1, and the resulting value is the channel attention C(F T,cnt ) of the teacher model.
[0044] Step (5-12) is specifically as follows. First, obtain the intermediate layer feature F S of the student model N S,cnt The height H S,cnt and the width W S,cnt and C S,cnt channels; then, perform global average pooling on the intermediate layer feature F S,cnt to obtain a feature map with dimensions 1×1×C S,cnt ; then, input this feature map into the first fully connected layer, the number of neurons in this layer is C S,cnt / 4 to obtain a feature map with dimensions 1×1×(C S,cnt / 4); then, apply the ReLU activation function to the feature map output by the first fully connected layer, and input the output after ReLU activation into the second fully connected layer, the number of neurons in this layer is C S,cnt to obtain a feature map with dimensions 1×1×C S,cnt ; finally, apply the Sigmoid activation function to the feature map output by the second fully connected layer, normalize the output value between 0 and 1, and the resulting value is the channel attention C(F S,cnt ) of the student model.
[0045] Preferably, step (5-13) is specifically as follows. Calculate the channel attention C(F T,cnt ) and C(F S,cntThe absolute difference between them is divided by the scaling factor, and then the Softmax function is applied to the result of the above division operation to obtain the channel attention weight W ch,cnt :
[0046] W ch,cnt = C S,cnt × Softmax(|C(F T,cnt ) ― C(F S,cnt )| / τ)
[0047] Step (5-15) adopts the following formula:
[0048]
[0049] Among them, Norm represents min-max normalization.
[0050] Preferably, in step (5-16), the attention loss L AT,cnt The calculation formula is:
[0051]
[0052] Among them, the function f(·) represents performing a convolution operation on the intermediate layer features of the aligned teacher model and the intermediate layer features of the student model.
[0053] The loss L OR,cnt of the student model has the following calculation formula:
[0054] L OR,cnt = γ1·L obj + γ2·L cls + γ3·L box
[0055] Among them, the value range of γ1 is between 0 and 1, preferably 0.7, the value range of γ2 is between 0 and 1, preferably 0.3, and the value range of γ3 is between 0 and 1, preferably 0.05. All three are weight coefficients, and L obj,cnt represents the target loss of the cnt-th iteration, L cls,cnt represents the classification loss of the cnt-th iteration, and L box,cnt represents the bounding box regression loss of the cnt-th iteration.
[0056] Preferably, the target loss L obj,cnt is equal to the binary cross-entropy loss between the target value predicted by this sample and the true target value y obj :
[0057]
[0058] The classification loss L cls,cntEqual to the classification value predicted for the sample and the true classification value y cls The binary cross-entropy loss between them is:
[0059]
[0060] The bounding box regression loss L box,cnt is equal to the predicted bounding box coordinates of the sample and the true coordinates y box The complete intersection over union loss between them is:
[0061]
[0062] where BCE(·) represents the binary cross-entropy loss function and CIoU(·) represents the complete intersection over union loss function.
[0063] According to another aspect of the present invention, there is provided a crop disease detection system based on knowledge distillation, including:
[0064] The first module is used to obtain the crop image to be detected.
[0065] The second module is used to perform image enhancement processing on the crop image to be detected obtained by the first module to obtain the enhanced image.
[0066] The third module is used to perform data preprocessing on the enhanced image obtained by the second module to obtain the preprocessed image.
[0067] The fourth module is used to input the preprocessed image obtained by the third module into a pre-trained crop disease detection model based on knowledge distillation to obtain the final detection result.
[0068] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0069] (1) Through steps (5-1) to (5-20), the present invention adopts the knowledge distillation method to extract the key knowledge of crop disease detection from a large-scale teacher model, and introduces an attention guidance mechanism to strengthen the attention of the student model to the key features of diseases, and constructs a lightweight disease detection model. This effectively solves the technical problem that the existing general object detection model has high complexity, large computational power, storage and energy requirements, exceeding the carrying range of existing agricultural equipment, thus restricting its deployment and application in the actual agricultural production environment;
[0070] (2) Since the present invention adopts steps (5-8) to steps (5-15), it designs a pixel attention, spatial and channel attention guidance mechanism. By calculating the gradient of each pixel in the intermediate layer feature map of the teacher model with respect to the total loss, its contribution to the total loss is clarified, thereby identifying the key features of the disease. Using these contribution values as weights, it guides the student model to learn the teacher model's attention to the significant features of the disease and amplifies the influence of important features. This effectively solves the technical problem that the general object detection model has insufficient discrimination and is difficult to achieve accurate recognition when facing crop diseases with similar features because it relies on convolutional and pooling operations to extract common features;
[0071] (3) Through steps (5-1) and (5-2), the present invention adopts an open-source dataset and data augmentation technology, significantly reducing the time and labor costs of making new data samples, while enriching the dataset to be detected. This effectively solves the overfitting problem that easily occurs when the scale of the training dataset of the general object detection model is limited, thereby significantly improving the generalization ability of the model to unknown data;
[0072] (4) Since the present invention adopts steps (5-1) to (5-20), it designs a crop disease detection method based on knowledge distillation, effectively breaking through the bottleneck of low efficiency and insufficient accuracy of traditional manual detection, realizing the automatic detection of crop diseases, and thus being able to solve the technical problem that the existing manual detection method overly relies on the subjective experience of farmers, resulting in poor consistency and reliability of disease detection results. In addition, the present invention greatly reduces the time and labor costs required for detection, enabling it to be efficiently applied to large-scale agricultural production, and having significant practical value and promotion prospects. Description of the Drawings
[0073] Figure 1 is a schematic diagram of the crop disease detection method based on knowledge distillation used in the present invention;
[0074] Figure 2 is a schematic diagram of the training process of the disease detection method used in the present invention;
[0075] Figure 3 is a schematic diagram of the attention guidance mechanism used in the present invention;
[0076] Figure 4 is a schematic diagram of the comparison of the strawberry disease recognition results between the present invention and the existing method;
[0077] Figure 5 is a schematic diagram of the comparison of the tomato recognition results between the present invention and the existing method;
[0078] Figure 6 is a schematic diagram of the comparison of the bean disease recognition results between the present invention and the existing method. Detailed implementation manners
[0079] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0080] For the convenience of explaining the model training and inference processes, the present invention uses YOLOv7 as the teacher model N T , and uses YOLOv7-tiny as the student model N S for illustration. It should be clear that the specific models selected here are only for illustration to explain the technical idea of the present invention. The teacher model and the student model can use other suitable models, and the present invention is not limited to the above examples.
[0081] As Figure 1 shown, the present invention provides a method for detecting crop diseases based on knowledge distillation, including the following steps:
[0082] (1) Obtain the crop image to be detected.
[0083] (2) Perform image enhancement processing on the crop image to be detected obtained in step (1) to obtain an enhanced image.
[0084] Specifically, for any crop image to be detected, one or any combination of the following 8 data enhancement methods can be randomly selected for processing: image HSV enhancement, including hue, saturation, and brightness enhancement, with the enhancement factors set to 0.015, 0.7, and 0.4 respectively; image translation, with the translation factor of 0.2; image scaling, with the scaling factor of 0.9; image shearing, with the shearing factor of 0.5; image perspective, with the perspective factor of 0.001; image left-right flipping, with the flipping factor of 0.5; multi-image stitching, with the stitching factor of 1.0; image mixing, with the mixing factor of 0.15.
[0085] The advantage of this step is that it significantly reduces the time and labor costs for making new data samples, while enriching the dataset to be detected. This not only improves the efficiency of data preparation, but also enhances the robustness of the crop disease detection model based on knowledge distillation, making it show higher accuracy and stability in the face of complex and diverse disease samples.
[0086] (3) Perform data preprocessing on the enhanced image obtained in step (2) to obtain a preprocessed image.
[0087] Specifically, first, resize the enhanced image to a resolution of 640×640; then, normalize the pixel values of the resized image from the range [0, 255] to the range [0, 1] to obtain the preprocessed image.
[0088] (4) Input the preprocessed image obtained in step (3) into a pre-trained crop disease detection model based on knowledge distillation to obtain the final detection result.
[0089] Specifically, the detection result obtained in this step exists in the form of detection boxes, and each detection box marks the predicted disease location and disease category.
[0090] The crop disease detection model based on knowledge distillation of the present invention uses a knowledge distillation model, which includes a teacher model N T and a student model N S . For ease of explanation, YOLOv7 is used as the teacher model N T , and YOLOv7-tiny is used as the student model N S for illustration. The model structures of YOLOv7 and YOLOv7-tiny are detailed in the paper "YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors" published by the team of the Institute of Information Science, Academia Sinica, Taiwan, China in the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR) in 2023.
[0091] As Figure 2 shown, the crop disease detection model based on knowledge distillation of the present invention is trained through the following steps:
[0092] (5-1) Download a mixed dataset composed of open-source strawberry datasets, tomato datasets, and pod datasets, and divide the dataset into a training set and a test set according to a ratio of 8:2.
[0093] The advantages of this sub-step are as follows. First, the open-source dataset is carefully annotated and organized, covering a rich variety of agricultural scenarios and target categories, which can effectively improve the generalization ability and accuracy of the model. Second, optimizing the model and comparing performance based on the open-source dataset can effectively accelerate technological innovation and algorithm improvement. In addition, the open-source dataset has a low acquisition cost, and there is no need to invest a large amount of additional resources in data collection and annotation. Finally, the diversity and scale advantages of the open-source dataset help the model to be fully trained in different scenarios, enhance the detection ability for complex environments and rare diseases, and further promote the application of crop disease target detection technology in actual agricultural scenarios.
[0094] (5-2) Perform image enhancement processing on the training set obtained in step (5-1) to obtain the enhanced training set.
[0095] Specifically, the enhancement processing process in this step is exactly the same as step (2) above and will not be elaborated here.
[0096] (5-3) Perform data preprocessing on the enhanced training set obtained in step (5-2) to obtain the preprocessed training set.
[0097] Specifically, the data preprocessing process in this step is exactly the same as step (3) above and will not be elaborated here.
[0098] (5-4) For each sample in the preprocessed training set obtained in step (5-3), input the sample into the teacher model N T for training, and after the training is completed, freeze the parameters of the teacher model N T .
[0099] (5-5) Set the counter cnt = 1, and initialize the total loss L cnt of the first iteration of the disease detection model to a random number between 0 and 1;
[0100] (5-6) Determine whether cnt is greater than the preset iteration threshold. If so, go to step (5-19); otherwise, go to step (5-7).
[0101] (5-7) For each sample in the preprocessed training set obtained in step (5-3), input the sample into the trained and parameter-frozen teacher model N T and the student model N S for forward propagation to obtain the intermediate layer features F T,cnt and F S,cnt corresponding to the sample for the cnt-th iteration.
[0102] (5 - 8) For each sample obtained in step (5 - 3), obtain the spatial attention of the intermediate - layer feature F of the cnt - th iteration corresponding to this sample obtained in step (5 - 7). T,cnt of the spatial attention;
[0103] Specifically, in this step, first, obtain the intermediate - layer feature F of the teacher model N T ; the height H T,cnt of the intermediate - layer feature F T,cnt , width W T,cnt and C T,cnt channels; then, in the first branch, first apply a convolution operation to the intermediate - layer feature F T,cnt with a kernel size of 1×9, and then apply another convolution operation to the output of the first convolution operation with a kernel size of 9×1, and finally output the feature map S1; subsequently, in the second branch, first apply a convolution operation to the intermediate - layer feature F T,cnt with a kernel size of 9×1, and then apply another convolution operation to the output of the first convolution operation with a kernel size of 1×9, and finally output the feature map S2; then, element - wise add the feature map S1 output by the first branch and the feature map S2 output by the second branch to obtain the fused feature map; finally, apply the Sigmoid activation function to the fused feature map to normalize the output value between 0 and 1 to obtain the spatial attention S(F T,cnt ) of the intermediate - layer feature F T,cnt .
[0104] (5 - 9) For each sample obtained in step (5 - 3), obtain the spatial attention of the intermediate - layer feature F of the cnt - th iteration corresponding to this sample obtained in step (5 - 7). S,cnt of the spatial attention;
[0105] Specifically, in this step, first, obtain the intermediate - layer feature F of the student model N S ; the height H S,cnt of the intermediate - layer feature F S,cnt , width W S,cnt and C S,cnt channels, and then, use the same method as the teacher model to perform convolution operations on the intermediate - layer feature F S,cnt in two parallel branches to obtain the feature maps S1 and S2 respectively. Then, element - wise add the feature maps S1 and S2 output by the two branches to obtain the fused feature map. Finally, apply the Sigmoid activation function to the fused feature map to normalize the output value between 0 and 1 to obtain the spatial attention S(F S,cnt ) of the intermediate - layer feature F S,cnt .
[0106] The advantages of the above sub-steps (5-8) and (5-9) are that, first, two parallel branches are designed to extract the intermediate layer features F T,cnt and F S,cnt Each branch only uses continuous convolution operations, which is simple and has low parameters. Secondly, the spatial attention mechanism can dynamically adjust the model's attention to the intermediate layer features F T,cnt and F S,cnt The attention to different spatial areas effectively suppresses background interference, accurately locates the key characteristic areas of crop diseases, and significantly improves the accuracy and efficiency of detection; finally, the spatial attention mechanism also enhances the model's perception of small disease targets, further improving the model's robustness in complex agricultural environments.
[0107] (5-10) The intermediate layer feature F obtained according to step (5-8) T,cnt The spatial attention S(F T,cnt ) and the intermediate layer features F obtained in steps (5-9) S,cnt The spatial attention S(F S,cnt ) Get the spatial attention weight W corresponding to the sample and the cnt-th iteration sp,cnt (The purpose is to use the spatial attention of the teacher model to guide the student model to better learn spatial features).
[0108] Specifically, this step is to calculate the spatial attention S(F T,cnt ) and S(F S,cnt ), and divided by the scaling factor τ (its value range is (0,1], preferably 0.5), and then the Softmax function is applied to the above division result to obtain the spatial attention weight W sp,cnt :
[0109] W sp,cnt =H S,cnt ×W S,cnt ×Softmax(|S(F T,cnt )―S(F S,cnt )| / τ)
[0110] The advantages of this step are as follows. First, by using the spatial attention of the teacher model to guide the student model to learn, it can efficiently transmit spatial feature information and significantly improve the student model's perception ability of spatial features. Then, by calculating the difference between the spatial attention of the teacher and student models and combining with the Softmax function, it can dynamically adjust the weight allocation, enabling the student model to more accurately learn key spatial features. Secondly, this method effectively enhances the robustness of the student model in complex scenarios, making it show higher stability when facing background interference or target occlusion. Finally, this step is simple to operate, has a low computational complexity, does not significantly increase the training cost of the model, and can significantly improve the detection performance of the model.
[0111] (5-11) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to this sample and the cnt-th iteration obtained in step (5-7) T,cnt of the channel attention;
[0112] Specifically, this step is as follows. First, obtain the intermediate layer feature F of the teacher model N T of the height H T,cnt of the width W T,cnt and C T,cnt channels; then, perform global average pooling operation on the intermediate layer feature F T,cnt to obtain a feature map with a dimension of 1×1×C T,cnt ; subsequently, input this feature map into the first fully connected layer (the number of neurons in this layer is C T,cnt / 4) to obtain a feature map with a dimension of 1×1×(C T,cnt / 4); then, apply the ReLU activation function to the feature map output by the first fully connected layer, and input the output after ReLU activation into the second fully connected layer (the number of neurons in this layer is C T,cnt ) to obtain a feature map with a dimension of 1×1×C T,cnt ; finally, apply the Sigmoid activation function to the feature map output by the second fully connected layer, normalize the output value between 0 and 1, and the obtained result is the channel attention C(F T,cnt ) of the teacher model. T,cnt ) of the teacher model.
[0113] (5-12) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to this sample and the cnt-th iteration obtained in step (5-7) S,cnt of the channel attention;
[0114] Specifically, this step is as follows. First, obtain the intermediate layer feature F of the student model N S of the height H S,cnt of the width W S,cnt and CS,cnt and C S,cnt channels; subsequently, perform global average pooling on the intermediate layer feature F S,cnt to obtain a feature map with dimensions 1×1×C S,cnt ; then, input this feature map into the first fully connected layer, where the number of neurons in this layer is C S,cnt / 4, to obtain a feature map with dimensions 1×1×(C S,cnt / 4); then, apply the ReLU activation function to the feature map output by the first fully connected layer, and input the output after ReLU activation into the second fully connected layer, where the number of neurons in this layer is C S,cnt , to obtain a feature map with dimensions 1×1×C S,cnt ; finally, apply the Sigmoid activation function to the feature map output by the second fully connected layer, normalize the output value between 0 and 1, and the resulting value is the channel attention C(F S,cnt ) of the student model.
[0115] The advantages of the above sub-steps (5-11) and (5-12) are as follows: First, global average pooling compresses the intermediate layer features into 1×1, reducing feature redundancy and enhancing the compactness and expressiveness of the features; Second, the fully connected layer further processes the feature map, learning the complex relationships between channels and enhancing the feature discrimination ability; Third, the ReLU activation function introduces non-linearity, enabling the model to learn more complex feature representations and improving the generalization performance; Finally, the Sigmoid activation function normalizes the output, clearly indicating the channel importance, realizing the weighting of important channels, and improving the detection accuracy and efficiency; At the same time, this mechanism has a small computational overhead and does not affect the training and inference efficiency of the model, making it more feasible in agricultural applications.
[0116] (5-13) For each sample obtained in step (5-3), according to the channel attention C(F T,cnt ) of the intermediate layer feature F obtained in step (5-11) and the channel attention C(F T,cnt ) of the intermediate layer feature F obtained in step (5-12), obtain the channel attention weight W S,cnt corresponding to this sample at the cnt-th iteration, so as to use the channel attention of the teacher model to guide the student model to better learn the channel features. T,cnt ch,cnt for the teacher model to guide the student model to better learn the channel features.
[0117] Specifically, this step calculates the absolute difference between the channel attention C(F T,cnt ) and C(F S,cnt ) of the intermediate layer features of the teacher model and the student model, divides it by the scaling factor, and then applies the Softmax function to the result of the above division operation to obtain the channel attention weight Wch,cnt :
[0118] W ch,cnt = C S,cnt × Softmax(|C(F T,cnt ) - C(F S,cnt )| / τ)
[0119] The advantages of this sub-step are as follows. First, by calculating the channel attention difference between the teacher model and the student model, the differences between the two are quantified, providing a clear learning direction for the student model. Second, dividing by the scaling factor can adjust the weight scale, making the weight distribution more reasonable and avoiding optimization problems caused by overly large or small values. Third, the Softmax function normalizes the difference into weights in the form of a probability distribution, ensuring the non-negativity and normalization of the weights, which is convenient for the model to perform weighted learning. Finally, using the channel attention of the teacher model to guide the student model to learn helps the student model focus on important channel features, improves the learning ability and detection performance, and enhances the robustness of the model in complex agricultural scenarios.
[0120] (5 - 14) For each sample obtained in step (5 - 3), calculate the total loss L of the disease detection model cnt for the intermediate layer features F T,cnt of the teacher model.
[0121] (5 - 15) For each sample obtained in step (5 - 3), perform min-max normalization on the gradients obtained in step (5 - 14), that is, normalize the gradient values between 0 and 1, and the resulting value is the pixel point attention weight W T,cnt of the intermediate layer features F pi,cnt corresponding to this sample and for the cnt-th iteration:
[0122]
[0123] where Norm represents min-max normalization.
[0124] The advantages of the above sub-steps (5 - 14) and (5 - 15) are as follows. First, by calculating the gradients, the pixel points that have the greatest impact on the loss can be directly located, providing a clear learning direction for the model and making the model more focused on the pixel points related to disease detection, thereby improving the detection accuracy. Then, normalizing the gradients avoids numerical instability problems and ensures the stability of the training process. Finally, the normalized weights can adaptively adjust the importance of pixel points, highlight key regions, suppress irrelevant information, provide accurate learning guidance for the student model, and improve its learning efficiency and detection performance.
[0125] (5-16) For each sample obtained in step (5-3), according to the spatial attention weight W of the cnt-th iteration corresponding to this sample obtained in step (5-10) sp,cnt , the channel attention weight W obtained in step (5-13) ch,cnt , and the pixel point attention weight W obtained in step (5-15) pi,cnt , calculate the attention loss L of the cnt-th iteration corresponding to this sample AT,cnt and the loss L of the student model OR,cnt .
[0126] The attention loss L in step (5-16) AT,cnt has the following calculation formula:
[0127]
[0128] Among them, the function f(·) represents performing a convolution operation on the intermediate layer features of the aligned teacher model and the intermediate layer features of the student model.
[0129] The advantages of this sub-step are as follows: First, by fusing spatial, channel, and pixel point attention weights, comprehensively considering the importance of different positions and channels in the feature map, the loss function can more accurately focus on key information, improving the detection accuracy and efficiency of the model for disease characteristics; Second, introducing the attention loss helps the student model better learn the feature representation of the teacher model, and effectively improves the performance of the student model by minimizing the weighted feature difference; Finally, this loss function design enhances the generalization ability of the model, enabling it to focus on the features important for the detection task and reducing background noise interference, thus showing higher robustness in complex agricultural scenarios.
[0130] The loss L of the student model in step (5-16) OR,cnt includes three parts: the object loss, the classification loss, and the bounding box regression loss. These three parts are weighted and summed according to their respective importance weights to obtain the total loss of the student model.
[0131] The loss L of the student model OR,cnt has the following calculation formula:
[0132] L OR,cnt = γ1·L obj + γ2·L cls + γ3·L box
[0133] Among them, the value range of γ1 is between 0 and 1, preferably 0.7, the value range of γ2 is between 0 and 1, preferably 0.3, and the value range of γ3 is between 0 and 1, preferably 0.05. All three are weight coefficients, and L obj,cnt represents the object loss of the cnt-th iteration, Lcls,cnt Denote the classification loss at the $cnt$-th iteration as $L$ box,cnt Denote the bounding box regression loss at the $cnt$-th iteration as $L$.
[0134] The object loss $L$ obj,cnt equals the binary cross-entropy loss between the predicted object value of this sample and the true object value $y$ obj :
[0135]
[0136] The classification loss $L$ cls,cnt equals the binary cross-entropy loss between the predicted classification value of this sample and the true classification value $y$ cls :
[0137]
[0138] The bounding box regression loss $L$ box,cnt equals the complete intersection over union loss between the predicted bounding box coordinates of this sample and the true coordinates $y$ box :
[0139]
[0140] where $BCE(\cdot)$ represents the binary cross-entropy loss function and $CIoU(\cdot)$ represents the complete intersection over union loss function.
[0141] The advantages of this sub-step are as follows. First, the loss of the student model is decomposed into object loss, classification loss, and bounding box regression loss, and weighted and summed according to weights, realizing multi-task learning, balancing different task objectives, avoiding a single task from dominating the training, and thus improving the overall performance of the model. Second, the object loss ensures that the model accurately identifies the presence or absence of the object, the classification loss helps to accurately distinguish different categories, and the bounding box regression loss optimizes the accuracy of the bounding box. These three losses respectively optimize the key links in the disease detection task, thus significantly improving the overall performance of the model in the object detection task. Finally, by setting weight coefficients for each part of the loss, the importance of each part of the loss can be flexibly adjusted according to specific task requirements, enabling the model to achieve the optimal performance balance in different scenarios, enhancing the generalization ability and practicality of the model.
[0142] (5-17) For each sample obtained in step (5-3), according to the attention loss $L$ corresponding to this sample calculated in step (5-16) and the loss $L$ of the student model AT,cnt obtain the total loss $L$ of the disease detection model corresponding to this sample at the $cnt$-th iteration OR,cnt cnt = αL AT,cnt + βL OR,cnt , where α and β represent loss weights, and the value ranges of both are between 0 and 1. Preferably, both are equal to 1.
[0143] The advantage of this sub-step is that by combining the attention loss L AT,cnt and the student model loss L OR,cnt to calculate the total loss L cnt , it is possible to optimize feature learning and task performance simultaneously, improving the generalization ability of the model. By setting the weights α and β, the contribution degrees of the attention loss and the student model loss in the total loss can be flexibly adjusted, enabling the model to better adapt to different training objectives. This design utilizes the knowledge of the teacher model to improve the performance of the student model through knowledge distillation, ensuring that it can achieve a high detection accuracy even under limited resources.
[0144] (5 - 18) For each sample obtained in step (5 - 3), use the total loss L cnt of the crop disease detection model at the cnt-th iteration obtained in step (5 - 17) S to perform backpropagation of gradients on the student model N S , and update the parameters of the student model N cnt through the Stochastic Gradient Descent (SGD) algorithm to minimize the total loss L
[0145] (5 - 19) Obtain the optimal parameters of the student model N S to obtain a preliminarily trained crop disease detection model based on knowledge distillation.
[0146] (5 - 20) Use the test set obtained in step (5 - 1) to test the preliminarily trained crop disease detection model based on knowledge distillation in step (5 - 19) until the obtained classification accuracy reaches the optimum, thereby obtaining a finally trained crop disease detection model.
[0147] Figure 3 Taking strawberry anthracnose detection as an example, it demonstrates the roles of spatial attention, channel attention, and pixel - point attention in disease detection. The figure shows that when detecting diseases, the importance of different spatial positions varies. The teacher model can provide a clearer foreground, clarifying the boundaries and positions of the diseases, while the minor differences in the background indicate that not all spatial positions have the same impact on the model performance. In addition, there are significant differences between the teacher model and the student model in channel attention, reflecting the context differences in semantic information processing. Figure 3In the last column, the teacher model precisely focuses on the diseased area through pixel attention, while the student model pays more attention to the irrelevant features around the diseased strawberries.
[0148] Based on the above analysis, the present invention proposes an attention guidance mechanism that includes spatial attention, channel attention, and pixel attention, and designs an attention-guided knowledge distillation loss function for model training. Its advantages are as follows:
[0149] (1) Improve the ability to identify key features: Through the spatial attention and channel attention guidance mechanisms, the student model can learn the teacher model's attention to key spatial positions and channel information, thereby more accurately identifying disease features.
[0150] (2) Enhance feature discrimination: The pixel attention mechanism calculates the contribution of each pixel to the total loss, identifies the significant disease features, and guides the student model to learn these features, enhancing the model's discrimination ability when facing diseases with similar features.
[0151] (3) Improve detection accuracy and robustness: By combining the pixel, spatial, and channel attention guidance mechanisms, the student model can more comprehensively learn the knowledge of the teacher model, significantly improving the accuracy and robustness of disease detection.
[0152] The specific steps of the relevant guidance mechanism and the detailed content of the knowledge distillation loss function have been described in the technical solution of the invention content and will not be elaborated here.
[0153] Example 1
[0154] To comprehensively evaluate the performance of the present invention, on the one hand, traditional object detection evaluation metrics are used, including Precision, Recall, F1 value, and Mean Average Precision (mAP), to measure the accuracy and robustness of the model in the detection task; on the other hand, the Parameter Reduction Rate (PRR) and the Flops Reduction Rate (FRR) are introduced to evaluate the optimization degree of the model's parameters and computational amount after distillation. Among them, PRR is used to measure the reduction degree of the model's parameters, and FRR is used to measure the reduction degree of the model's computational amount (FLOPs). Their calculation formulas are as follows:
[0155]
[0156] Test Example 1:
[0157] The experiments were conducted on three datasets of strawberries, tomatoes, and pods, and the performance of the teacher model YOLOv7, the student model YOLOv7-tiny, the present invention, and three existing methods (FGD, IRKD, GKD-BMFI) was compared. The test metrics included Precision, Recall, F1 value, mean average precision (mAP), as well as PRR and FRR. The experimental results are shown in Table 1.
[0158] Table 1 Performance Comparison between the Present Invention and Existing Methods
[0159]
[0160] The three datasets of strawberries, tomatoes, and pods cover a wide range of crop disease characteristics in the real world, comprehensively verifying the superior performance of the present invention. The experimental results show that the present invention achieves the best or nearly the best results in all indicators, demonstrating strong generalization ability and robustness. This not only proves the effectiveness of the method on these three specific datasets but also indicates its potential for excellent performance in other crop disease detection scenarios, highlighting its broad application prospects in agricultural practice.
[0161] In addition, by means of knowledge distillation technology, the present invention significantly reduces the number of parameters of the student model YOLOv7-tiny from 36.51 million of the teacher model YOLOv7 to 6.03 million, while reducing the computational workload (FLOPs) from 103.3 billion to 13.2 billion. The parameter reduction rate (PRR) and the computational efficiency improvement rate (FRR) reach 83.48% and 87.2% respectively. This optimization result shows that the present invention significantly reduces the computational complexity of the model through knowledge distillation technology while maintaining high detection accuracy. This makes it show significant advantages in resource-constrained agricultural scenarios, especially suitable for real-time deployment and operation on mobile devices or embedded systems.
[0162] Test Example 2:
[0163] Visualize the disease detection results of the above three datasets, and the results are as Figures 4 to 6 shown. Among them, the first row is the ground truth label, and the second to fourth rows are the detection results of the teacher model, the student model, and the present invention respectively. The detection results are presented in the form of a heatmap, and the redder the area, the higher the attention of the model.
[0164] From Figures 4 to 6 the ground truth label, it can be seen that there are significant differences in the disease characteristics of different crops: Figure 4 The leaf spot and powdery mildew diseases in Figure 5 are dense and overlapping; Figure 6The angular leaf spot and bean rust in it are major targets. At the same time, some disease characteristics are similar. For example, leaf spot, target spot, and black spot. All three of these diseases are mainly manifested by the lesions on the leaves, and the morphology and color of the lesions are relatively similar, making them easy to confuse. These diverse disease characteristics pose relatively high requirements for the accuracy and robustness of the detection model.
[0165] Generally speaking, the analysis and visualization results in Table 1 show that the present invention significantly improves the performance of the student model. This improvement is mainly attributed to its attention guidance mechanism. This mechanism significantly enhances the model's recognition accuracy for disease regions of different sizes by identifying the significant regions in the feature map and combining gradient-based pixel-level attention. At the same time, the integration of spatial attention and channel attention further optimizes the overall performance of the model.
[0166] These features enable the present invention to exhibit significant application potential in agricultural practices, providing powerful methods and tools for improving the accuracy and efficiency of crop disease diagnosis systems, especially suitable for complex and diverse disease detection scenarios.
[0167] Those skilled in the art can easily understand that the above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A crop disease detection method based on knowledge distillation, characterized in that: The following steps are involved: (1) Obtain the image of the crop to be detected. (2) Performing image enhancement processing on the image of the crop to be detected obtained in step (1) to obtain an enhanced image. (3) Performing data preprocessing on the enhanced image obtained in step (2) to obtain a preprocessed image. (4) The preprocessed image obtained in step (3) is input into a pre-trained crop disease detection model based on knowledge distillation to obtain the final detection result.
2. The crop disease detection method based on knowledge distillation according to claim 1, characterized in that: The data enhancement processing in step (2) adopts any one of the following eight methods or any combination of any multiple of them: image HSV enhancement, including hue, saturation and brightness enhancement, and the enhancement factors are set to 0.015, 0.7 and 0.4 respectively; For image translation, the translation factor is 0.2; for image scaling, the scaling factor is 0.9; for image shearing, the shearing factor is 0.5; for image perspective, the perspective factor is 0.001; for image left-right flipping, the flipping factor is 0.5; for multi-image splicing, the splicing factor is 1.0; for image mixing, the mixing factor is 0.
15.
3. The crop disease detection method based on knowledge distillation according to claim 1 or 2, characterized in that: The detection result obtained in step (4) is in the form of a detection frame, each of which indicates the predicted disease location and disease category. The crop disease detection model based on knowledge distillation in step (4) adopts a knowledge distillation model, which includes a teacher model N T and a student model N S .
4. The crop disease detection method based on knowledge distillation according to any one of claims 1 to 3, characterized in that: The crop disease detection model based on knowledge distillation is trained through the following steps: (5-1) Download the open source mixed dataset consisting of strawberry dataset, tomato dataset and bean pod dataset, and divide the dataset into training set and test set in a ratio of 8:
2. (5-2) Performing image enhancement processing on the training set obtained in step (5-1) to obtain an enhanced training set. (5-3) Performing data preprocessing on the enhanced training set obtained in step (5-2) to obtain a preprocessed training set. (5-4) For each sample in the preprocessed training set obtained in step (5-3), the sample is input into the teacher model N T Perform training and freeze the teacher model N after training is completed T Parameters. (5-5) Set the counter cnt = 1 and initialize the total loss L of the first iteration of the disease detection model cnt is a random number between 0 and 1; (5-6) Determine whether cnt is greater than a preset iteration threshold, if so, proceed to step (5-19), otherwise proceed to step (5-7); (5-7) For each sample in the preprocessed training set obtained in step (5-3), the sample is input into the trained and frozen parameter teacher model N obtained in step (5-4) T and student model N S Perform forward propagation to obtain the intermediate layer feature F corresponding to the sample and the cnt-th iteration T,cnt and F S,cnt . (5-8) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to the sample obtained in step (5-7) at the cnt-th iteration T,cnt spatial attention; (5-9) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to the sample obtained in step (5-7) at the cnt-th iteration S,cnt spatial attention; (5-10) The intermediate layer feature F obtained according to step (5-8) T,cnt The spatial attention S(F T,cnt ) and the intermediate layer features F obtained in steps (5-9) S,cnt The spatial attention S(F S,cnt ) Get the spatial attention weight W corresponding to the sample and the cnt-th iteration sp,cnt (The purpose is to use the spatial attention of the teacher model to guide the student model to better learn spatial features). (5-11) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to the sample obtained in step (5-7) at the cnt-th iteration T,cnt Channel attention; (5-12) For each sample obtained in step (5-3), obtain the intermediate layer feature F corresponding to the sample obtained in step (5-7) at the cnt-th iteration S,cnt Channel attention; (5-13) For each sample obtained in step (5-3), the intermediate layer feature F obtained in step (5-11) T ,cnt The channel attention C(F T,cnt ) and the intermediate layer features F obtained in step (5-12) S,cnt The channel attention C(F T,cnt ) Get the channel attention weight W corresponding to the sample and the cnt-th iteration ch,cnt , so as to utilize the channel attention of the teacher model to guide the student model to better learn the channel features. (5-14) For each sample obtained in step (5-3), calculate the total loss L of the disease detection model cnt For the intermediate layer features F of the teacher model T,cnt gradient. (5-15) For each sample obtained in step (5-3), the gradient obtained in step (5-14) is normalized to the minimum and maximum values between 0 and 1. The result is the intermediate layer feature F corresponding to the sample at the cnt iteration. T,cnt The pixel attention weight W pi,cnt . (5-16) For each sample obtained in step (5-3), the spatial attention weight W corresponding to the sample at the cnt-th iteration obtained in step (5-10) is sp,cnt , the channel attention weight W obtained in step (5-13) ch,cnt , and the pixel attention weight W obtained in step (5-15) pi,cnt Calculate the attention loss L corresponding to the sample and the cnt-th iteration AT,cnt and the loss L of the student model OR,cnt . (5-17) For each sample obtained in step (5-3), the attention loss L corresponding to the sample at the cnt-th iteration calculated in step (5-16) is AT,cnt and the loss L of the student model OR,cnt Get the total loss L of the disease detection model corresponding to the sample and the cnt-th iteration cnt =αL AT,cnt +βL OR,cnt , where α and β represent loss weights, and both range from 0 to 1. Preferably, both are equal to 1. (5-18) For each sample obtained in step (5-3), the total loss L of the crop disease detection model of the cntth iteration obtained in step (5-17) is cnt For student models N S Perform gradient backpropagation and update the student model N through the stochastic gradient descent algorithm SGD S Parameters so that the total loss L cnt Minimize, set counter cnt=cnt+1, and return to step (5-6). (5-19) Get the student model N S The optimal parameters of are obtained, thus obtaining a preliminarily trained crop disease detection model based on knowledge distillation. (5-20) Using the test set obtained in step (5-1), the crop disease detection model based on knowledge distillation that was initially trained in step (5-19) is tested until the classification accuracy is optimal, thereby obtaining the final trained crop disease detection model.
5. The crop disease detection method based on knowledge distillation according to claim 4, characterized in that: Steps (5-8) are as follows: first, obtain the teacher model N T The intermediate layer features F T,cnt Height H T,cnt , Width W T,cnt and C T,cnt channels; then, in the first branch, the intermediate layer features F T,cnt Apply a convolution operation with a kernel size of 1×9, then apply another convolution operation with a kernel size of 9×1 to the output of the first convolution operation, and finally output the feature map S1; then, in the second branch, first perform the intermediate layer feature F T,cnt Apply a convolution operation with a convolution kernel size of 9×1, then apply another convolution operation to the output of the first convolution operation with a convolution kernel size of 1×9, and finally output the feature map S2; then, add the feature map S1 output by the first branch and the feature map S2 output by the second branch element-wise to obtain the fused feature map; finally, apply the Sigmoid activation function to the fused feature map and normalize the output value to between 0 and 1 to obtain the intermediate layer feature F T,cnt The spatial attention S(F T,cnt ). Steps (5-9) are as follows: first, obtain the student model N S The intermediate layer features F S,cnt Height H S,cnt , Width W S,cnt and C S,cnt channels, and then, using the same method as the teacher model, the intermediate layer features F S,cnt The convolution operation of two parallel branches is performed to obtain feature maps S1 and S2 respectively. After that, the feature maps S1 and S2 output by the two branches are element-wise added to obtain the fused feature map. Finally, the Sigmoid activation function is applied to the fused feature map to normalize the output value between 0 and 1 to obtain the intermediate layer feature F. S,cnt The spatial attention S(F S,cnt ).
6. The crop disease detection method based on knowledge distillation according to claim 5, characterized in that: Steps (5-10) are as follows: Calculate the spatial attention S(F T,cnt ) and S(F S,cnt ), and divided by the scaling factor τ (its value range is (0,1], preferably 0.5), and then the Softmax function is applied to the above division result to obtain the spatial attention weight W sp,cnt : W sp,cnt =H S,cnt ×W S,cnt ×Softmax(|S(F T,cnt )―S(F S,cnt )| / τ) Steps (5-11) are as follows: first, obtain the teacher model N T The intermediate layer features F T,cnt Height H T,cnt , Width W T,cnt and C T,cnt channels; then, the intermediate layer feature F T,cnt Perform a global average pooling operation to obtain a 1×1×C T,cnt Then, the feature map is input into the first fully connected layer (the number of neurons in this layer is C T,cnt / 4) to obtain a dimension of 1×1×(C T,cnt / 4) feature map; then, the ReLU activation function is applied to the feature map output by the first fully connected layer, and the output after ReLU activation is input to the second fully connected layer (the number of neurons in this layer is C T,cnt ) to obtain a dimension of 1×1×C T,cnt Finally, the Sigmoid activation function is applied to the feature map output by the second fully connected layer to normalize the output value to between 0 and 1. The result is the channel attention C(F T,cnt ). Steps (5-12) are as follows: first, obtain the student model N S The intermediate layer features F S,cnt Height H S,cnt , Width W S,cnt and C S′cnt channels; then , for the intermediate layer feature F S,cnt Perform a global average pooling operation to obtain a 1×1×C S′cnt The feature map is then input into the first fully connected layer, which has C neurons. S′cnt / 4, to obtain a dimension of 1×1×(C S′cnt / 4) feature map; then, the ReLU activation function is applied to the feature map output by the first fully connected layer, and the output after ReLU activation is input to the second fully connected layer, which has C neurons. S′cnt , to obtain a dimension of 1×1×C S′cnt Finally, the Sigmoid activation function is applied to the feature map output by the second fully connected layer to normalize the output value to between 0 and 1. The result is the channel attention C(F S,cnt ).
7. The crop disease detection method based on knowledge distillation according to claim 6, characterized in that: Step (5-13) is to calculate the channel attention C(F T,cnt ) and C(F S,cnt ), and divided by the scaling factor, and then applied the Softmax function to the above division result to obtain the channel attention weight W ch,cnt : W ch,cnt =C S,cnt ×Softmax(|C(F T,cnt )―C(F S,cnt )| / τ) Steps (5-15) are performed using the following formula: Among them, Norm means minimum and maximum normalization.
8. The crop disease detection method based on knowledge distillation according to claim 7, characterized in that: The attention loss L in step (5-16) AT,cnt The calculation formula is: The function f(·) represents the convolution operation on the intermediate layer features of the aligned teacher model and the intermediate layer features of the student model. The loss L of the student model OR,cnt The calculation formula is: L OR,cnt =γ1·L obj +γ2·L cls +γ3·L box Among them, the value range of γ1 is between 0 and 1, preferably 0.7, the value range of γ2 is between 0 and 1, preferably 0.3, and the value range of γ3 is between 0 and 1, preferably 0.
05. All three are weight coefficients. obj,cnt represents the target loss of the cnt-th iteration, L cls,cnt represents the classification loss of the cnt-th iteration, L box,cnt represents the bounding box regression loss at the cnt-th iteration.
9. The crop disease detection method based on knowledge distillation according to claim 8, characterized in that: Target loss L obj,cnt Equal to the target value predicted by the sample and the true target value y obj The binary cross entropy loss between: Classification loss L cls,cnt Equal to the classification value predicted for this sample and the true classification value y cls The binary cross entropy loss between: Bounding box regression loss L box,cnt Equal to the bounding box coordinates predicted for this sample and the real coordinate y box The complete intersection-over-union loss between Among them, BCE(·) represents the binary cross entropy loss function, and CIoU(·) represents the complete intersection-over-union loss function.
10. A crop disease detection system based on knowledge distillation, characterized in that: include: The first module is used to obtain images of crops to be detected. The second module is used to perform image enhancement processing on the image of the crop to be detected obtained by the first module to obtain an enhanced image. The third module is used to perform data preprocessing on the enhanced image acquired by the second module to obtain a preprocessed image. The fourth module is used to input the preprocessed image obtained by the third module into a pre-trained crop disease detection model based on knowledge distillation to obtain the final detection result.