Grape leaf disease real-time detection method based on improved RTDETR model

By improving the feature extraction, interaction, and fusion modules of the RTDETR model, the problems of computational resource consumption and detection speed in grape leaf disease detection of the YOLO series models were solved, achieving efficient and real-time disease detection.

CN120912852AActive Publication Date: 2025-11-07BEIFANG UNIV OF NATITIES

Patent Information

Application Number
CN202510825368.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-11-07
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

Existing YOLO series models consume computational resources as the target density increases in grape leaf disease detection, and the NMS post-processing mechanism limits the detection speed, making it difficult to achieve real-time detection.

Method used

The improved RTDETR model uses PVNet for feature extraction, SPPELAN for feature interaction, and ASF-YOLO for feature fusion, enhancing feature information exchange and robustness, reducing computational cost, and maintaining high detection accuracy.

Benefits of technology

While reducing computational load, the speed and accuracy of grape leaf disease detection were significantly improved, the problem of missed detection of small targets was solved, and the robustness of the model in complex environments was enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912852A_ABST
    Figure CN120912852A_ABST
Patent Text Reader

Abstract

The invention discloses a grape leaf disease real-time detection method based on an improved RTDETR model, and the method comprises the steps: improving a feature extraction module, a feature interaction module and a feature fusion module of an original RTDETR model, transmitting image data in a training set into the feature extraction module for feature extraction, and obtaining the feature information of a disease; the extracted feature information is input into a feature interaction module based on SPPELAN design to be processed, and information exchange and fusion between features are enhanced; inputting the information processed by the feature interaction module into a feature fusion module based on ASF-YOLO reconstruction for integration; and the integrated features are transmitted to a decoder detection network to obtain a disease detection result, and a test set is used for testing. According to the method, the problems of calculation redundancy, insufficient real-time performance, small target missing detection and the like in grape leaf disease detection are solved, and relatively high detection precision and speed are still kept while the calculation complexity is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of deep learning disease detection, in particular to a grape leaf disease real-time detection method based on an improved RTDETR model. BACKGROUND

[0002] Grape is an important economic crop, and its growth cycle stages will inevitably be affected by diseases, which seriously damages the yield and quality of grape and causes considerable economic losses. According to the report released by the International Organisation of Vine and Wine, the global grape total production in 2023 was 74.7 million tons, of which the yield loss due to natural disasters, diseases and other factors reached 4.4 million tons. In 2024, the outbreak of fungal diseases led to a 23.5% drop in French wine production, setting a new record low since 1957. In this context, how to detect and identify grape leaf diseases in a timely and accurate manner is particularly important, and a real-time detection method is undoubtedly the key to solving this problem. On the one hand, it provides a basis for timely prevention and control measures, effectively controlling the spread of diseases and reducing the impact on grape yield and quality. On the other hand, it greatly improves the efficiency of disease monitoring and saves manpower and time costs.

[0003] With the rapid rise of smart agriculture, image processing and deep learning-based detection methods have received widespread welcome in target identification. Among them, the YOLO series model is more representative. This model has achieved a good balance between lightweight and detection accuracy, but it also has some problems. For example, the YOLO series model relies on the non-maximum suppression (NMS) post-processing mechanism to eliminate redundant detection boxes. This process requires sorting and thresholding the candidate boxes, which leads to an increase in computational resource consumption with the target density, especially in scenarios where leaf lesions are dense and overlapping. NMS significantly inhibits the detection speed, becoming a key obstacle to the deployment of real-time detection systems. SUMMARY

[0004] The present application aims to overcome the shortcomings and deficiencies of the prior art and proposes a grape leaf disease real-time detection method based on an improved RTDETR model. By utilizing the advantages of the RTDETR model, the impact of NMS on detection speed is eliminated, while maintaining high detection accuracy, speed and recall rate with significantly reduced computational load.

[0005] In order to achieve the above object, the technical scheme provided by the present application is: a grape leaf disease real-time detection method based on an improved RTDETR model, wherein the improved RTDETR model is improved on the feature extraction module, the feature interaction module and the feature fusion module of the original RTDETR model, wherein the improvement on the feature extraction module is: replacing the original backbone network HGNetV2 with PVNet, which is a fusion of VanillaNet and PConv and is uniformly named as PVNet; the improvement on the feature interaction module is: using the scale invariance advantage of SPPELAN to replace the AIFI module in the RTDETR model; and the improvement on the feature fusion module is: using the collaborative optimization characteristics of the multi-scale feature fusion and attention mechanism of ASF-YOLO to construct the feature fusion module, so as to enhance the representation ability of the target region.

[0006] The specific implementation of the grape leaf disease real-time detection method includes the following steps:

[0007] 1) Obtain image data of grape leaf diseases, and label the disease categories and positions to construct a structured annotation dataset conforming to the VOC standard, then perform data enhancement to expand the dataset, and finally divide the dataset into a training set, a validation set and a test set;

[0008] 2) input the training set into the improved RTDETR model for training, the process being: first, obtain the feature information of the corresponding disease through the feature extraction module, input the extracted feature information into the feature interaction module for processing to enhance the information exchange and fusion between the features, then input the information processed by the feature interaction module into the feature fusion module for integration, and finally pass the integrated features to the decoder detection network of the improved RTDETR model to obtain the detection result of the grape leaf disease image, including disease category and position information; after multiple rounds of iterative training, the model performance indicators are evaluated with the validation set during the training, and when the performance of the validation set no longer improves or reaches the preset number of rounds, the training is stopped and the model parameters are saved to obtain the optimal model;

[0009] 3) input the test set into the optimal model for forward propagation prediction to obtain the accurate detection result of the grape leaf disease image.

[0010] Further, in step 1), disease samples are selected from the PlantVillage open-source disease image library, the Labelimg labeling tool is used to label the disease categories and positions to construct a structured annotation dataset conforming to the VOC standard, and then part of the samples are selected by data enhancement, including transforming the brightness, saturation, enhancing the contrast and adding noise of the samples to simulate disease samples under special scenarios, so as to expand the dataset, and finally the dataset is divided into a training set, a validation set and a test set.

[0011] Further, in step 2), the feature extraction module uses VanillaNet as the basic network, fuses VanillaNet with PConv to form a new backbone network, and uniformly names it as PVNet, aiming to improve the calculation efficiency while maintaining efficient feature extraction; using the dynamic training mechanism of VanillaNet, the single-layer convolution is split into two layers and an adjustable activation function is inserted in the early training stage, gradually fusing the nonlinear calculation into the weight, ensuring that it is restored to a single-layer efficient structure during inference, and a series of activation functions are designed to capture local context information through neighborhood weighting, enhancing the ability to distinguish the edge of the lesion and the complex background; then using the selective channel processing mechanism of PConv, only local channels are used to perform spatial feature extraction, while the remaining channels directly pass the original data, reducing the parameter calculation amount while avoiding the problem of decreased feature expression ability caused by excessive compression of channels; the above new backbone network includes a Stem module and a Stage module, wherein the Stem module includes a convolution module, a normalization module and an activation function module, and the Stage module includes a convolution module, a normalization module, an activation function module and a pooling module; the overall process of the feature extraction module is: the input image is quickly down-sampled by 4x4 convolution, then the channel is adjusted by 1x1 convolution, and the feature response is adjusted by dynamic activation function Activation, enhancing the sensitivity to the edge of the lesion, and then passing through multiple Stage modules to obtain the final image resolution, and each Stage module is expressed by a mathematical formula as follows:

[0012] F Stage (X)=Activation(MaxPool2d(BN(Conv 1×1 (LeakyReLU(BN(PConv 3×3 (X)))))))

[0013] In the formula, F Stage (X) represents the Stage module itself, X represents the input image, PConv 3×3 represents a 3x3 partial convolution, BN represents normalization, LeakyReLU represents an activation function, Conv 1×1 represents a 1x1 convolution, and MaxPool2d represents a two-dimensional maximum pooling operation, and Activation represents a dynamic activation function.

[0014] The feature interaction module is to gradually expand the receptive field of the feature information obtained by the feature extraction module through the SPPELAN multi-level pooling design, while capturing pixel-level lesion details and regional-level semantic features, solving the problem of insufficient adaptability to multi-size targets; the feature interaction module uses the multi-scale feature fusion advantage of SPPELAN to replace the AIFI module in the RTDETR model, and fuses the feature maps at different stages of the feature extraction module through cross-level residual connection, uses the high-resolution information of the shallow layer to compensate for the detail loss of the deep semantic features, enhances the information exchange and fusion between features, and significantly improves the robustness to leaf shape or occluded areas;

[0015] The feature fusion module is to integrate the information processed by the feature interaction module, and redesign the feature fusion module of the RTDETR model using ASF-YOLO. First, the SSFF module is used to capture different grape leaf lesion shape features; then the TFE module is used to fuse multi-scale feature maps to enrich the detail information of the lesion; finally, the CPAM is introduced to integrate the SSFF module and the TFE module, to realize adaptive channel weight allocation and spatial position focusing, to improve the robustness of detection in complex field environments and enhance the representation ability of target regions; finally, the integrated feature map is selected by IoU-aware Query Selection to select a fixed number of features as target queries, and then through Decoder&Head, the prediction head is mapped to confidence and bounding box to get the final detection result, including disease class and location information.

[0016] Further, in step 2), through multiple rounds of iterative training, the model performance indicators are evaluated during the period using the validation set, and an important indicator, FPS, is focused on, which refers to the number of images that the model can process per unit time, which is closely related to real-time performance. The higher the FPS value, the faster the model responds. By monitoring the evaluation indicators, it is determined whether the model has overfitting phenomenon. When the performance indicators on the validation set no longer improve for consecutive multiple training rounds, or the number of training rounds reaches the pre-set maximum number of rounds, the training is stopped and the model parameters are saved. At this time, it is determined as the optimal model; wherein, the formula of FPS is as follows:

[0017]

[0018] AllTime=pre+inf+post

[0019] In the formula, ms represents millisecond, AllTime represents total time, pre represents preprocessing time Preprocess, inf represents inference time Inference, and post represents post-processing time Postprocess.

[0020] Compared with the prior art, the present application has the following advantages and beneficial effects:

[0021] 1、The present application is aimed at the difficulty of real-time detection of grape leaf diseases, and is improved on the basis of the RTDETR model. In the feature extraction module, VanillaNet is used as the basic network, VanillaNet is combined with PConv to form a new backbone network PVNet, which aims to reduce the calculation amount while maintaining efficient feature extraction.

[0022] 2、The present application is based on the advantage of SPPELAN in the feature interaction module, and constructs a feature fusion module that takes into account global semantics and local details, solves the problem of insufficient adaptability to multi-size targets, and significantly improves the recognition ability of the model to leaf shape or occluded areas.

[0023] 3、The present application completely discards the original module of the RTDETR model in the feature fusion module, and introduces the SSFF module and the TFE module based on ASF-YOLO to effectively improve the robustness of the model to leaf texture, illumination and other interference in the complex field environment.

[0024] In summary, the present application first reduces the computational complexity by reconstructing the feature extraction module, secondly enhances the information exchange and fusion between features through the improved feature interaction module, and finally effectively solves the problem of small target missed detection in the disease detection task through the improved feature fusion module. In addition, this method also shows high detection accuracy and speed on the validation set and test set. BRIEF DESCRIPTION OF DRAWINGS

[0025] Figure 1 To improve the structure diagram of the RTDETR model. In the figure, Backbone is the improved feature extraction module (its new backbone network is PVNet), RepC3 is the reparameterization convolution module, Zoom_cat is the multi-scale feature fusion module, corresponding to the TFE module of ASF-YOLO, Conv is the ordinary convolution module, Upsample is the up sampling operation module, Concat is the fusion operation, ScalSeq is the scale sequence feature fusion module, corresponding to the SSFF module of ASF-YOLO, Add is the addition operation, corresponding to the CPAM module of ASF-YOLO.

[0026] Figure 2is a structural diagram of the feature extraction module; in the figure, Conv2d 4x4 / s4 is a two-dimensional convolution with a convolution kernel size of 4x4 and a step size of 4, BN is normalization, LeakyReLU is an activation function, Conv2d 1x1 is a two-dimensional convolution with a convolution kernel size of 1x1, Activation is a dynamic activation function, PConv 3x3 is a partial convolution with a convolution kernel size of 3x3, and MaxPool2d 2x2 / s2 is a two-dimensional maximum pooling operation with a pooling kernel size of 2x2 and a step size of 2.

[0027] Figure 3 is a structural diagram of SPPELAN; in the figure, Conv is a convolution with a convolution kernel size of 1x1, a step size of 1, and a padding of 0, MaxPool2d is a two-dimensional maximum pooling operation with a pooling kernel size of 5x5, a step size of 1, and a padding of 2, and concat is a fusion operation.

[0028] Figure 4 is a structural diagram of the feature fusion module (SSFF+TFE+CPAM); in the figure, x is input, Split is a data shunting operation module, Conv is a 1x1 convolution, SiLU is an activation function, Add Dim is a dimension expansion operation module, Upsample is an up-sampling operation module, 3D Conv is a three-dimensional convolution, 3D BN is a three-dimensional batch normalization, 3D Max Pool is a three-dimensional maximum pooling, Remove Dim is a dimension removal operation module, Concatenate is a splicing operation, Max Pooling is a maximum pooling operation, Avg Pooling is an average pooling operation, Add is an addition operation, keep is an unchanged operation, Squeeze last dim is a dimension compression operation module, Transpose last two dims is a dimension transformation operation module, Transpose back and Unsqueeze are dimension adjustment operation modules, Sigmoid is an activation function, Multiply is a feature fusion operation module, feature map is a feature map, Dim h and Dim w are dimension change operations, Expand is a feature expansion operation module, and Output is output. DETAILED DESCRIPTION

[0029] The application will be further described below in conjunction with specific embodiments, but the embodiments of the application are not limited thereto.

[0030] As Figures 1 to 4As shown, the embodiment discloses a grape leaf disease real-time detection method based on an improved RTDETR model, which is improved on the original RTDETR model in feature extraction module, feature interaction module and feature fusion module. The improvement on the feature extraction module is: replacing the original backbone network HGNetV2 with PVNet, which is the fusion of VanillaNet and PConv, and is uniformly named as PVNet. The improvement on the feature interaction module is: using the scale invariance advantage of SPPELAN to replace the AIFI module in the RTDETR model. The improvement on the feature fusion module is: using the collaborative optimization characteristics of multi-scale feature fusion and attention mechanism of ASF-YOLO to construct the feature fusion module, and enhancing the representation ability of the target region.

[0031] The specific implementation of the grape leaf disease real-time detection method includes the following steps:

[0032] 1) Obtain image data of grape leaf diseases through the PlantVillage open-source disease image library, and select 6238 images labeled by Labelimg labeling tool for disease category and position to construct a structured annotation dataset conforming to VOC standards. Then, select part of the samples to expand the dataset by data enhancement, including transforming the brightness, saturation, enhancing the contrast and adding noise of the samples to simulate disease samples in special scenarios, and divide the dataset into a training set of 4990 images, a validation set of 623 images and a test set of 625 images in a ratio of 8:1:1, as shown in Table 1.

[0033] Table 1: Division of dataset

[0034] Disease type Total number of images Training set Validation set Test set Black measles 2038 1660 206 172 Black rot 2072 1621 226 225 Leaf spot 2128 1709 191 228

[0035] 2) Input the training set into the improved RTDETR model for training. The process is as follows: first, obtain the feature information of the corresponding disease through the feature extraction module, input the extracted feature information into the feature interaction module for processing, enhance the information exchange and fusion between features, then input the information processed by the feature interaction module into the feature fusion module for integration, and finally pass the integrated features to the decoder detection network of the improved RTDETR model to obtain the detection result of the grape leaf disease image, including disease category and position information. After multiple rounds of iterative training, the model performance indicators are evaluated with the validation set. When the performance of the validation set no longer improves or reaches the preset number of rounds, the training is stopped and the model parameters are saved to obtain the optimal model. The specific situation is as follows:

[0036] The feature extraction module uses VanillaNet as a basic network, fuses VanillaNet with PConv to form a new backbone network, and is uniformly named as PVNet, aiming to improve the calculation efficiency while maintaining efficient feature extraction; the dynamic training mechanism of VanillaNet is used, the single-layer convolution is split into two layers and the adjustable activation function is inserted in the early training stage, the nonlinear calculation is gradually fused into the weight, ensuring that the inference is restored to a single-layer efficient structure, and a series of activation functions are designed, the local context information is captured through neighborhood weighting, and the ability to distinguish the edge of the lesion and the complex background is enhanced; then the selective channel processing mechanism of PConv is used, only the spatial feature extraction is performed on the local channel, and the original data is directly transmitted in the remaining channels, reducing the parameter calculation amount while avoiding the problem of decline of feature expression ability caused by excessive compression of channels; the above new backbone network includes a Stem module and a Stage module, wherein the Stem module includes a convolution module, a normalization module and an activation function module, and the Stage module includes a convolution module, a normalization module, an activation function module and a pooling module; in addition, taking 640*640 as an example, the overall process of the feature extraction module is that: the input image is quickly down-sampled to 160*160 resolution through 4*4 convolution, then the channel is adjusted through 1*1 convolution, and the feature response is adjusted through the dynamic activation function Activation, the sensitivity to the edge of the lesion is enhanced, then a plurality of Stage modules are passed, and finally the resolution of the image is down-sampled from 160*160 to 20*20, and each Stage module is expressed by a mathematical formula as follows:

[0037] F Stage (X)=Activation(MaxPool2d(BN(Conv 1×1 (LeakyReLU(BN(PConv 3×3 (X)))))))

[0038] In the formula, F Stage (X) represents the Stage module itself, X represents the input image, PConv 3×3 represents a 3*3 partial convolution, BN represents normalization, LeakyReLU represents an activation function, Conv 1×1 represents a 1*1 convolution, MaxPool2d represents a two-dimensional maximum pooling operation, and Activation represents a dynamic activation function.

[0039] The feature interaction module is to gradually expand the receptive field of the feature information obtained by the feature extraction module through the SPPELAN multi-level pooling design, while capturing pixel-level lesion details and regional semantic features, solving the problem of insufficient adaptability to multi-size targets; the feature interaction module uses the multi-scale feature fusion advantage of SPPELAN to replace the AIFI module in the RTDETR model, and fuses the feature maps at different stages of the feature extraction module through cross-level residual connection, uses the high-resolution information of the shallow layer to compensate for the detail loss of the deep semantic features, enhances the information exchange and fusion between features, and significantly improves the robustness to leaf shape or occluded areas;

[0040] The feature fusion module is to integrate the information processed by the feature interaction module, and redesign the feature fusion module of the RTDETR model using ASF-YOLO. First, the SSFF module is used to capture different grape leaf lesion shape features; then the TFE module is used to fuse multi-scale feature maps to enrich the detail information of the lesion; finally, the CPAM is introduced to integrate the SSFF module and the TFE module, to realize adaptive channel weight allocation and spatial position focusing, to improve the robustness of detection in complex field environments and enhance the representation ability of the target area; finally, the integrated feature map is selected as the target query through IoU-aware Query Selection, and then through Decoder&Head, the prediction head is mapped to confidence and bounding box to get the final detection result, including disease class and location information;

[0041] Through multiple iterations of training, the model performance indicators are evaluated during the training using the validation set, and an important indicator, FPS, is focused on. FPS refers to the number of images that the model can process per unit time, which is closely related to real-time performance. The higher the FPS value, the faster the model responds. By monitoring the evaluation indicators, it is determined whether the model has overfitting phenomenon. When the performance indicators on the validation set no longer improve for multiple consecutive training rounds, or the number of training rounds reaches the pre-set maximum number of rounds, the training is stopped and the model parameters are saved. At this time, it is determined as the optimal model. The formula of FPS is as follows:

[0042]

[0043] AllTime=pre+inf+post

[0044] In the formula, ms represents milliseconds, AllTime represents the total time, pre represents the preprocessing time Preprocess, inf represents the inference time Inference, and post represents the post-processing time Postprocess.

[0045] 3) The test set is input to the optimal model for forward propagation prediction, that is, the accurate detection result of the grape leaf disease image can be obtained.

[0046] The improved RTDETR model is compared with YOLOV8L, YOLOV9C, YOLOV10L, YOLO11L, YOLOV12L, RTDETR-Mamba-L and RTDETR-L models, and the experimental results are shown in Table 2.

[0047] Table 2 Analysis of experimental results

[0048]

[0049] In the above experimental results, it can be clearly seen that the improved RTDETR model achieves better performance than other mainstream models, which fully proves the effectiveness and robustness of the method of the present application. In the follow-up research, the method of the present application is prepared to be applied to other disease detection fields to fully verify the generalization ability of the method.

[0050] The above examples are the preferred embodiments of the present application, but the embodiments of the present application are not limited by the above examples, and any changes, modifications, substitutions, combinations, simplifications made without departing from the spirit and principles of the present application should be equivalent replacement methods, and are all included in the protection scope of the present application.

Claims

1. A method for real-time detection of grape leaf diseases based on an improved RTDETR model, characterized in that, The improved RTDETR model is improved on the feature extraction module, feature interaction module and feature fusion module of the original RTDETR model, wherein the improvement on the feature extraction module is: replacing the original backbone network HGNetV2 with PVNet, which is a fusion of VanillaNet and PConv, and is uniformly named as PVNet; the improvement on the feature interaction module is: using the scale invariance advantage of SPPELAN to replace the AIFI module in the RTDETR model; the improvement on the feature fusion module is: using the collaborative optimization characteristics of the multi-scale feature fusion and attention mechanism of ASF-YOLO to construct the feature fusion module, and enhancing the representation ability of the target region; The specific implementation of the grape leaf disease real-time detection method includes the following steps: 1) Obtain image data of grape leaf diseases, and label the disease categories and positions to construct a structured annotation dataset conforming to the VOC standard, then perform data enhancement to expand the dataset, and finally divide the dataset into a training set, a validation set and a test set; 2) input the training set into the improved RTDETR model for training, the process is: first, obtain the feature information of the corresponding disease through the feature extraction module, input the extracted feature information into the feature interaction module for processing, enhance the information exchange and fusion between features, then input the information processed by the feature interaction module into the feature fusion module for integration, finally pass the integrated features to the decoder detection network of the improved RTDETR model to obtain the detection result of the grape leaf disease image, including disease category and position information; after multiple rounds of iterative training, the model performance indicators are evaluated with the validation set during the training, when the validation set performance no longer improves or reaches the preset number of rounds, stop training and save the model parameters, and obtain the optimal model; 3) input the test set into the optimal model for forward propagation prediction, and the accurate detection result of the grape leaf disease image can be obtained.

2. The method for real-time detection of grape leaf diseases based on the improved RTDETR model according to claim 1, characterized in that, In step 1), disease samples are selected from the PlantVillage open source disease image library, and the Labelimg labeling tool is used to label the disease categories and positions to construct a structured annotation dataset conforming to the VOC standard, then some samples are selected by data enhancement, including transforming the brightness, saturation, enhancing the contrast and adding noise of the samples to simulate disease samples in special scenarios, thereby expanding the dataset, and finally dividing the dataset into a training set, a validation set and a test set.

3. The method for real-time detection of grape leaf diseases based on the improved RTDETR model according to claim 1, characterized in that, In step 2), the feature extraction module uses VanillaNet as the basic network, fuses VanillaNet and PConv to form a new backbone network, and is uniformly named as PVNet, aiming to improve the calculation efficiency while maintaining efficient feature extraction; The dynamic training mechanism of VanillaNet is used, the single-layer convolution is split into two layers and the adjustable activation function is inserted in the early stage of training, the nonlinear calculation is gradually integrated into the weight, the single-layer efficient structure is restored during inference, a series of activation functions are designed, the local context information is captured through neighborhood weighting, and the discrimination ability of the lesion edge and the complex background is enhanced; the selective channel processing mechanism of PConv is used, spatial feature extraction is only performed on local channels, and the original data is directly transmitted in the remaining channels, thereby reducing the parameter calculation amount and avoiding the problem of decreased feature expression ability caused by excessive compression of channels; the above new backbone network includes a Stem module and a Stage module, wherein the Stem module includes a convolution module, a normalization module and an activation function module, and the Stage module includes a convolution module, a normalization module, an activation function module and a pooling module; the overall process of the feature extraction module is that: the input image is quickly down-sampled through 4*4 convolution, then the channel is adjusted through 1*1 convolution, the feature response is adjusted through the dynamic activation function Activation, the sensitivity to the lesion edge is enhanced, and then a plurality of Stage modules are used to obtain the final image resolution, and each Stage module is expressed by a mathematical formula as follows: F Stage (X) = Activation(MaxPool2d(BN(Conv 1×1 (LeakyReLU(BN(PConv 3×3 (X))))))) where F Stage (X) represents the Stage module itself, X represents the input image, PConv 3×3 represents a 3x3 partial convolution, BN represents normalization, LeakyReLU represents an activation function, Conv 1×1 represents a 1x1 convolution, MaxPool2d represents a two-dimensional max pooling operation, and Activation represents a dynamic activation function. The feature interaction module is used to gradually expand the receptive field while capturing pixel-level lesion details and region-level semantic features by using the SPPELAN multi-level pooling design, thereby solving the problem of insufficient adaptability to multi-size targets; the feature interaction module uses the multi-scale feature fusion advantage of SPPELAN to replace the AIFI module in the RTDETR model, fuses feature maps at different stages of the feature extraction module through cross-level residual connection, compensates for the detail loss of deep semantic features by using shallow high-resolution information, enhances the information exchange and fusion between features, and significantly improves the robustness to leaf shape or occluded regions; The feature fusion module is used to integrate the information processed by the feature interaction module, the feature fusion module of the RTDETR model is redesigned by using the ASF-YOLO, the SSFF module is used to capture different grape leaf lesion shape features; then the TFE module is used to fuse multi-scale feature maps and enrich the detail information of the lesion; finally, the CPAM is introduced to integrate the SSFF module and the TFE module, realize adaptive channel weight distribution and spatial position focusing, improve the robustness of detection in the complex field environment, and enhance the representation ability of the target region; finally, the integrated feature map is subjected to IoU-aware QuerySelection to select a fixed number of features as target queries, and then subjected to Decoder&Head to map the confidence and the bounding box to obtain the final detection result, including the disease category and the position information.

4. The method for real-time detection of grape leaf diseases based on the improved RTDETR model according to claim 1, characterized in that, In step 2), through multiple rounds of iterative training, during which the model performance indicators are evaluated with the validation set, and focus on an important indicator FPS, which refers to the number of images that the model can process per unit time, closely related to real-time performance, the higher the FPS value, the faster the model response speed, by monitoring the evaluation indicators, to determine whether the model has overfitting phenomenon, when the performance indicators on the validation set no longer improve in continuous multiple training rounds, or the number of training rounds reaches the pre-set maximum number of rounds, stop training and save the model parameters, at this time, it is determined as the optimal model; Wherein, the formula of FPS is as follows: AllTime=pre+inf+post In the formula, ms represents millisecond, AllTime represents total time, pre represents preprocessing time Preprocess, inf represents inference time Inference, and post represents post-processing time Postprocess.

Citation Information

Patent Citations

  • Grape leaf disease detection method based on deep learning

    CN117315648A

  • Grape leaf disease detection method

    CN118864401A

  • Tomato leaf disease detection method based on SSP-DETR model

    CN119785157A

Cited By

  • Underwater target detection method based on RT-SonarNet

    CN121982507A