A Hierarchical Leaf Disease Detection Method and System Based on Deep Learning
Through a hierarchical apple leaf disease detection method based on deep learning, the backbone network and feature pyramid network are used to extract and fusion disease characteristics, and combine the pre- and post-region proposal networks to generate proposal boxes, solving the problem of low detection accuracy of apple leaf disease in the prior art, achieving higher detection accuracy and recall.
Patent Information
- Application Number
- CN202310353202.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-04
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-04-04
AI Technical Summary
In the prior art, the accuracy of apple leaf disease detection is low. Especially in complex natural environments, the loss of focus and blurred spots in the diseased area and the lack of obvious texture lead to poor detection results.
The hierarchical blade disease detection method based on deep learning is adopted to extract disease characteristics at different levels through the backbone network, and feature fusion and extraction are combined with the feature pyramid network to generate a pyramid feature map. Then, the front area proposal network is used to generate a blade proposal box, the rear area proposal network refines the disease proposal box, and finally the disease classification and positioning is performed through the detection head of the area of interest.
It improves the recall and detection accuracy of apple leaf disease detection in complex natural environments, and can more accurately identify and locate diseases on apple leafs.
Smart Images

Figure CN116664480B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a hierarchical leaf disease detection method and system based on deep learning. Background Art
[0002] Traditional apple disease detection mostly relies on the personal experience diagnosis of farmers, with poor reliability and timeliness. Disease detection relying solely on manpower can no longer meet the needs of the rapid development of orchards. The development of deep learning has attracted wide attention to convolutional neural networks (CNNs). Using a convolutional neural network to replace the human visual function, an image sensor is used to obtain the image information of plant diseases, and then the image storage information is converted into a multi-dimensional matrix. By performing computational processing and recognition on the data, it is possible to accurately analyze the main features in the collected targets, extract the effective information therein, and then reasonably monitor the growth state of plants. However, the growth environment of apples in natural orchards is relatively complex. Due to the vast orchard, in the images of dense disease areas obtained by the camera, a large number of out-of-focus and blurred disease spots are extremely likely to appear. At the same time, the disease spots of some diseases occupy a small number of pixels and the texture information is not obvious. Therefore, the complex environment will seriously affect the accuracy of apple leaf disease spot detection. Summary of the Invention
[0003] The purpose of the present invention is to overcome the above-mentioned shortcomings of the prior art and provide a hierarchical leaf disease detection method and system based on deep learning to solve the problem of low accuracy of leaf disease spot detection in the prior art.
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] A hierarchical leaf disease detection method based on deep learning, comprising:
[0006] Input the image to be detected, extract disease features at different levels through a backbone network, perform feature fusion and feature extraction on the disease features at different levels through a feature pyramid network to obtain a pyramid feature map;
[0007] Input the pyramid feature map into a pre-region proposal network to obtain disease proposal boxes and leaf proposal boxes of all categories; the pre-region proposal network consists of 5 feature extraction convolutional layers with a size of 3×3;
[0008] In each leaf proposal box, prefabricated anchor boxes are generated by a position anchor box generator, and the center coordinates and sizes of the prefabricated anchor boxes are regressed through a subsequent region proposal network to obtain the regression parameters of the position and the disease categories, and disease proposal boxes are generated; the subsequent region proposal network includes an ROI Align module, an attention module, and a feature convolutional layer; the ROI Align module obtains diseased leaf features from the leaf proposal boxes; the attention module adjusts the weights in the diseased leaf features to obtain the adjusted diseased leaf features; the feature convolutional layer obtains disease proposal boxes based on the adjusted diseased leaf features;
[0009] After all the proposal boxes are extracted for features through the region of interest detection head, the classification and position of the diseases and diseased leaves are obtained.
[0010] A further improvement of the present invention lies in:
[0011] Preferably, the pyramid feature map generates prefabricated anchor boxes in the front region proposal network, and the center position and width and height of the threshold anchor boxes are regressed through 4 convolutions with a size of 1×1 to obtain proposal boxes; whether there are diseases or leaves in the proposal boxes is discriminated through 1 convolution with a size of 1×1.
[0012] Preferably, low-level features are extracted from the front region network to obtain a low-level feature aggregation module; the formula of the low-level feature aggregation module is:
[0013]
[0014] where deconv represents a deconvolution layer, maxpool represents a max pooling layer, L 1 and L 0 are both feature maps.
[0015] Preferably, the ROI Align module obtains diseased leaf features based on the leaf proposal boxes and the low-level feature aggregation module.
[0016] Preferably, the subsequent region proposal network obtains the regression parameters of the position and the disease categories, adjusts the offset of the prefabricated anchor boxes, and the grid generation formula is:
[0017]
[0018] where P w,h is the width and height of N leaf proposal boxes generated by the front region proposal network, ROI w,h is the width and height of the ROI Align module in the horizontal and vertical directions, and S i,j represents the grid point step size in the horizontal and vertical directions.
[0019] Preferably, the region of interest detection head extracts position features from the disease proposal boxes. After aligning the position features, the aligned position features are fed into a fully connected layer for re-regression of the disease proposal boxes to obtain the final category of the disease and the final bounding box.
[0020] Preferably, both the front region proposal network and the rear region proposal network are trained by backpropagation through a loss function to obtain the trained front region proposal network and rear region proposal network.
[0021] Preferably, the loss function is:
[0022]
[0023] τ represents the level of the region proposal network, and τ is equal to 2; is the stage regression loss, is the classification loss; is the number of pre-set anchor boxes, is equal to the number of images in a batch.
[0024] Preferably, after the region of interest detection head extracts features, the top 100 diseased leaf bounding boxes and disease bounding boxes with the highest scores are obtained. The bounding box with the highest score is obtained through non-maximum suppression to obtain the classification and location of the disease.
[0025] A hierarchical leaf disease detection system based on deep learning, comprising:
[0026] A feature extraction unit, configured to input an image to be detected, extract disease features at different levels through a backbone network, perform feature fusion and feature extraction on the disease features at different levels through a feature pyramid network, and obtain a pyramid feature map;
[0027] A diseased leaf generation unit, configured to input the pyramid feature map into a front region proposal network to obtain disease proposal boxes and leaf proposal boxes of all categories; the front region proposal network consists of 5 feature extraction convolutional layers with a size of 3×3;
[0028] A disease generation unit, configured to generate pre-set anchor boxes through a position anchor box generator in each leaf proposal box, and obtain the regression parameters of the position and the disease category by regressing the center coordinates and sizes of the pre-set anchor boxes through a rear region proposal network, and generate disease proposal boxes; the rear region proposal network includes an ROI Align module, an attention module, and a feature convolutional layer; the ROI Align module obtains diseased leaf features from the leaf proposal boxes; the attention module adjusts the weights in the diseased leaf features to obtain adjusted diseased leaf features; the feature convolutional layer obtains disease proposal boxes based on the adjusted diseased leaf features;
[0029] The classification detection unit is used to obtain the classification and location of diseases and diseased leaves after extracting features from all proposed boxes through the region of interest detection head.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] The present invention proposes a hierarchical apple leaf disease detection method based on deep learning for apple leaf disease detection in natural environments. First, the front-end region proposal network generates proposed boxes in the entire image and filters out leaf proposed boxes, and the back-end region proposal network generates lesion proposed boxes based on the leaf proposed boxes. Secondly, a low-level feature aggregation module is designed to better utilize the bridging features generated by the front-end region proposal network. Then, multi-level ROI Align blocks and GCNet are introduced into the back-end region proposal network to scale the aggregated features to the same size and focus more on lesions. Finally, a position anchor box generator is proposed to make it easier for the preset anchor boxes to capture the target lesions according to the position of the diseased leaves. In a complex natural environment, this hierarchical apple leaf disease detection method can improve the recall rate and detection accuracy of the detection task.
[0032] The present invention proposes a hierarchical apple leaf disease detection method based on deep learning. Aiming at the problem of low recall rate and accuracy in apple leaf disease detection under complex natural backgrounds, a hierarchical apple leaf disease detection method is proposed, which includes the hierarchical detection idea of the front-end region proposal network and the back-end region proposal network, and designs a low-level feature aggregation module and a position anchor box generator. The method aims to generate lesion proposed boxes based on the leaf proposed boxes to improve the final detection performance of apple leaf diseases in complex natural environments. The low-level feature aggregation module makes full use of the lesion semantic information, the max pooling layer maximizes the response of the corresponding region, prevents strong responses from being weakened by surrounding neurons, and the deconvolution layer retains more semantic information. At the same time, the position anchor box generator generates dense anchor boxes in the leaf proposed boxes, making it easier to capture the true value annotation boxes of the lesions and increasing the number of positive samples in the training samples. The experimental results show that this method achieves 49.0% AR 100 and 34.0% mAP on the test set. It can achieve competitive performance in the apple leaf detection task in complex environments and has broad application prospects in agricultural actual production.
[0033] In addition, the present invention has the following beneficial effects:
[0034] Furthermore, the present invention designs and optimizes the existing computer vision image method, establishes a hierarchical detection framework, and through the cascaded hierarchical detection structure of the front-end region proposal network and the back-end region proposal network, enables the diseased leaves and lesions to be detected step by step hierarchically, effectively improving the apple leaf disease detection effect in complex natural environments.
[0035] Furthermore, the present invention constructs a low-level feature fusion module. The low-level feature fusion module aggregates the features with high response to the lesion features, makes full use of the semantic information of the lesions, and significantly improves the real-time diagnosis effect of apple lesions.
[0036] Furthermore, the present invention constructs a position anchor box generator. Based on the leaf proposal boxes generated by the previous region proposal network, dense pre-set anchor boxes are generated to better capture the annotation boxes of the lesions, increase the proportion of positive samples of the lesions, and contribute to network training. Description of the Drawings
[0037] Figure 1 is the flowchart of the dataset preprocessing of the present invention;
[0038] Figure 2 is the flowchart of the model construction of the present invention;
[0039] Figure 3 is the flowchart of the apple leaf disease diagnosis of the present invention;
[0040] Figure 4 is the diagnosis result diagram of the present invention: (a) is Alternaria leaf blotch and diseased leaves; (b) is rust and diseased leaves; (c) is green peach aphid; (d) is mosaic disease; (e) is brown spot. Detailed Embodiments
[0041] The following further describes the present invention in detail with reference to the drawings and specific embodiments.
[0042] Embodiment 1
[0043] One embodiment of the present invention discloses a hierarchical apple leaf disease detection method based on deep learning, which can diagnose the specific categories of five apple leaf diseases and two types of diseased leaves, and at the same time accurately locate the positions of the lesions; this method takes advantage of the deep learning technology in image processing, extracts the disease image features through operations such as convolution, pooling, and activation, and effectively realizes the detection of apple leaf diseases.
[0044] The present invention designs and implements a hierarchical apple leaf disease detection method on the Ubuntu platform. The whole method consists of multiple sub-methods, including dataset preprocessing, leaf proposal generation of the previous region proposal network, lesion proposal generation of the subsequent region proposal network, and low-level feature aggregation module, etc.
[0045] The method specifically includes the following steps:
[0046] Step 1, input the image to be detected into the whole model.
[0047] Step 2, identify the disease
[0048] Step 2.1, Obtain pyramid features
[0049] Input the image to be detected into the ResNet50 backbone network. Through the backbone network, feature maps with different receptive fields are extracted. ResNet50 has 5 stages. Obtain the output of the last convolutional layer in the 5 stages. Input the output feature maps of the last convolutional layer of each stage into the Feature Pyramid Network; the Feature Pyramid Network performs feature fusion and feature extraction on these 5 feature maps to generate 5 pyramid feature maps and obtain 5 pyramid features.
[0050] Step 2.2, Obtain leaf proposal boxes through the pre - region proposal network.
[0051] In Step 2.2, input the 5 pyramid features generated in Step 2.1 into the pre - region proposal network. Through the pre - region proposal network, generate proposal boxes for all targets, and then filter out the proposal boxes related to leaves. The specific process is as follows:
[0052] For the 5 pyramid feature maps, generate pre - set anchor boxes corresponding to the original ratio. The pre - region proposal network extracts features through 5 convolutional layers with a size of 3×3 for feature extraction. Regress the center position, width, and height of the first pre - set anchor box through 4 convolutional layers with a size of 1×1 to obtain proposal boxes; generate the confidence of the existence of objects in the proposal boxes through 1 convolutional layer with a size of 1×1. Here, the objects specifically refer to diseases and diseased leaves.
[0053] Then obtain all the disease proposal boxes and leaf proposal boxes on the image to be detected, filter out all the leaf proposal boxes, and define the proposal boxes with an area greater than 60×60 pixels as leaf proposal boxes according to the average area of the true annotation boxes of diseased leaves in the leaf proposal boxes. Each proposal box is a 5 - dimensional matrix [C, X1, Y1, X2, Y2], where X1, Y1 represent the upper - left coordinates of the proposal box, and X2, Y2 represent the lower - right coordinates of the proposal box, and C represents the confidence of the existence of the target.
[0054] In Step 2.3, reuse the feature maps generated by the 3×3 convolutional layer for feature extraction in the pre - region proposal network in Step 2.2, which is called the bridging feature. The lower layer of the bridging feature responds more to disease spots, and the low - level features play an important role in improving the performance of disease spot detection. To make full use of the low - level semantic features, select the 0th and 1st layer features in the lower layer of the feature map. The specific details of the low - level feature aggregation module are as follows: Bridging feature map L 0 Pass through a max - pooling layer with a kernel size of 3×3 and padding of 1, and the obtained feature map is added to the feature map L 1 to obtain F' 1 ; Bridging feature map L 1After passing through a transposed convolutional layer with a kernel size of 3×3 and a padding of 1, the resulting feature map is added to the feature map L 0 to obtain F'. 0 The feature F' passing through the low-level feature aggregation module 0 and F' 1 will be sent to the progressive module for further processing. The feature aggregation formula is as follows:
[0055]
[0056] where deconv represents the transposed convolutional layer and maxpool represents the max pooling layer.
[0057] Step 2.4 Based on the leaf proposal boxes generated by the previous region proposal network, through the position anchor box generator, refine the proposal boxes of leaf disease spots to further obtain more disease proposal boxes.
[0058] Specifically, within the position coordinates of each leaf proposal box obtained in Step 2.2, generate smaller preset anchor boxes through the position anchor box generator, and calculate the regression parameters and the confidence of whether there is a disease for the preset anchor boxes through the subsequent region proposal network to generate disease proposal boxes.
[0059] The subsequent region proposal network is a convolutional neural network, including the ROI Align module, the attention module, and the feature convolutional layer.
[0060] In the subsequent region proposal network, sort the 5D matrix in the leaf-related proposal boxes generated in Step 2.2 according to the confidence of whether there is an object, that is, sort according to the above C, to obtain the top 50 leaf proposal boxes. The subsequent region proposal network is based on the leaf proposal boxes and the aggregated feature F' 0 and F' 1 , and uses a multi-level ROI Align module to combine the four-dimensional position matrix [X1, Y2, X2, Y2] to extract the features corresponding to diseased leaves on the aggregated feature F' 0 and F' 1 and scale them to the same scale. The stride S of the multi-level ROI Align module is set to 20.
[0061] Input the diseased leaf features of the same scale into the attention mechanism module. The attention mechanism module processes the diseased leaf features of the same scale, adjusts the weights in the diseased leaf features, and makes the leaf features of the same scale more focused on the small disease spots with colors significantly different from the leaves.
[0062] Input the adjusted diseased leaf features into the feature convolutional layer. The feature convolutional layer of the subsequent region proposal network contains two heads, and each head further includes a classification layer and a regression layer. Output the disease proposal boxes through the feature convolutional layer. The convolutional layer performs regression on the preset anchor boxes to obtain regression parameters. Calculate the disease proposal boxes by passing the disease anchor boxes through the regression parameters.
[0063] The convolutional feature extraction network in the subsequent region proposal network also adopts 4 convolutions with a size of 1×1 and 1 convolution with a size of 1×1 to fine-tune the small anchor boxes generated by the position anchor box generator. Among them, 4 regression convolutions with a size of 1×1 calculate the offset of the center position and width and height of the preset anchor boxes, and 1 classification convolution with a size of 1×1 calculates the confidence of whether there is a disease in the small anchor boxes.
[0064] Step 2.6 Save the disease and diseased leaf proposal boxes generated by the front region proposal network and the subsequent region proposal network. Select the top 500 proposal boxes according to the confidence for each layer and send them into the region of interest detection head.
[0065] The region of interest detection head contains an ROI Align module and two fully connected layers. Based on the position information of the proposal boxes, the ROI Align module extracts the corresponding features in the bridging features and performs feature alignment. Each disease and diseased leaf proposal box corresponds to a 7×7 feature map. The aligned features are unfolded into one dimension with a size of 1×49 and sent into two fully connected layers. One fully connected layer performs re-regression on the proposal boxes to obtain the bounding boxes. The other fully connected layer performs fine classification of the diseases and diseased leaves to obtain the specific categories of the diseases and diseased leaves. The outputs of the above two fully connected layers also form a 5-dimensional matrix [C, X1, Y2, X2, Y2], corresponding to the category information and position information of the bounding boxes.
[0066] Embodiment 2
[0067] The embodiment of the present invention also discloses a hierarchical apple leaf disease detection system based on deep learning, including:
[0068] A feature extraction unit, which includes a backbone network and a pyramid network. The backbone network is used to extract 5 feature maps with different receptive fields from the picture to be detected. The backbone network inputs the feature maps into the pyramid network. After the pyramid network performs feature fusion and feature extraction on the feature maps with different receptive fields, 5 pyramid features are obtained.
[0069] A diseased leaf generation unit, which is used to generate disease proposal boxes and leaf proposal boxes through the front region proposal network and screen out the leaf proposal boxes among them.
[0070] A disease generation unit, which is used to generate disease proposal boxes through the subsequent region proposal network and the position anchor box generator.
[0071] The classification detection unit inputs the disease proposal boxes generated by the diseased leaf generation unit and the disease generation unit into the region of interest detection head to obtain the positions of the proposal boxes, extracts the features at the corresponding positions. After feature alignment, the features are sent to the fully connected layer to obtain the final disease classification and positions.
[0072] Example 3
[0073] One embodiment of the present invention discloses a construction method of a hierarchical apple leaf disease detection system based on deep learning.
[0074] Firstly, the dataset preprocessing simulates the influence of the shooting angle factor in the actual production scenario on the collected apple disease images. Secondly, the front region proposal network generates leaf proposal boxes based on the feature maps at 5 scales. Then, the rear region proposal network generates disease spot proposals according to the leaf proposal boxes. Finally, the region of interest detection head performs different sizes of disease spot types and positions based on the position information and corresponding features of the proposal boxes.
[0075] The construction process of the hierarchical apple leaf disease detection system based on deep learning mainly includes 4 steps: dataset preprocessing, model construction, model training, and disease detection. See Figure 3 . After completing these 4 steps, the final diagnosis result is obtained. The method specifically includes the following steps:
[0076] Step 1, Dataset preprocessing
[0077] Dataset preprocessing refers to constructing the dataset format for object detection, which includes three parts: images, disease annotation information on the images, and dataset division. See Figure 2 , and the entire dataset preprocessing can be divided into three sub-steps: video data frame splitting, data annotation, and data division.
[0078] Step 1.1 Video data frame splitting. Collect apple leaf disease videos, including 5 common apple leaf diseases and 2 types of apple diseased leaves, a total of 7 types of diseases. The diseases are Alternaria leaf spot, rust, green peach aphid, mosaic disease, and brown spot disease respectively. The diseased leaves are Alternaria leaf spot diseased leaves and rust diseased leaves respectively. There are a total of 257 videos, which are divided into 30921 frames of apple leaf disease images.
[0079] Step 1.2 Data annotation. Annotate the location and category of diseases in the images, including the location coordinates, category, image name, and size of the diseases, to form the disease information for each image. Store the disease information for each image in its respective XML document. All the diseased images and the corresponding disease information for each image form a dataset. The diseased images are RGB three-channel data, and are annotated using the CVAT online semi-automatic annotation and manual post-correction methods to mark the disease type and location. With the help of botanical experts, the specific category of apple diseases can be determined for difficult-to-distinguish samples.
[0080] Step 1.3 Data division. During the experiment, the image data is randomly divided into a training set and a validation set in a ratio of 7:1, and the file information is stored in the form of txt. Each line records an image name.
[0081] Step 2, Model construction
[0082] First, extract disease features at different levels based on the backbone network ResNet50, and the Feature Pyramid Network performs feature fusion on the feature maps at different levels. Secondly, a Region Proposal Network is preposed. Anchor boxes are set at the center of each cell of the disease feature map, and then regression is performed. Proposal boxes are generated for all diseases and all diseased leaves, and the proposal boxes related to the leaves are filtered out. Then a Region Proposal Network is postposed. Based on the leaf-related proposal boxes generated by the preposed Region Proposal Network, the candidate proposal boxes of apple disease spots are refined. At the same time, within the leaf-related proposal boxes, a position anchor box generator generates smaller and denser preset anchor boxes, and the postposed Region Proposal Network regresses the center coordinates and sizes of the anchor boxes, calculates the regression parameters and target categories, and generates the related proposal boxes of apple disease spots. Finally, based on the position information of the proposal boxes, the Region of Interest Detection Head extracts the features at the corresponding positions and sends them into two fully connected layers for final class classification and anchor box regression respectively. See Figure 3 , and the specific steps of feature extraction are as follows:
[0083] Step 2.1 Input the apple leaf disease dataset generated in Step 1 into the ResNet50 backbone network. Through the backbone network, feature maps with different receptive fields are extracted to learn disease-related features. ResNet50 has 5 stages, and the output of the last convolutional layer in the 5 stages is obtained. The output feature maps of the last convolutional layer in each stage are input into the Feature Pyramid Network. The Feature Pyramid Network performs feature fusion and feature extraction on these 5 feature maps to generate 5 pyramid feature maps and obtain 5 pyramid features.
[0084] Step 2.2 Send the 5 pyramid features generated in Step 2.1 into the preposed Region Proposal Network. Through the preposed Region Proposal Network, proposal boxes are generated for all targets, and then the leaf proposal boxes are filtered out. The specific process is as follows:
[0085] For 5 pyramid feature maps, corresponding preset anchor boxes are generated at the original ratio. The prior region proposal network extracts features through 5 convolutional layers with a size of 3×3 for feature extraction. The center position, width, and height of the first preset anchor box are regressed through 4 convolutional layers with a size of 1×1 to obtain proposed boxes. Whether there is an object is determined through a convolutional layer with a size of 1×1 for the proposed boxes, obtaining a 5-dimensional matrix [C, X1, Y2, X2, Y2]. Among them, X1 and Y1 represent the upper left coordinates of the proposed box, X2 and Y2 represent the lower right coordinates of the proposed box, and C represents the confidence level of the existence of the target.
[0086] The prior region proposal network extracts proposed boxes for all diseases and diseased leaves, and then filters out the proposed boxes related to leaves. According to the average area of the true annotation boxes of these leaves, a proposed box with an area greater than 60×60 pixels is defined as a leaf proposed box. Considering that proposed boxes of other diseases may also be filtered out, but in a training batch, the ratio of positive and negative samples during anchor box sampling is 1:3, so the proposed boxes of other diseases are used as negative samples. As long as the leaf proposed boxes are filtered out, the neural network training can proceed normally in step 3.
[0087] In step 2.3, the feature maps generated by the 3×3 convolutional layers for feature extraction in the prior region proposal network in step 2.2 are reused, which is called bridge features. The lower layers of the bridge features respond more to disease spots, and the low-level features play an important role in improving the detection performance of disease spots. In order to make full use of the low-level semantic features, the features of layer 0 and layer 1 in the lower-level feature maps are selected. The specific details of the low-level feature aggregation module are as follows: Bridge feature map L 0 Pass through a max-pooling layer with a kernel size of 3×3 and a padding of 1, and the obtained feature map is added to the feature map L 1 to obtain F' 1 ; Bridge feature map L 1 Pass through a transposed convolutional layer with a kernel size of 3×3 and a padding of 1, and the obtained feature map is added to the feature map L 0 to obtain F' 0 . The feature F' 0 and F' 1 obtained after passing through the low-level feature aggregation module will be sent to the progressive module for the next step of processing. The feature aggregation formula is as follows:
[0088]
[0089] Among them, deconv represents the transposed convolutional layer, and maxpool represents the max-pooling layer.
[0090] Step 2.4 then refines the candidate proposal boxes of apple lesions based on the leaf-related proposal boxes generated by the front region proposal network through a post-region proposal network. Within the leaf-related proposal box position coordinates [X1, Y2, X2, Y2], the position anchor box generator generates smaller and denser pre-set anchor boxes, regresses the center coordinates and size of the anchor boxes through the convolutional neural network, calculates the regression parameters and target categories, and generates relevant proposal boxes for apple lesions.
[0091] In the post-region proposal network, the position information matrix of the leaf-related proposal box generated in step 2.2 is sorted according to the confidence of whether there is an object, and the top 50 leaf proposal boxes are obtained. The post-region proposal network is based on the leaf proposal box and the aggregated feature F' 0 and F' 1 , using the multi-level ROI Align module combined with the position matrix [X1, Y2, X2, Y2], in the aggregate feature F' 0 and F' 1 The features corresponding to the diseased leaves are extracted and scaled to the same scale. The step size S of the multi-level ROI Align module is set to 20.
[0092] The spots on the diseased leaves show different color and texture attributes compared to other parts of the leaves. The attention module can make the leaf features of the same scale more focused on small spots whose colors are obviously different from the leaves. The Global Correlation Network (GCNet) adopts the mechanism of the NL module in context information modeling, which can make full use of global context information. At the same time, it references the SE block in the conversion stage. The detailed architecture of GCNet can be written as Formula 4.
[0093]
[0094] Where i is the index of the query location and j enumerates all possible locations. (·) Represents a linear transformation matrix (e.g. 1x1 convolution). is the weight of global attention, δ(·)=W v2 ReLU(LN(W v1 (·))) indicates bottleneck transformation.
[0095] The convolutional layer of the post-region proposal network contains two heads, each of which also includes a classification layer and a regression layer. In addition, the weights of the convolutional layer are shared, and the convolutional feature extraction network in the post-region proposal network also uses 4 1×1 convolutions and 1 1×1 convolution to fine-tune the small anchor box generated by the position anchor box generator.
[0096] Among them, the position anchor box generator generates anchor boxes based on the position information contained in the leaf proposal boxes. In the position anchor box generator, the length of the grid is set to the integer part of the width and height in the ROI Align module divided by the width and height of the leaf-related proposal, and the center position of the anchor point is the offset of the grid point corresponding to the upper right corner coordinates (X1, Y1) of each proposal box. The grid point generation process is as follows:
[0097]
[0098] Where P w,h is the width and height of the N leaf proposal boxes generated by the front region proposal network, and ROI w,h is the width and height of the ROIAlign module in the horizontal and vertical directions, and S i, represents the grid point step size in the horizontal and vertical directions.
[0099] Step 2.6 saves the disease proposal boxes generated by the front region proposal network and the rear region proposal network and sends them into the region of interest detection head. Based on the position information of the proposal boxes, the region of interest detection head extracts the features at the corresponding positions and aligns them. The aligned features are sent into two fully connected layers to perform the final fine classification of the disease and diseased leaves of the anchor box and the regression of the anchor box again, and obtain the bounding boxes with specific disease and diseased leaf categories.
[0100] Through the above process, a new hierarchical apple leaf disease detection model is obtained. This model uses the ResNet50 backbone network, the front region proposal network, the rear region proposal network and the region of interest detection head to extract features from apple disease pictures in complex natural scenes, generate a series of preset anchor boxes, classify and regress the anchor boxes, and then can predict the disease category and the position of the disease spot.
[0101] Step 3, model training
[0102] For the hierarchical detection structure, the ground truth annotation boxes to be matched by different region proposal networks are screened. The data in the training set are input into the neural network model batch by batch until the data in the training set are trained and both the loss function and the prediction accuracy of the model tend to be stable, and the final neural network model is obtained.
[0103] The apple leaf disease dataset divided in Step 1 is sent into the model constructed in Step 2 batch by batch for training to generate a model weight file that fits the apple leaf disease dataset. During the process, the images in the validation set are used to monitor the fitting situation of the model weights to the annotation boxes during the training process to avoid phenomena such as underfitting or overfitting. The specific training details are as follows:
[0104] In the training phase of the present invention, the optimizer adopts the stochastic gradient descent method, and the momentum is set to 0.9. The initial learning rate for the first 100,000 iterations is 10 -4 , and the learning rate drops to 10 -5 after 100,000 to 150,000 iterations, and drops to 10 -6 after 150,000 iterations, with a total of 200,000 iterations.
[0105] The obtained gradient values will perform backpropagation on the backbone network, feature pyramid network, pre-region proposal network, post-region proposal network, and region of interest detection head simultaneously, updating the weights of the entire network to make the final output of the network gradually approach the annotation box and the true category.
[0106] Step 3.1 Hierarchical detection target matching
[0107] To achieve the purpose of hierarchical detection, a hierarchical target class list and an annotation box filter are added to the hierarchical detection framework. The annotation box filter can obtain the annotation boxes that each level of the region proposal network needs to match according to the target class list of different layers. In the experiment, the pre-region proposal network matches the annotation boxes of seven categories, and the post-region proposal network matches the annotation boxes of spot leaf blight and rust. The two layers of annotation boxes will be used to calculate the classification loss and regression loss of the first layer and the second layer respectively.
[0108] Step 3.2 Default box and annotation box matching
[0109] Step 3.2.1 For each annotation box in the training set image, select the pre-set anchor box with the largest IOU to match it, which is the positive sample. IOU is the result obtained by dividing the overlapping part of two boxes A and box B by the set part of the two regions. The calculation formula is as follows:
[0110]
[0111] For the remaining unmatched pre-set anchor boxes, if the IOU with any annotation box is greater than the threshold, then the pre-set anchor box matches the corresponding annotation box. The anchor box with an IOU value higher than 0.5 with the annotation box is a positive sample, and the label is 1 during training. The anchor box with an IOU value lower than 0.3 with the annotation box is a negative sample, and the label is 0 during training. Each layer of the region proposal network performs positive and negative sample matching according to the annotation boxes screened by the hierarchical target class list.
[0112] Step 3.2.2 To ensure that the positive and negative samples are as balanced as possible, sample the negative samples. When sampling, sort them in descending order according to the confidence error (the smaller the confidence of predicting the background, the greater the error), and select the top-k with larger errors as the training negative samples to ensure that the ratio of positive and negative samples is close to 1:3.
[0113] Step 3.3 Calculate the loss function
[0114] The annotation boxes in the front region proposal network match seven categories, and the annotation boxes in the rear region proposal network match the annotation boxes of apple scab and rust. The two layers of annotation boxes will be used to calculate the classification loss and regression loss in the first layer and the second layer respectively.
[0115] The loss function is mainly used in the training stage of the model. After the training data of each batch is sent into the model, the predicted values are output through forward propagation. Then the loss function will calculate the difference value between the predicted values and the true values, that is, the loss value. The predicted values refer to the proposed boxes and categories, and the true values are the annotation boxes and annotation categories. After obtaining the loss value, the model updates each parameter through backpropagation to reduce the loss between the true value and the predicted value, so that the predicted values generated by the model approach the true value, thus achieving the purpose of learning.
[0116] The network needs to calculate two parts of losses, namely the region proposal network loss and the region of interest detection head loss.
[0117] The region proposal network loss is as follows: Specifically, in the process of proposed box sampling, the top K proposed boxes with the highest confidence are sent to the region of interest detection head for further training. In the two-layer region proposal network, K is set to 500. The entire region proposal network can be trained with the backpropagation of the multi-task loss. The specific formula is as follows:
[0118]
[0119] τ represents the layer of the region proposal network. In the present invention, τ is equal to 2. is the stage regression loss, is the classification loss. is the number of preset anchor boxes, is equal to the number of images in the batch. These two parameters are used for normalization. Because in the front region proposal network, is about 10 times of, so λ is set to 10. In the rear region proposal network, is about 28 times of, so λ 1 is set to 28 to ensure the balance of the two loss terms and enable the network to perform balanced training on the regression and classification tasks. is about times of 2 is set to 28 to ensure the balance of the two loss terms and enable the network to perform balanced training on the regression and classification tasks.
[0120] In implementation, the GIoU loss and binary cross-entropy loss are respectively used as the regression loss and classification loss. For two anchor boxes A and B, the minimum convex set C (the minimum bounding box of A and B) can be calculated. With the minimum convex set, the GIOU loss can be calculated:
[0121]
[0122]
[0123] The binary cross-entropy loss is shown in formula (8), where y i is either 1 or 0, representing the presence or absence of a target in the proposed box. p(y i ) is the confidence that the proposed box predicts the presence of a target. n represents the total number of predicted boxes.
[0124]
[0125] The loss of the region of interest detection head is relatively simple. The regression loss is the same as that of the region proposal network regression loss, and the GIOU loss is adopted. λ also represents the loss balancing term, which is set to 10. The classification loss involves specific category calculations. Sj is the j-th value of the output vector S of the fully connected layer, representing the probability that the predicted sample belongs to the j-th category. y j represents the true category of the labeled box, and there are a total of 7 categories.
[0126]
[0127]
[0128]
[0129] Step 4, disease diagnosis and prediction.
[0130] Using the hierarchical detection model generated in Step 3, input the disease image to obtain the category and location information of five apple leaf diseases and two diseased leaves, and after a series of post-processing, display the position and disease category of the disease box in the disease image. The specific process is as follows:
[0131] Step 4.1 Input the image to be diagnosed into the backbone network. Through the feature pyramid network, the front and rear region proposal networks, generate proposed boxes hierarchically. The region of interest detection head classifies and re-regresses the proposed boxes to obtain the top 100 bounding boxes with the highest scores corresponding to each disease and diseased leaf category, that is, the position information and category information of the bounding boxes.
[0132] Step 4.2 Non-maximum suppression. Process the bounding boxes in Step 4.1 through non-maximum suppression to further screen the bounding boxes and make the detection interface cleaner. In the present invention, the non-maximum suppression threshold is set to 0.6. The specific process is as follows:
[0133] Traverse the bounding boxes of each class, and calculate the intersection over union (IoU) between the bounding box with the highest confidence in this class and the remaining bounding boxes respectively. If the IoU is greater than the set threshold of 0.6, it means that the two boxes may represent the same target, then discard the low-score bounding box and retain the box with a higher score. Finally, find the box with the largest score from the remaining bounding boxes, and repeat the above process.
[0134] Step 4.3 Stack the bounding boxes generated in Step 4.2 according to the class labels and confidence levels, sort them according to the confidence threshold, and filter out the bounding boxes with a confidence level higher than the threshold. In the present invention, the confidence threshold is set to 0.5, and finally, the disease class and the location of the lesion are marked on the image.
[0135] Example 4
[0136] The detection of apple leaf diseases at the layer level requires four steps: dataset preprocessing, model construction, model training, and disease prediction, which are implemented based on the Ubuntu platform. (1) Dataset preprocessing: First, the apple leaf disease videos collected in natural scenes are framed to obtain disease image data. Then, the target objects in the apple leaf disease images are labeled, and the position coordinates, categories, image names, widths, and heights of the target objects are stored in an XML document. During the training process, the training set and the validation set are divided according to a ratio of 7:1. (2) Model construction: First, a batch of data in the training set is input into the ResNet50 backbone network, and feature maps with different disease characteristics are extracted through the backbone network. The output feature maps of the last convolutional layer of the 5 stages of ResNet are all input into the Feature Pyramid Network to generate 5 pyramid feature maps. Then, through a pre-region proposal network, default anchor boxes are set at the center of each cell of the disease feature map for default anchor box classification and regression. This pre-region proposal network extracts the proposal boxes of all classes of diseases, and then filters out the proposal boxes related to the leaves. Subsequently, through a post-region proposal network, based on the leaf-related proposal boxes generated by the pre-region proposal network, the candidate proposal boxes of apple lesions are refined. Smaller and denser preset anchor boxes are generated through the position anchor box generator, and at the same time, the regression parameters and target classes are calculated to generate the proposal boxes related to apple lesions. Finally, the Region of Interest (ROI) detection head extracts the features at the corresponding positions based on the position information of all the proposal boxes and aligns their features. The aligned features are fed into two fully connected layers for final class classification and anchor box regression respectively. (3) Model training: The disease dataset training set constructed by dataset preprocessing is divided into small batches and input into the hierarchical detection model in turn. According to the hierarchical detection list, the labeled boxes to be matched by each hierarchical region proposal network are assigned. Secondly, the positive and negative samples are divided according to the Intersection over Union (IOU) size between each labeled box and the prior box in the image, the loss function is calculated, and the parameters of the detection model are updated. (4) Disease prediction: First, load the weight file of the hierarchical detection model generated after training and input the apple leaf disease image in the natural scene. According to the forward-propagated data of the model, calculate the center position offset and width-height offset of each preset anchor box, and then convert them into the format of the upper-left and lower-right coordinates. Then, through non-maximum suppression with a threshold set to 0.6, the boxes with a score higher than 0.5 are selected as the final prediction results and visualized in the image, such asFigure 4 as shown
[0137] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hierarchical leaf disease detection method based on deep learning, characterized in that, it includes: Input the image to be detected, extract disease features at different levels through the backbone network, perform feature fusion and feature extraction on the disease features at different levels through the Feature Pyramid Network to obtain a pyramid feature map; Input the pyramid feature map into the pre-region proposal network to obtain disease proposal boxes and leaf proposal boxes of all categories; The pre-region proposal network consists of 5 feature extraction convolutional layers with a size of 3×3; Extract low-level features from the pre-region network to obtain a low-level feature aggregation module; The formula of the low-level feature aggregation module is: Among them, F 0 ′ and F 1 ′ are both underlying feature aggregation modules, deconv represents a transposed convolutional layer, maxpool represents a max pooling layer, L 1 and L 0 are both feature maps; In each leaf proposal box, prefabricated anchor boxes are generated by a position anchor box generator, and the center coordinates and sizes of the prefabricated anchor boxes are regressed through a subsequent region proposal network to obtain the regression parameters of the position and the disease categories, and disease proposal boxes are generated; the subsequent region proposal network includes an ROI Align module, an attention module, and a feature convolutional layer; the ROI Align module combines a four-dimensional position matrix on the aggregated features F 0 ′ and F 1 ′ to obtain diseased leaf features from the leaf proposal boxes; the attention module adjusts the weights in the diseased leaf features to obtain the adjusted diseased leaf features; the feature convolutional layer obtains disease proposal boxes based on the adjusted diseased leaf features; The post-region proposal network obtains the regression parameters of the position and the disease category, adjusts the offset of the prefabricated anchor box, and the grid generation formula is: Among them, P w,h is the width and height of N blade proposal boxes generated by the pre-region proposal network, and ROI w,h is the width and height of the ROIAlign module in the horizontal and vertical directions, and S i,j represents the grid point step size in the horizontal and vertical directions; After all the proposal boxes extract features through the Region of Interest (ROI) detection head, the classification and position of the disease and diseased leaves are obtained.
2. The hierarchical leaf disease detection method based on deep learning according to claim 1, characterized in that, Prefabricated anchor boxes are generated in the pre-region proposal network for the pyramid feature map, and the center position and width and height of the threshold anchor box are regressed through 4 convolutional layers with a size of 1×1 to obtain proposal boxes; and whether there is a disease or a leaf in the proposal box is discriminated through 1 convolutional layer with a size of 1×1.
3. The hierarchical leaf disease detection method based on deep learning according to claim 1, characterized in that, The ROI Align module obtains diseased leaf features based on the leaf proposal box and the low-level feature aggregation module.
4. The hierarchical leaf disease detection method based on deep learning according to claim 1, characterized in that, The Region of Interest (ROI) detection head extracts position features from the disease proposal box. After aligning the position features, the aligned position features are sent to a fully connected layer for re-regression of the disease proposal box to obtain the final category and the final bounding box of the disease.
5. The hierarchical leaf disease detection method based on deep learning according to claim 1, characterized in that, Both the pre-region proposal network and the post-region proposal network are trained by backpropagation through a loss function to obtain the trained pre-region proposal network and post-region proposal network.
6. The hierarchical leaf disease detection method based on deep learning according to claim 5, characterized in that, The loss function is: τ represents the level of the Region Proposal Network, and τ is equal to 2; is the stage regression loss, is the classification loss; is the number of pre-set anchor boxes, which is equal to the number of images in the batch.
7. The hierarchical leaf disease detection method based on deep learning according to claim 1, characterized in that, After the Region of Interest (ROI) detection head extracts features, the top 100 diseased leaf bounding boxes and disease bounding boxes with the highest scores are obtained, and the bounding box with the highest score is obtained through non-maximum suppression to obtain the classification and position of the disease.
8. A hierarchical leaf disease detection system based on deep learning for implementing the detection method according to claim 1, characterized in that, it includes: A feature extraction unit for inputting the image to be detected, extracting disease features at different levels through the backbone network, performing feature fusion and feature extraction on the disease features at different levels through the Feature Pyramid Network to obtain a pyramid feature map; A diseased leaf generation unit for inputting a pyramid feature map into a pre - region proposal network to obtain disease proposal boxes and leaf proposal boxes for all categories; The pre - region proposal network consists of 5 feature extraction convolutional layers with a size of 3×3; A disease generation unit for generating pre - defined anchor boxes through a position anchor box generator in each leaf proposal box, and obtaining the regression parameters of the position and disease categories by regressing the center coordinates and sizes of the pre - defined anchor boxes through a post - region proposal network to generate disease proposal boxes; The post - region proposal network includes an ROIAlign module, an attention module, and a feature convolutional layer; The ROIAlign module obtains diseased leaf features from the leaf proposal boxes; The attention module adjusts the weights in the diseased leaf features to obtain adjusted diseased leaf features; The feature convolutional layer obtains disease proposal boxes based on the adjusted diseased leaf features; A classification and detection unit for obtaining the classification and positions of diseases and diseased leaves after extracting features from all disease proposal boxes through a region of interest detection head.
Citation Information
Patent Citations
Tunnel disease detection method and device and electronic equipment
CN114723709A