Detection method and device for plants with diseases and insect pests
Through the GSConv module and CA attention mechanism, combined with SIoU Loss, a pest detection model is constructed, which solves the accuracy and speed of pest and plant detection in the existing technology, and achieves efficient pest and plant identification.
Patent Information
- Application Number
- CN202510589943.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-08-08
AI Technical Summary
The existing pest and disease plant detection methods based on neural network architecture have inference delays and insufficient recognition accuracy in actual applications, resulting in missed detection and missed detection, and the inability to effectively distinguish pest and disease plants from normal plants, causing economic losses.
The GSConv module and CA attention mechanism are adopted, combined with SIoU Loss, a pest detection model is constructed, and the detection accuracy and speed are improved through feature extraction, feature transformation and border loss function training.
While lightening the model, it improves the accuracy and speed of disease and pest and plant detection, adapts to multiple detection architectures, and enhances the universality of the detection model.
Smart Images

Figure CN120451481A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural technology, and in particular to a method and device for detecting plant diseases and insect pests. Background Art
[0002] With the intensification of global agriculture and the worsening of climate change, pests and diseases are posing an increasingly severe threat to crop production. Globally, crop losses due to pests and diseases reach 20%-40% annually, severely threatening food security, farmers' livelihoods, and ecological balance. Furthermore, increased international trade is exacerbating the risk of the spread of invasive pests. Invasive species such as the fall armyworm and the red imported fire ant have already devastated agriculture and ecosystems in many regions. Therefore, the development of scientific and effective pest and disease prevention methods is urgent.
[0003] Currently, computer vision and deep learning-based methods are being introduced to assist in the identification of pests and diseases, enabling effective control. Compared to manual screening, these improved methods are more time-efficient and labor-efficient, and through large-scale application, they are improving the yield and quality of agricultural products.
[0004] However, existing neural network-based approaches suffer from significant inference delays and low recognition accuracy in practical applications, leading to the inability to distinguish between pest-infested and healthy plants, resulting in significant economic losses. Therefore, ensuring the accuracy of overall detection while minimizing false positives and missed detections of pest-infested plants in complex scenarios remains a pressing challenge in this field. Summary of the Invention
[0005] To address the low accuracy of plant pest and disease detection in existing technologies, the present invention proposes a method and device for detecting plant pests and diseases. This method integrates the GSConv module and the CA attention mechanism to achieve lightweight models while ensuring inference accuracy. Furthermore, the SIoU loss algorithm is combined to ensure both rapid convergence and high detection accuracy. Furthermore, the method and device are highly versatile and adaptable to a variety of convolutional neural network-based detection architectures, including but not limited to the YOLO series and Faster R-CNN.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A method for detecting plant diseases and insect pests, comprising the following steps:
[0008] S1: Obtain agricultural pest and disease plant dataset from the Internet;
[0009] S2: Input the agricultural pest and disease plant dataset into the constructed pest and disease detection model for training;
[0010] S3: Input the agricultural plant data set to be detected into the trained pest and disease detection model to output the pest and disease plants.
[0011] Preferably, in S1, the obtained agricultural plant disease and insect pest dataset is divided into a training set and a test set according to a ratio of 8:2.
[0012] Preferably, in S2, the pest and disease detection model includes a feature extraction unit, an FPN unit, a CA attention mechanism unit and a training unit;
[0013] Among them, the feature extraction unit is used to extract the initial feature map from the agricultural pest and disease plant data set;
[0014] The FPN unit is used to perform feature transformations of different scales on the initial feature map to obtain rich feature information;
[0015] The CA attention mechanism unit is used to encode the acquired feature information to obtain an enhanced feature map;
[0016] The training unit is used to construct the bounding box loss function and train the enhanced feature map.
[0017] Preferably, the S2 includes:
[0018] S2-1: Input the agricultural pest and disease plant dataset into the constructed feature extraction unit and output the initial feature map;
[0019] S2-2: Input the initial feature map into the FPN unit to perform feature transformation at different scales to obtain feature information;
[0020] S2-3: Input the acquired feature information into the CA attention mechanism unit for encoding to obtain the enhanced feature map;
[0021] S2-4: Construct a bounding box loss function as a training unit to train the enhanced feature map, thereby completing the training of the pest and disease detection model.
[0022] Preferably, in S2-1, the feature extraction unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fusion layer, and a transformation layer; the output end of the first convolutional layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer, respectively; the output end of the second convolutional layer and the output end of the third convolutional layer are connected to the input end of the fusion layer, respectively; and the output end of the fusion layer is connected to the input end of the transformation layer; wherein,
[0023] The first convolutional layer is a 1*1 convolution; the second and third convolutional layers are dual-branch structures. The second convolutional layer uses a 3*3 convolution, and the third convolutional layer is a 1*1 convolution, which is used for identity mapping, that is, the input and output channels remain unchanged; the fusion layer is used to perform a concat operation on the feature information output by the second and third convolutional layers to obtain fused features; finally, the transformation layer uses a shuffle function transformation to promote information exchange between channels, thereby outputting the initial feature map.
[0024] Preferably, in S2-2, the FPN unit outputs feature information of three different scales, which are 80*80, 40*40, and 20*20 respectively.
[0025] Preferably, in S2-3, for the input x, i.e., feature information, the CA attention mechanism unit uses pooling kernels of (H, 1) and (1, W) to encode the information in the horizontal and vertical directions, and the c-th channel with a height of h and a width of w can be expressed as:
[0026]
[0027] In formulas (1) and (2), and represents the feature representation obtained by global average pooling in the horizontal and vertical directions respectively; W and H represent the width and height of the input feature map respectively; x c (h,i) and x c (j,w) represents the feature value at a specific position in the input feature map;
[0028] Through the transformation of formulas (1) and (2), feature information in two directions is extracted and a feature map is obtained. The feature map is then converted into an attention map through encoding and finally multiplied by the original feature map to enhance the feature information extraction capability:
[0029] f=δ(F1([z h ,z w ])) (3)
[0030] g h =σ(F h (f h )) (4)
[0031] g w =σ(F w (f w )) (5)
[0032]
[0033] In formulas (3), (4), (5), and (6), f represents the feature map obtained after downsampling, and the two tensors in the horizontal and vertical directions are f h and f w ;δ represents the downsampling process; F1 represents 1×1 convolution; z h Represents the global average pooling in the horizontal direction; z w represents the global average pooling in the numerical direction; σ represents the activation function; F h Represents the encoding function in the horizontal direction; F w Represents the encoding function in the vertical direction; g h g represents the attention map obtained by transforming the tensor of the feature map in the horizontal direction; w Represents the attention map obtained by transforming the tensor of the feature map in the vertical direction; y c (i, j) represents the feature map of the final output; x c (i, j) represents the feature value of position (i, j) in the cth channel of the original input feature map; represents the attention weight of the c-th channel at horizontal position i; represents the attention weight of the c-th channel at horizontal position j.
[0034] Preferably, in S2-4, the border loss function includes Angle cost, Distance cost, Shape cost, and IoU cost, and the construction method is:
[0035] (1) In order to make the Angle cost function converge quickly, the angle perception component is introduced and defined as follows:
[0036]
[0037]
[0038] In formulas (7), (8), (9), and (10), Λ represents the penalty term for angle alignment, which is used to measure the angle deviation between the predicted box and the true box; X represents the intermediate variable in the calculation process; o is the center point distance between the predicted box and the true box; c h represents the height of the rectangle with o as the diagonal; α represents the angle formed by the diagonal and the width; and is the center coordinate of the real box; and Represents the center coordinate of the predicted box; β is the angle deviation to ensure diagonal alignment;
[0039] (2) Use the distance cost function to define the distance as follows:
[0040]
[0041] In formulas (11) and (12), Δ represents the distance loss, which is used to quantify the deviation between the predicted box and the true box in the center point coordinates; c w represents the width of the rectangle with o as the diagonal; t is the summation variable, which means the losses in the x and y directions are calculated and accumulated separately; γ is the adjustment coefficient of the distance loss; ρ t is the direction-dependent distance error term; ρ x represents the normalized distance error in the x direction; ρ y Represents the normalized distance error in the y direction;
[0042] (3)Shape cost function definition:
[0043]
[0044] In formulas (13) and (14), Ω is the shape loss, which is used to measure the difference between the predicted box and the true box in width w and height h; w t represents the difference in width or height; θ is the adjustment parameter of shape loss; θ w and θ h Represents the relative error of width and height respectively; w gt and h gt Indicates the width and height of the real box; w and h are the width and height of the predicted box;
[0045] (4) IoU cost function definition:
[0046] IoU reflects the ratio of the intersection to the union when the predicted box intersects the real box. The formula is as follows:
[0047]
[0048] In formula (15), b represents the center coordinate of the prediction box; ∩ represents the intersection; b gt Represents the center coordinates of the prediction box;
[0049] (5) Construct the SIoU cost function, that is, the regression loss of the border is:
[0050]
[0051] In formula (16), L box Represents the regression loss of the bounding box; IoU represents the IoU cost function.
[0052] Preferably, in S2-4, the test set is input into the trained pest and disease detection model for testing:
[0053]
[0054] In formulas (17), (18), (19), and (20), R precision Indicates the precision rate, which indicates the proportion of positive samples to the total number of samples; TP indicates the total number of correctly classified positive samples; FP indicates the total number of incorrectly classified positive samples; R recall represents the recall rate, which indicates the proportion of all detected positive samples to the positive samples in the data set; FN represents the total number of misclassified negative samples; AP represents the comprehensive evaluation index of single-category detection. The higher the AP value, the better the detection effect of a certain category; P(R) represents the functional relationship between the precision and the recall rate on the recall-precision curve; mAP represents the comprehensive evaluation of the entire network; C represents the total number of detection categories; AP(c) represents the comprehensive evaluation index of category c.
[0055] The present invention also provides a device for detecting plant diseases and insect pests, comprising a data acquisition module, a plant disease and insect pest detection module and a display module;
[0056] The data acquisition module is used to obtain the agricultural pest and disease plant data set and the agricultural plant data set to be tested;
[0057] The plant disease and insect pest detection module is used to build and train a plant disease and insect pest detection model, and perform plant disease and insect pest detection on the agricultural plant dataset to be tested;
[0058] The display module is used to display plants with diseases and insect pests.
[0059] In summary, due to the adoption of the above technical solution, compared with the prior art, the present invention has at least the following beneficial effects:
[0060] The present invention provides a method and device for detecting plant diseases and insect pests. First, a feature extraction unit is used to extract the final feature map to speed up the detection speed. Then, a CA attention mechanism unit is used to enhance the detection head's ability to capture information of different scales and improve detection accuracy. Finally, SIoU Loss is introduced as a bounding box loss to accelerate convergence and improve the detection accuracy of plant diseases and insect pests. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Schematic diagram of a method for detecting plant diseases and insect pests according to an exemplary embodiment of the present invention.
[0062] Figure 2 Schematic diagram of the structure of a pest and disease detection model according to an exemplary embodiment of the present invention.
[0063] Figure 3 FIG. 4 is a schematic diagram of a feature extraction unit according to an exemplary embodiment of the present invention.
[0064] Figure 4 Schematic diagram of the CA attention mechanism unit principle according to an exemplary embodiment of the present invention.
[0065] Figure 5 Schematic diagram of a device for detecting plant diseases and insect pests according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0066] The present invention will be further described in detail below with reference to the examples and specific implementation methods. However, this should not be understood as limiting the scope of the present invention to the following examples, as all technologies implemented based on the present invention fall within the scope of the present invention.
[0067] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0068] like Figure 1 As shown, the present invention provides a method for detecting plant diseases and insect pests, which specifically includes the following steps:
[0069] S1: Obtain the agricultural pest and disease plant dataset from the Internet (AI Challenger 2018, Agricultural Pest and Disease Research Library).
[0070] In this embodiment, the obtained agricultural pest and disease dataset is divided into a training set and a test set according to an 8:2 ratio.
[0071] S2: Input the agricultural pest and disease plant dataset into the constructed pest and disease detection model for training.
[0072] In this embodiment, Figure 2 As shown in the figure, the pest and disease detection model includes a feature extraction unit, an FPN unit, a CA attention mechanism unit, and a training unit. The output of the feature extraction unit is connected to the input of the FPN unit, the output of the FPN unit is connected to the input of the CA attention mechanism unit, and the output of the CA attention mechanism unit is connected to the input of the training unit.
[0073] Among them, the feature extraction unit is used to extract the initial feature map from the agricultural pest and disease plant data set;
[0074] The FPN unit is used to perform feature transformations of different scales on the initial feature map to obtain rich feature information;
[0075] The CA attention mechanism unit is used to encode the acquired feature information to obtain an enhanced feature map; the core of the CA attention mechanism unit is to enhance the output and thus obtain better detection results;
[0076] The training unit is used to construct the bounding box loss function and train the enhanced feature map.
[0077] S2-1: Input the agricultural pest and disease plant dataset into the constructed feature extraction unit and output the initial feature map.
[0078] In this embodiment, the feature extraction unit is based on the general YOLOv7 model, and the ELAN module in Backbone is replaced by the GSConv module (including the first convolutional layer, the second convolutional layer, and the third convolutional layer) for extracting the main features of the network and reducing the redundant parameters of the model. Figure 3 shown.
[0079] The feature extraction unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fusion layer and a transformation layer; the output end of the first convolutional layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer respectively, the output end of the second convolutional layer and the output end of the third convolutional layer are connected to the input end of the fusion layer respectively, and the output end of the fusion layer is connected to the input end of the transformation layer.
[0080] Among them, the first convolution layer is 1*1 convolution; the second and third convolution layers are dual-branch structures. The second convolution layer uses a 3*3 convolution with a larger kernel to realize feature extraction in order to reduce the number of model parameters; the third convolution layer is a depth-separable convolution, which is also a 1*1 convolution and is used for identity mapping, that is, the input and output channels remain unchanged; the fusion layer is used to perform concat operation on the feature information output by the second and third convolution layers to realize the fusion of feature information of different scales and obtain fused features; finally, the transformation layer uses shuffle function transformation to promote information exchange between channels, thereby outputting the initial feature map.
[0081] S2-2: Input the initial feature map into the FPN (Feature Pyramid Networks) unit for feature transformation at different scales to obtain rich feature information.
[0082] In this embodiment, the FPN unit performs feature transformations of different scales on the final feature map through multiple convolution modules and up and down sampling processing, thereby obtaining rich feature information and reducing feature redundancy.
[0083] The FPN unit outputs feature information of three different scales, with sizes of 80*80, 40*40, and 20*20 respectively.
[0084] S2-3: Introduce the CA coordinate attention mechanism unit to enhance the detection ability of targets of different scales, and encode the acquired feature information to obtain an enhanced feature map.
[0085] In this embodiment, in order to enhance the extraction of feature information, a CA attention mechanism unit is introduced into the detection head.
[0086] like Figure 4 As shown in the figure, for the input x (feature information), the pooling kernels of (H, 1) and (1, W) are used to encode the information in the horizontal and vertical directions. The c-th channel with a height of h and a width of w can be expressed as:
[0087]
[0088] In formulas (1) and (2), and Represents the feature representation obtained by global average pooling in the horizontal and vertical directions respectively. W and H represent the width and height of the input feature map respectively. c (h,i) and x c (j,w) represents the feature value at a specific position in the input feature map.
[0089] By transforming formulas (1) and (2), we can extract feature information in two directions (height and width) and obtain feature maps. Then, we convert the feature maps into attention maps through encoding, and finally multiply them by the original feature maps to enhance the feature information extraction capability. The formula is as follows:
[0090] f=δ(F1([z h ,z w ])) (3)
[0091] g h =σ(F h (f h )) (4)
[0092] g w =σ(F w (f w )) (5)
[0093]
[0094] In formulas (3), (4), (5), and (6), f represents the feature map obtained after downsampling, and the two tensors in the horizontal and vertical directions are f h and f w;δ represents the downsampling process; F1 represents 1×1 convolution; z h represents the feature encoding obtained by global average pooling in the horizontal direction (width direction), that is, the average value of all width features at each height h; z w represents the feature encoding obtained by global average pooling in the vertical direction (height direction), that is, the average value of all height features at each width w. σ represents the activation function (such as ReLU), which is used to perform nonlinear transformation on the encoded features and generate attention weights; F h It is the horizontal encoding function used to process the horizontal feature f h ; F w It is the encoding function in the vertical direction, which has similar functions and processes the vertical features f w ;g h g represents the attention map obtained by transforming the tensor of the feature map in the horizontal direction; w Represents the attention map obtained by transforming the tensor of the feature map in the vertical direction; y c (i,j) is the final output feature map by converting the original feature x c (i,j) and the attention weights in the horizontal and vertical directions and Multiplying together, we can enhance the important positions; c (i, j) is the eigenvalue of position (i, j) in the cth channel of the original input feature map; is the attention weight of the c-th channel at horizontal position i; is the attention weight of the c-th channel at horizontal position j.
[0095] S2-4: Construct a bounding box loss function as a training unit and train the final feature map to complete the training of the pest and disease detection model.
[0096] This paper introduces SIoU Loss as a bounding box loss. It is mainly composed of four cost functions: Anglecost, Distance cost, Shape cost, and IoU cost. Among them, the last three elements have been studied extensively in previous work and have produced positive effects, but there is still room for improvement. Therefore, the Angle cost is added. This addition ensures the prediction effect and enables the predicted box to be quickly moved to the nearest coordinate axis. Finally, only the X or Y coordinate needs to be regressed. In general, the Angle cost penalty greatly reduces the degrees of freedom of the loss, making it easier to converge.
[0097] (1) In order to make the angle cost function converge quickly, we will first try to minimize α, if α≤π / 4, otherwise minimize β=π / 4-α. To achieve this first, the angle perception component is introduced and defined as follows:
[0098]
[0099] In formulas (7), (8), (9), and (10), Λ represents the penalty term for angle alignment, which is used to measure the angle deviation between the predicted box and the true box; x is the intermediate variable in the calculation process; o is the center point distance between the predicted box and the true box; c h represents the height of the rectangle with o as the diagonal; α represents the angle formed by the diagonal and the width; and is the center coordinate of the real box; and represents the center coordinate of the predicted box; β is the angle deviation to ensure diagonal alignment.
[0100] (2) Use the distance cost function to define the distance as follows:
[0101]
[0102] In formulas (11) and (12), Δ represents the distance loss, which is used to quantify the deviation between the predicted box and the true box in the center point coordinates; c w represents the width of the rectangle with o as the diagonal; t is the summation variable, which means the losses in the x and y directions are calculated and accumulated separately; γ is the adjustment coefficient of the distance loss; ρ t is the direction-dependent distance error term; ρ x represents the normalized distance error in the x direction (horizontally); ρ y Represents the normalized distance error in the y direction (vertical);
[0103] From formulas (7)-(12), it can be seen that when α (in formula 8) is small, the contribution of distance cost is small, but as the angle gradually approaches , the contribution of Distance cost becomes larger and larger.
[0104] (3)Shape cost function definition:
[0105]
[0106] In formulas (13) and (14), Ω is the shape loss, which is used to measure the difference between the predicted box and the real box in width w and height h. Through exponential transformation and parameter θ weighting, the shape deviation of large-sized targets is highlighted; w tRepresents the difference in width or height; θ is the adjustment parameter of shape loss, which controls the sensitivity to the difference. The larger the value, the stronger the penalty for small deviations; w and θ h Represents the relative error of width and height respectively; w gt and h gt represents the width and height of the real box; w and h represent the width and height of the predicted box.
[0107] In this embodiment, θ reflects the degree of attention paid to the shape cost. The value of θ for each data set is uniquely determined. In this paper, θ=4 is calculated using a genetic algorithm.
[0108] (4) IoU cost function definition:
[0109] In this embodiment, IoU reflects the ratio of the intersection to the union when the predicted box intersects the real box. The formula is as follows:
[0110]
[0111] In formula (15), b represents the center coordinate of the prediction box; ∩ represents the intersection; b gt Represents the center coordinates of the prediction box;
[0112] (5) Construct the SIoU cost function, that is, the regression loss of the border is:
[0113]
[0114] In formula (16), L box Represents the regression loss of the bounding box; IoU represents the IoU cost function.
[0115] In this embodiment, the above-mentioned border loss is taken as part of the total loss, and the feature maps of each layer are trained and adjusted, not just the feature maps of S2-3. Because the neural network is a whole, the weights of the neural network will be adjusted according to the distance to the target, thereby obtaining better feature output and ultimately achieving better detection results.
[0116] In this example, the trained detection model is tested using a test set and the test results are evaluated. The test set is input into the trained pest detection model; the test set passes through the detection model, and finally outputs important results such as precision, recall, AP, and mAP. The evaluation parameters are calculated using the following formulas:
[0117]
[0118] In formulas (17), (18), (19), and (20), R precisionIndicates the precision rate, which indicates the proportion of positive samples to the total number of samples; TP indicates the total number of correctly classified positive samples; FP indicates the total number of incorrectly classified positive samples; R recall represents the recall rate, which indicates the proportion of all detected positive samples to the positive samples in the data set; FN represents the total number of misclassified negative samples; AP represents the comprehensive evaluation index of single-category detection. The higher the AP value, the better the detection effect of a certain category; P(R) represents the functional relationship between the precision and recall rate on the recall-precision curve; mAP represents the comprehensive evaluation of the entire network; C represents the total number of detection categories; AP(c) represents the comprehensive evaluation index of category c.
[0119] The complexity of a model is measured by the number of parameters or computational effort. Generally speaking, the fewer parameters a model has, the faster the detection speed. Speed is usually evaluated using FPS, which represents the number of frames per second.
[0120] In this embodiment, the above process is a target detection problem, and the quality of the model can be judged based on precision, recall, AP, and mAP.
[0121] S3: Input the agricultural plant data set to be detected into the trained pest and disease detection model to output the pest and disease plants.
[0122] In this embodiment, the agricultural plant dataset to be detected includes normal plants and plants with diseases and insect pests. Therefore, it is necessary to detect the plants with diseases and insect pests, thereby assisting in disease and insect pest prevention and control, and helping to improve the accuracy of decision-making.
[0123] Based on the above-mentioned method for detecting pests and diseases in plants, Figure 5 As shown, the present invention further provides a device for detecting plant diseases and insect pests, comprising a data acquisition module 1 , a plant disease and insect pest detection module 2 and a display module 3 .
[0124] The output end of the data acquisition module 1 is connected to the input end of the disease and insect pest plant detection module 2 , and the output end of the disease and insect pest plant detection module 2 is connected to the input end of the display module 3 .
[0125] The data acquisition module 1 is used to obtain a dataset of agricultural pests and diseases plants and a dataset of agricultural plants to be tested;
[0126] Plant pest and disease detection module 2 is used to build a plant pest and disease detection model and complete training, and perform plant pest and disease detection on the agricultural plant dataset to be detected;
[0127] The display module 3 is used to display the plants with diseases and insect pests.
[0128] The present invention also provides an electronic device, comprising a processor, wherein the processor is configured to execute a computer program stored in a memory, so that the electronic device implements the steps of a method for detecting plant diseases and insect pests in the above embodiment.
[0129] The present invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed on a processor, the computer program implements the steps of a method for detecting plant diseases and insect pests in the above embodiment.
[0130] A computer program includes computer program code, which may be in source code form, object code form, executable files, or some intermediate form. Computer-readable media may include at least any entity or device capable of carrying computer program code to an electronic device, recording media, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunications signals, and software distribution media. Examples include USB flash drives, removable hard drives, magnetic disks, or optical disks. In some jurisdictions, due to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunications signals.
[0131] Those skilled in the art will appreciate that the above-mentioned embodiments are specific examples for implementing the present invention, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A method for detecting plant diseases and insect pests, characterized in that: The specific steps include: S1: Obtain agricultural pest and disease plant dataset from the Internet; S2: Input the agricultural pest and disease plant dataset into the constructed pest and disease detection model for training; S3: Input the agricultural plant data set to be detected into the trained pest and disease detection model to output the pest and disease plants.
2. A method for detecting plant diseases and insect pests according to claim 1, characterized in that: In S1, the obtained agricultural pest and disease plant dataset is divided into a training set and a test set according to a ratio of 8:
2.
3. A method for detecting plant diseases and insect pests according to claim 1, characterized in that: In S2, the pest and disease detection model includes a feature extraction unit, an FPN unit, a CA attention mechanism unit and a training unit; Among them, the feature extraction unit is used to extract the initial feature map from the agricultural pest and disease plant data set; The FPN unit is used to perform feature transformations of different scales on the initial feature map to obtain rich feature information; The CA attention mechanism unit is used to encode the acquired feature information to obtain an enhanced feature map; The training unit is used to construct the bounding box loss function and train the enhanced feature map.
4. A method for detecting plant diseases and insect pests according to claim 3, characterized in that: The S2 includes: S2-1: Input the agricultural pest and disease plant dataset into the constructed feature extraction unit and output the initial feature map; S2-2: Input the initial feature map into the FPN unit to perform feature transformation at different scales to obtain feature information; S2-3: Input the acquired feature information into the CA attention mechanism unit for encoding to obtain the enhanced feature map; S2-4: Construct a bounding box loss function as a training unit to train the enhanced feature map, thereby completing the training of the pest and disease detection model.
5. A method for detecting plant diseases and insect pests according to claim 4, characterized in that: In S2-1, the feature extraction unit includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fusion layer, and a transformation layer; the output end of the first convolutional layer is connected to the input end of the second convolutional layer and the input end of the third convolutional layer, respectively; the output end of the second convolutional layer and the output end of the third convolutional layer are connected to the input end of the fusion layer, respectively; and the output end of the fusion layer is connected to the input end of the transformation layer; wherein, The first convolutional layer is a 1*1 convolution; the second and third convolutional layers are dual-branch structures. The second convolutional layer uses a 3*3 convolution, and the third convolutional layer is a 1*1 convolution, which is used for identity mapping, that is, the input and output channels remain unchanged; the fusion layer is used to perform a concat operation on the feature information output by the second and third convolutional layers to obtain fused features; finally, the transformation layer uses a shuffle function transformation to promote information exchange between channels, thereby outputting the initial feature map.
6. A method for detecting plant diseases and insect pests according to claim 4, characterized in that: In S2-2, the FPN unit outputs feature information of three different scales, which are 80*80, 40*40, and 20*20 respectively.
7. A method for detecting plant diseases and insect pests according to claim 4, characterized in that: In S2-3, for the input x, i.e., feature information, the CA attention mechanism unit uses the pooling kernels of (H, 1) and (1, W) to encode the information in the horizontal and vertical directions. The c-th channel with a height of h and a width of w can be expressed as: In formulas (1) and (2), and represents the feature representation obtained by global average pooling in the horizontal and vertical directions respectively; W and H represent the width and height of the input feature map respectively; x c (h,i) and x c (j,w) represents the feature value at a specific position in the input feature map; Through the transformation of formulas (1) and (2), feature information in two directions is extracted and a feature map is obtained. The feature map is then converted into an attention map through encoding and finally multiplied by the original feature map to enhance the feature information extraction capability: f=δ(F1([z h ,z w ])) (3) g h =σ(F h (f h )) (4) g w =σ(F w (f w )) (5) In formulas (3), (4), (5), and (6), f represents the feature map obtained after downsampling, and the two tensors in the horizontal and vertical directions are f h and f w ;δ represents the downsampling process; F1 represents 1×1 convolution; z h Represents the global average pooling in the horizontal direction; z w represents the global average pooling in the numerical direction; σ represents the activation function; F h Represents the encoding function in the horizontal direction; F w Represents the encoding function in the vertical direction; g h g represents the attention map obtained by transforming the tensor of the feature map in the horizontal direction; w Represents the attention map obtained by transforming the tensor of the feature map in the vertical direction; y c (i, j) represents the feature map of the final output; x c (i, j) represents the feature value of position (i, j) in the cth channel of the original input feature map; represents the attention weight of the c-th channel at horizontal position i; represents the attention weight of the c-th channel at horizontal position j.
8. A method for detecting plant diseases and insect pests according to claim 4, characterized in that: In S2-4, the border loss function includes Angle cost, Distance cost, Shape cost, and IoU cost, and the construction method is: (1) In order to make the Angle cost function converge quickly, the angle perception component is introduced and defined as follows: In formulas (7), (8), (9), and (10), Λ represents the penalty term for angle alignment, which is used to measure the angle deviation between the predicted box and the true box; X represents the intermediate variable in the calculation process; o is the center point distance between the predicted box and the true box; c h represents the height of the rectangle with o as the diagonal; α represents the angle formed by the diagonal and the width; and is the center coordinate of the real box; and Represents the center coordinate of the predicted box; β is the angle deviation to ensure diagonal alignment; (2) Use the distance cost function to define the distance as follows: In formulas (11) and (12), Δ represents the distance loss, which is used to quantify the deviation between the predicted box and the true box in the center point coordinates; c w represents the width of the rectangle with o as the diagonal; t is the summation variable, which means the losses in the x and y directions are calculated and accumulated separately; γ is the adjustment coefficient of the distance loss; ρ t is the direction-dependent distance error term; ρ x represents the normalized distance error in the x direction; ρ y Represents the normalized distance error in the y direction; (3)Shape cost function definition: In formulas (13) and (14), Ω is the shape loss, which is used to measure the difference between the predicted box and the true box in width w and height h; w t represents the difference in width or height; θ is the adjustment parameter of shape loss; θ w and θ h Represents the relative error of width and height respectively; w gt and h gt Indicates the width and height of the real box; w and h are the width and height of the predicted box; (4) IoU cost function definition: IoU reflects the ratio of the intersection to the union when the predicted box intersects the real box. The formula is as follows: In formula (15), b represents the center coordinate of the prediction box; ∩ represents the intersection; b gt Represents the center coordinates of the prediction box; (5) Construct the SIoU cost function, that is, the regression loss of the border is: In formula (16), L box Represents the regression loss of the bounding box; IoU represents the IoU cost function.
9. A method for detecting plant diseases and insect pests according to claim 4, characterized in that: In S2-4, the test set is input into the trained pest and disease detection model for testing: In formulas (17), (18), (19), and (20), R precision Indicates the precision rate, which indicates the proportion of positive samples to the total number of samples; TP indicates the total number of correctly classified positive samples; FP indicates the total number of incorrectly classified positive samples; R recall Recall rate, which indicates the proportion of all positive samples detected to the positive samples in the data set; FN represents the total number of misclassified negative samples; AP represents the comprehensive evaluation index of single-category detection. The higher the AP value, the better the detection effect of a certain category. P(R) represents the functional relationship between the precision and recall on the recall-precision curve. mAP represents the comprehensive evaluation of the entire network. C represents the total number of detection categories. AP(c) represents the comprehensive evaluation index of category c.
10. A device for detecting plant diseases and insect pests based on the method according to any one of claims 1 to 9, characterized in that: It includes data acquisition module, pest and disease plant detection module and display module; The data acquisition module is used to obtain the agricultural pest and disease plant data set and the agricultural plant data set to be tested; The plant disease and insect pest detection module is used to build and train a plant disease and insect pest detection model, and perform plant disease and insect pest detection on the agricultural plant dataset to be tested; The display module is used to display plants with diseases and insect pests.