Crop disease and insect pest recognition method and device and nonvolatile storage medium

By introducing attention modules into deep learning models, extracting and weighting crop image features, the problem of lack of real-time crop pest recognition methods in the prior art is solved, and a high accuracy and timely recognition effect is achieved.

CN119963921APending Publication Date: 2025-05-09CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125620.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-26
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

There is a lack of real-time crop pest identification methods for large-scale dispersed areas in the prior art, resulting in a lag in crop identification and low recognition accuracy.

Method used

A crop pest and disease recognition method is adopted to obtain the image to be identified and input it into a deep learning model trained with attention modules, feature extraction and recognition are performed, and the attention weighted feature map is used for prediction to obtain recognition results.

Benefits of technology

Real-time identification of crop diseases and pests in large-scale dispersed areas has been achieved, the accuracy and timeliness of identification have been improved, and the problems of lag and low accuracy have been solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963921A_ABST
    Figure CN119963921A_ABST
Patent Text Reader

Abstract

The invention discloses a crop disease and insect pest recognition method and device and a nonvolatile storage medium. The method comprises the following steps: acquiring a to-be-recognized image; the to-be-recognized image is input into a crop disease and pest recognition model to obtain a recognition result, the crop disease and pest recognition model is obtained by training an initial crop disease and pest recognition model through historical crop disease and pest images, and the initial crop disease and pest recognition model is a deep learning model provided with an attention module; the attention module is used for determining the weight of each channel in the feature map corresponding to the to-be-recognized image. The technical problems of hysteresis and low recognition accuracy of crop recognition due to lack of a real-time crop disease and insect pest recognition method for a large-scale scattered area in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and in particular, to a method, device and non-volatile storage medium for identifying crop diseases and insect pests. Background Art

[0002] In the related technologies, the identification, detection and forecasting of crop diseases and pests mainly rely on the plant protection department, and a modern early warning system has not yet been established. Due to the lack of prevention and control funds and professional staff, it is difficult to provide professional technical guidance in a timely manner in some remote areas, and accurate and scientific pest and disease prevention and control methods are difficult to be applied in a timely manner to various pest and disease occurrence areas, which easily leads to delays in disasters. In addition, farmers produce on a family basis and crops are planted in a relatively scattered manner. When pests and diseases break out, there is a phenomenon of blindly using pesticides, which not only aggravates the drug resistance of pests and diseases in the region, but also seriously harms the ecological environment. In the related technologies, there is a lack of real-time crop disease and pest identification methods for large-scale dispersed areas, resulting in a lag in crop identification and a low accuracy rate of identification.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a method, device and non-volatile storage medium for identifying crop diseases and pests, so as to at least solve the technical problem in the related art that there is a lack of a real-time identification method for crop diseases and pests in large-scale dispersed areas, resulting in a lag in crop identification and a low identification accuracy.

[0005] According to one aspect of an embodiment of the present application, a method for identifying crop pests and diseases is provided, comprising: obtaining an image to be identified; inputting the image to be identified into a crop pest and disease identification model to obtain an identification result, wherein the crop pest and disease identification model is obtained by training an initial crop pest and disease identification model with historical crop pest and disease images, and the initial crop pest and disease identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weights of each channel in a feature map corresponding to the image to be identified.

[0006] In some embodiments of the present application, an image to be identified is input into a crop disease and pest identification model to obtain a recognition result, including: inputting the image to be identified into the input layer of the crop disease and pest identification model; performing feature extraction on the image to be identified through the convolution layer of the crop disease and pest identification model to obtain a feature map; processing the feature map based on an attention module to determine an attention weight, wherein the attention weight includes a horizontal attention weight and a vertical attention weight; identifying the image to be identified based on the fully connected layer and the attention weight of the crop disease and pest identification model to obtain a recognition result.

[0007] In some embodiments of the present application, a feature map is processed based on an attention module to obtain an attention weight, including: performing global average pooling on the feature map in the horizontal and vertical directions of the coordinate axis, respectively, to obtain a first horizontal feature vector corresponding to the horizontal direction and a first vertical feature vector corresponding to the vertical direction, wherein the horizontal direction corresponds to the width of the feature map, and the vertical direction corresponds to the height of the feature map; fusing the first horizontal feature vector and the first vertical feature vector to obtain a fused feature vector; processing the fused feature vector to obtain an attention weight.

[0008] In some embodiments of the present application, a horizontal feature vector and a vertical feature vector are fused to obtain a fused feature vector, including: splicing a first horizontal feature vector and a first vertical feature vector to obtain a spliced ​​vector; performing convolution dimensionality reduction on the spliced ​​vector based on a convolution kernel, and activating the spliced ​​vector after convolution dimensionality reduction based on a first preset activation function to obtain a fused feature vector.

[0009] In some embodiments of the present application, the fused feature vector is processed to obtain the attention weight, including: dividing the fused feature vector along the spatial dimension to obtain a second horizontal feature vector and a second vertical feature vector; activating the second horizontal feature vector and the second vertical feature vector respectively based on a second preset activation function to obtain the horizontal attention weight and the vertical attention weight.

[0010] In some embodiments of the present application, the image to be identified is identified based on the fully connected layer and attention weights of the crop disease and pest identification model to obtain a recognition result, including: using the attention weights to weight each channel in the feature map to obtain a weighted feature map; predicting the weighted feature map based on the fully connected layer to obtain a recognition result.

[0011] In some embodiments of the present application, a weighted feature map is predicted based on a fully connected layer to obtain a recognition result, including: using a fully connected layer to perform classification prediction on the weighted feature map to obtain an output vector, wherein each element of the output vector corresponds to a predicted value of a preset type of crop disease and pest; processing the output vector through a third preset activation function to obtain a probability value corresponding to each element; and determining the preset type of crop disease and pest corresponding to the element with the highest probability value as the recognition result.

[0012] In some embodiments of the present application, a crop disease and pest recognition model is obtained by training an initial crop disease and pest recognition model with historical crop disease and pest images in the following manner: obtaining historical crop disease and pest images, wherein the historical crop disease and pest images are divided into different categories based on the types of crop disease and pests; determining a training set based on the historical crop disease and pest images, wherein the training set is used to train the initial crop disease and pest recognition model; training the initial crop disease and pest recognition model based on the training set until a preset number of iterations is reached to obtain a crop disease and pest recognition model.

[0013] In some embodiments of the present application, the method also includes: sending identification information associated with the recognition result to a user terminal, wherein the identification information includes at least an image to be identified, details of the recognition result, characteristics of the pests and diseases, and control measures, and the details of the recognition result include at least an image of the pest and disease part and the corresponding type of pest and disease.

[0014] According to another aspect of an embodiment of the present application, a device for identifying crop diseases and pests is also provided, including: an acquisition module for acquiring an image to be identified; and a recognition module for inputting the image to be identified into a crop disease and pest recognition model to obtain a recognition result, wherein the crop disease and pest recognition model is obtained by training an initial crop disease and pest recognition model with historical crop disease and pest images, and the initial crop disease and pest recognition model is a deep learning model provided with an attention module, wherein the attention module is used to determine the weights of each channel in a feature map corresponding to the image to be identified.

[0015] According to another aspect of an embodiment of the present application, a non-volatile storage medium is further provided, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above method for identifying crop diseases and pests.

[0016] According to another aspect of an embodiment of the present application, an electronic device is further provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the above method for identifying crop pests and diseases is executed when the program is run.

[0017] According to another aspect of the embodiments of the present application, a computer program product is also provided, including computer instructions, which implement the above method for identifying crop diseases and insect pests when executed by a processor.

[0018] In an embodiment of the present application, an image to be identified is obtained; the image to be identified is input into a crop disease and pest identification model to obtain a recognition result, wherein the crop disease and pest identification model is obtained by training an initial crop disease and pest identification model with historical crop disease and pest images, and the initial crop disease and pest identification model is a deep learning model provided with an attention module, wherein the attention module is used to determine the weights of each channel in a feature map corresponding to the image to be identified, and image recognition is performed on the image to be identified by the crop disease and pest identification model to directly obtain a recognition result, thereby realizing real-time crop diseases and pests, thereby solving the technical problems in the related art of lacking a real-time crop disease and pest identification method for large-scale scattered areas, resulting in lag in crop identification and low recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 It is a hardware structure block diagram of a computer terminal for implementing a method for identifying crop pests and diseases according to an embodiment of the present application;

[0021] Figure 2 is a flow chart of a method for identifying crop pests and diseases provided in an embodiment of the present application;

[0022] Figure 3 A depth-separable convolutional graph is provided according to an embodiment of the present application;

[0023] Figure 4 is a schematic diagram of a point-by-point convolution operation provided according to an embodiment of the present application;

[0024] Figure 5 is an inverse residual structure diagram provided according to an embodiment of the present application;

[0025] Figure 6 It is a Bottleneck structure diagram of MobileNetV3 provided according to an embodiment of the present application;

[0026] Figure 7 It is a CA attention mechanism structure diagram provided according to an embodiment of the present application;

[0027] Figure 8 It is a Bottleneck structure diagram in a CA-MobileNetV3 model provided according to an embodiment of the present application;

[0028] Fig. 9It is a structural schematic diagram of a crop disease and insect pest identification device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0030] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or reject automated decision-making results; if the user chooses to reject, the expert decision-making process will be entered.

[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0033] Attention mechanism: The attention mechanism is a mechanism that enables a neural network to focus on a subset of its inputs (or features). Attention can be applied to any type of input regardless of its shape. In the case of limited computing power, the attention mechanism is a resource allocation solution that is the main means to solve the problem of information overload, allocating computing resources to more important tasks.

[0034] MobileNet: A lightweight convolutional neural network model.

[0035] In the related art, the identification, detection and forecasting of crop diseases and insect pests mainly rely on the plant protection department, and a modern early warning system has not yet been established. Therefore, there is a lack of real-time identification methods for crop diseases and insect pests in large-scale dispersed areas in the related art, resulting in lag in crop identification and low identification accuracy. In order to solve this problem, the present application provides a relevant solution in the embodiment, which is described in detail below.

[0036] According to an embodiment of the present application, an embodiment of a method for identifying crop diseases and pests is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] The method embodiments provided in the embodiments of the present application can be executed in a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal for implementing a method for identifying crop pests and diseases. Figure 1 As shown, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.

[0038] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0039] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for identifying crop pests and diseases in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the above-mentioned method for identifying crop pests and diseases is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0040] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0041] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0042] In the above operating environment, an embodiment of the present application provides an embodiment of a method for identifying crop diseases and pests. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0043] like Figure 2 FIG. 1 is a flow chart of a method for identifying crop pests and diseases according to an embodiment of the present application, comprising:

[0044] Step S202, obtaining an image to be recognized.

[0045] The following is a specific embodiment: the image data to be identified is an image taken by the user who wants to identify crop pests and diseases in response to the upload instruction of the user terminal. The crop leaf pest identification system (for example, the common pest identification system for apple trees, the system is developed based on the Django framework (an open source framework written in Python), mainly including image preprocessing and pest identification functions) includes an image input module, an image recognition module and a result display module. The image input module receives the above-mentioned image to be identified, sends it to the crop pest and disease identification model of the image recognition module for identification, obtains the identification result, and displays the identification result through the result display module. In specific applications, the user only needs to upload the crop leaf pest and disease image to be queried (i.e., the above-mentioned image to be identified) in the identification interface of the image input module, click the query button to send the image to the server, the server receives the image uploaded by the user as the input of the crop pest and disease identification model, and displays the identification result to the user in the display interface of the result display module (i.e., the user terminal, the user terminal includes but is not limited to the web page, mobile phone APP).

[0046] Step S204, input the image to be identified into a crop disease and pest identification model to obtain an identification result, wherein the crop disease and pest identification model is obtained by training an initial crop disease and pest identification model with historical crop disease and pest images, and the initial crop disease and pest identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be identified.

[0047] After obtaining the recognition result, the identification information associated with the recognition result is sent to the user terminal, wherein the recognition information at least includes the image to be recognized, the recognition result details, the characteristics of the pests and diseases and the prevention and control measures. The recognition result details at least include the image of the pest and disease part and the corresponding pest and disease type.

[0048] In step S204, the image to be identified is input into the crop pest and disease identification model, and there are many ways to implement the identification result, for example: input the image to be identified into the input layer of the crop pest and disease identification model; extract features of the image to be identified through the convolution layer of the crop pest and disease identification model to obtain a feature map; process the feature map based on the attention module to determine the attention weight, wherein the attention weight includes the horizontal attention weight and the vertical attention weight; identify the image to be identified based on the fully connected layer and the attention weight of the crop pest and disease identification model to obtain a recognition result.

[0049] There are many ways to process the feature map based on the attention module to obtain the attention weight, for example: perform global average pooling on the feature map in the horizontal direction and the vertical direction of the coordinate axis respectively to obtain the first horizontal feature vector corresponding to the horizontal direction and the first vertical feature vector corresponding to the vertical direction, wherein the horizontal direction corresponds to the width of the feature map and the vertical direction corresponds to the height of the feature map; fuse the first horizontal feature vector and the first vertical feature vector to obtain a fused feature vector; process the fused feature vector to obtain the attention weight. Fusion of the horizontal feature vector and the vertical feature vector to obtain the fused feature vector can be achieved in the following ways: splicing the first horizontal feature vector and the first vertical feature vector to obtain a spliced ​​vector; performing convolution dimensionality reduction on the spliced ​​vector based on the convolution kernel, and activating the spliced ​​vector after convolution dimensionality reduction based on the first preset activation function to obtain the fused feature vector.

[0050] In the above steps, there are multiple ways to process the fused feature vector to obtain the attention weight, for example: split the fused feature vector along the spatial dimension to obtain the second horizontal feature vector and the second vertical feature vector; activate the second horizontal feature vector and the second vertical feature vector respectively based on the second preset activation function to obtain the horizontal attention weight and the vertical attention weight. There are multiple ways to obtain the recognition result based on the fully connected layer and attention weight of the crop pest recognition model and the image to be recognized, for example: use the attention weight to weight each channel in the feature map to obtain a weighted feature map; predict the weighted feature map based on the fully connected layer to obtain the recognition result: use the fully connected layer to classify and predict the weighted feature map to obtain an output vector, wherein each element of the output vector corresponds to a predicted value of a preset crop pest type; process the output vector through the third preset activation function to obtain the probability value corresponding to each element; determine the preset crop pest type corresponding to the element with the highest probability value as the recognition result.

[0051] The following are specific embodiments:

[0052] The crop pest and disease recognition model can be a MobileNet model improved by a hybrid attention mechanism (named CA-MobileNet model). The CA-MobileNet model is a third-generation MobileNet model (MobileNetV3 model) that introduces an attention module (called CA attention mechanism module). The CA-MobileNet model is an improved MobileNetV3 model structure that uses an attention mechanism module (Squeeze-and-Excitation Networks, referred to as SENet). The SENet module (or SE module) does not consider the position information of the image, so the CA attention mechanism module is introduced into MobileNetV3 to obtain the CA-MobileNet model. Compared with the original MobileNetV3 model, the CA-MobileNet model has improved convergence speed and recognition accuracy, and uses the rotation and Gaussian noise addition method to expand the pest and disease data set. The expanded crop pest and disease leaf data set is used for training, and the final model recognition accuracy reaches 99%. The MobileNetV3 neural network is a new generation of lightweight neural network model improved on the basis of the first generation MobileNet model (MobileNetV1) and the second generation MobileNet model (MobileNetV2). MobileNetV3 combines the depthwise separable convolution of MobileNetV1, the inverted residual structure and linear bottleneck of MobileNetV2, adds an attention mechanism module based on MobileNetV1 and MobileNetV2, redesigns the structure of the time-consuming layer and updates the activation function from swish to h-swish. The expression of the activation function swish is y=x·σ(x), where σ(x) is the S-type function (sigmoid function), x represents the input value, and y represents the output value. The activation function h-swish is a variant of the swish function, and its expression is Among them, the expression of the ReLU6 function (a rectified linear unit activation function limited between 0 and 6) is min(max(0, x+3), 6), which means taking the maximum value between 0 and x+3, and then taking the minimum value with 6, x represents the input value, and y represents the output value. Compared with the swish function, the calculation of h-swish is more efficient because it avoids the calculation of the sigmoid function, which involves exponential operations, is relatively complex to implement on hardware and has high computational cost. While h-swish mainly involves simple linear operations and ReLU6 function operations, it can calculate the results faster in resource-constrained environments such as mobile devices. Due to its computational efficiency advantage, h-swish is widely used in deep learning models on mobile terminals or with demanding computing resources. For example, in some lightweight neural network architectures, it is used to accelerate the inference process of the model while maintaining good model accuracy, and has good performance in lightweight models for tasks such as target detection and image segmentation. The MobileNetV3 model structure is shown in Table 1 below.

[0053] Table 1

[0054]

[0055]

[0056] The first column of the table is the size of the input vector of each feature layer of MobileNetV3; the second column is the operation performed by each feature layer. Conv2D,33 is a 2D convolution layer (Conv2D). "33" usually indicates that the size of the convolution kernel is 3x3, which means that in this layer, the input feature map will be convolved with a 3x3 convolution kernel to generate a new feature map. Conv2D,11 is similar to Conv2D,33, and "11" indicates that the size of the convolution kernel is 1x1; Bneck,33 is the bottleneck layer (Bottleneck) structure layer in MobileNetV3, where "33" also indicates a 3x3 depth-separable convolution; Bneck,55 is similar to "Bneck,33", but the "55" here indicates that a 5x5 convolution kernel is used; Pool,77 refers to a pooling layer (Pooling Layer) in the neural network model, which specifically refers to a 7x7 global average pooling (GlobalAverage Pooling) operation. The third and fourth columns represent the number of channels after the inner inverse residual structure is upgraded and the number of channels of the Bottleneck output in the bottleneck layer (Bottleneck) structure; SE represents whether the attention mechanism module is introduced in this layer; NL represents the type of activation function, where HS represents h-swish and RE represents Rectified Linear Unit (ReLU); the last column Strides represents the step size of the convolution in the Bottleneck structure.

[0057] Before inputting the image to be identified into the input layer of the crop pest and disease recognition model, it is necessary to adjust the size of the image to be identified to the size required by the crop pest and disease recognition model. For example, use bilinear interpolation to adjust the image size to the input size required by the CA-MobileNet model. The image to be identified is input into the input layer of the crop pest and disease recognition model. The input layer will be initialized according to the size and number of channels of the image. Then, the convolution layer of the crop pest and disease recognition model is used to extract features of the image to be identified. The convolution layer uses multiple convolution kernels to slide the image data and calculate features to generate a multi-channel feature map. The feature map will pass through multiple Bottleneck structures. Each Bottleneck structure may contain depthwise separable convolution, point-by-point convolution, SE module, etc. to extract deeper feature information. Figure 3As shown, it is a depth-separable convolution graph provided according to an embodiment of the present application, which shows the process of using multiple convolution kernels in the convolution layer to slide the image data and calculate features to generate a multi-channel feature map. The convolution layer performs a depthwise (DW) operation. DW is a special convolution operation, which is mainly used in convolutional neural networks in deep learning. When processing multi-channel input data (such as the red, green, and blue (RGB) channels of an image), DW convolution is to perform convolution operations on each channel separately. RGB is a color model in image processing. Specifically: for an input feature map (i.e., the above-mentioned image to be identified) H×W×C, (H is the height, W is the width, and C is the number of channels), the convolution kernel (Filters) size of DW convolution is k×k×1 (k is a constant). DW convolution will have C such convolution kernels, and each convolution kernel only convolves one channel of the input feature map, so that C convolution feature maps are obtained, and the height × width of each feature map is H. ′ ×W ′ This process can be seen as performing spatial convolution on each channel independently without mixing the information between channels. The 3channelinput in the figure represents the feature map of 3 channels (H×W×3) (i.e. the image to be recognized above). Each channel is convolved through a convolution kernel (Filters), and finally 3 feature maps are obtained.

[0058] Each of the above convolution kernels only performs convolution operation with one channel of the input feature map, and the number of channels of each convolution kernel is 1, so the number of DW convolution kernels is equal to the number of channels of the input feature map, and the number of channels of the output feature matrix (also called output feature map) (each channel corresponds to a feature map) is equal to the number of channels of the input feature map and the number of convolution kernels. If you want to change the number of feature maps, you need to add a pointwise convolution (PW) convolution operation after the DW convolution operation. Figure 4As shown, it is a schematic diagram of a point-by-point convolution operation provided by the present application. PW convolution is an ordinary convolution with a convolution kernel size of 1. Combining DW convolution and PW convolution into one is called depthwise separable convolution. Theoretically, the computational complexity of depthwise separable convolution is only 1 / 9 of that of ordinary convolution operation. Therefore, the use of depthwise separable convolution can greatly reduce the computational complexity of the model. The main feature of PW convolution is that the size of the convolution kernel is 1×1. In depthwise separable convolution, DW convolution is performed first. It performs spatial convolution on each channel independently and does not mix the information between channels. Then PW convolution is performed to linearly combine the features of each channel after DW convolution and fuse the information of different channels, so that the model can learn more complex feature representations. This combination method is widely used in some lightweight neural network architectures, which can reduce the number of parameters and the amount of computation while maintaining good model performance. Figure 4 In , the input feature map (Maps*3) size ((height, width, number of channels) is H×W×3), the PW convolution kernel (Filters*4) used is 1×1×3, and the number of output channels is 4. For any position of the input feature map (for example, the upper left corner), the operation of the PW convolution at this position is: each depth component of the convolution kernel is element-wise multiplied with the pixel value of the same position on the corresponding channel of the input feature map (that is, the 1×1×3 area (that is, 3 elements) corresponding to the position), and then the 3 results are added to obtain a single output value. This output value belongs to the first channel of the output feature map. Then, the second output channel is calculated in the same way, so that two values ​​of the output feature map at the upper left corner are obtained. Repeating this process for all positions of the entire input feature map, you can get an output feature map (Maps*4) of (height×width×number of channels) H×W×4, including 4 feature maps, each with a height×width of H×W.

[0059] After obtaining the feature map, the feature map is processed based on the attention module (i.e., the CA attention mechanism module mentioned above) to determine the attention weight: the feature map is globally averaged pooled in the horizontal and vertical directions of the coordinate axis to obtain the first horizontal feature vector corresponding to the horizontal direction and the first vertical feature vector corresponding to the vertical direction. Specifically, the attention module (e.g., the CA attention mechanism) performs global average pooling on the feature map along the horizontal direction (i.e., the X-axis direction) and the vertical direction (i.e., the Y-axis direction) of the coordinate axis, and divides the input features into two parts (i.e., the first horizontal feature vector corresponding to the horizontal direction and the first vertical feature vector corresponding to the vertical direction). Different characteristics of the input features can be obtained in one spatial direction, and accurate position information is retained in the other spatial direction. When performing global average pooling, the global average pooling is divided into two steps. For example, the feature map of size (height × width × number of channels) H×W×C is pooled separately, and two pooling kernels (H, 1) and (1, W) are used to perform pooling operations along the horizontal coordinate and vertical coordinate directions of the feature map to obtain the first horizontal feature vector corresponding to the horizontal direction. The first vertical eigenvector corresponding to the vertical direction Among them, h and H represent height, which helps the network locate the target area more accurately.

[0060] The first horizontal feature vector and the first vertical feature vector are concatenated to obtain a concatenated vector. Specifically, the first horizontal feature vector corresponding to the horizontal direction is concatenated to obtain a concatenated vector. The first vertical eigenvector corresponding to the vertical direction The splicing is performed along the spatial dimension to obtain a splicing vector. The spatial dimension may refer to the width (W) and height (H) dimensions of the image. The splicing vector is convolved and reduced in dimension based on the convolution kernel, and the splicing vector after convolution and reduction is activated based on the first preset activation function to obtain a fused feature vector, that is, after the 1×1 convolution operation, the activation function is used for activation. Specifically: The concatenated vector is convolved and activated:

[0061]

[0062] Among them, F1 represents convolution dimensionality reduction using a 1×1 convolution kernel, δ represents the first preset activation function, and f represents the concatenated vector after convolution dimensionality reduction. (That is, the size of f is C / r×1×(H+W), where r is the number of channels of the convolution kernel).

[0063] Then, the fused feature vector is segmented along the spatial dimension to obtain a second horizontal feature vector and a second vertical feature vector; based on a second preset activation function, the second horizontal feature vector and the second vertical feature vector are activated respectively to obtain a horizontal attention weight and a vertical attention weight, specifically:

[0064] Split f along the spatial dimension and then convolve it and use Sigmoid activation:

[0065] g h =σ(F h (f h ))

[0066] g w =σ(F w (f w ))

[0067] Among them, f h represents the second level eigenvector, f w represents the second vertical eigenvector, F h and F w They represent 1×1 convolution for dimension increase, σ represents the Sigmoid activation function (i.e., the second preset activation function mentioned above), and f w ∈R C×1×w ,f h ∈R C ×H×1 , g h represents the attention weight in the vertical direction, g w Represents the attention weight in the horizontal direction.

[0068] Use the attention weights to weight each channel in the feature map to obtain a weighted feature map:

[0069]

[0070] Among them, y c (i, j) represents the vector value (i.e., eigenvalue) with coordinates (i, j) in the Cth channel of the weighted feature map, x c (i,j) represents the coordinate of the Cth channel of the input vector (i.e., feature map) as the (i,j) vector value (i.e., eigenvalue). Represents the vector value at coordinate (i, 1) in the feature vector of size H×1×C (i.e., the attention weight in the vertical direction at coordinate (i, 1)). Represents the vector value at coordinate (1, j) in the feature vector of size 1×W×C (i.e., the attention weight in the horizontal direction at coordinate (1, j)).

[0071] Based on the fully connected layer, the weighted feature map is predicted to obtain the recognition result: the weighted feature map is classified and predicted using the fully connected layer to obtain an output vector, wherein each element of the output vector corresponds to a predicted value of a preset crop pest type; the output vector is processed by the third preset activation function to obtain the probability value corresponding to each element; the preset crop pest type corresponding to the element with the highest probability value is determined as the recognition result. Specifically: first, a 256-dimensional fully connected layer is added after the last layer of the initial crop pest recognition model. The main purpose of this is to reduce the dimension of the model output. Then a layer of ReLU activation layer is added, by setting all negative values ​​to zero and keeping the positive values ​​unchanged, thereby increasing the nonlinearity of the model. In order to enhance the generalization ability of the model, another layer of dropout regularization (Dropout Regularization, referred to as Dropout) layer is added. The Dropout layer makes a part of the neurons randomly dormant to reduce the model parameters. Dropout is a regularization method. By randomly dormant a certain proportion of neurons during the training process, the model's dependence on training data can be reduced, thereby improving the performance of the model on unseen data. In this embodiment, the Dropout ratio is set to 0.3, which means that during training, 30% of the neurons will be temporarily "turned off" during each forward propagation process, and will not participate in the calculation or update the weights. This is beneficial to increase the generalization ability of the model.

[0072] Finally, the output dimension of the fully connected layer is set to match the number of different types of crop pests and diseases to be identified. For example, four different types of crop leaf pests and diseases need to be identified, and the output dimension of the fully connected layer is 4. The purpose of this layer is to map the feature vector processed as above to the predicted values ​​of different types of pests and diseases. The weighted feature map is classified and predicted using the fully connected layer to obtain an output vector. In order to display the classification prediction results in the form of probability, a soft maximum function (Softmax Function, referred to as Softmax) activation function is added to the output layer as a Softmax layer. The Softmax function converts the output vector of the fully connected layer into a probability vector, in which each element represents the probability that the corresponding pest and disease type belongs to this category. The Softmax function (i.e., the third preset activation function mentioned above) can convert each element of the output vector into a value between 0 and 1, and the sum of these values ​​is 1. Finally, the preset crop pest and disease type corresponding to the element with the highest probability value is determined as the recognition result.

[0073] like Figure 5As shown, it is an inverse residual structure diagram provided according to an embodiment of the present application. The inverted residual block is a module structure used in a convolutional neural network (CNN) of deep learning, showing the steps of "dimensionality increase-convolution-dimensionality reduction". The convolution of the image to be identified first increases the channel dimension, and then the DW convolution is used for feature extraction, and finally the convolution is used to reduce the number of channels. The reason for using the inverse residual structure is that the high-dimensional information loses less information after passing through the ReLU activation function. The inverse residual structure is used to increase the dimension and then use the ReLU6 activation function, which can retain the image features as much as possible. The last dimensionality reduction convolution layer of the inverse residual structure uses a linear activation function. By first increasing the dimension and then performing spatial convolution, features can be better extracted and fused. The dimensionality increase operation allows the network to have more "space" to learn features, especially when processing complex image data and other tasks, it can better capture the relationship and spatial information between different channels. Specifically, as shown in the figure, the input feature map is first dimensionalized by a point-to-point convolution layer, that is, the number of channels is increased to obtain a 3x3 feature map. The input feature map is first passed through a 1x1 convolution layer for dimensionality increase, that is, the number of channels is increased. Relu6, depthwise convolution (Dwise) means that the feature map after dimensionality increase will be passed to a depthwise separable convolution layer, each input channel is depthwise convolved with a corresponding 3x3 convolution kernel, without mixing channel information, and then the feature map after depthwise convolution is fused in the channel dimension by pointwise convolution, using ReLU6 or h-swish as the activation function. The feature map after depthwise separable convolution is then passed through a 1x1 convolution layer for dimensionality reduction, returning to the same number of channels as the input feature map.

[0074] like Figure 6 As shown in the figure, the Bottleneck structure diagram of MobileNetV3 shows the Bottleneck structure diagram of the MobileNetV3 model. As shown in the figure, the steps are as follows: First, the input feature map is adjusted through a 1x1 convolution layer to adjust the number of channels, and NL (NL represents the nonlinear activation function (h-swish)) is used as the activation function, and then a depthwise separable convolution (Dwise) is applied for downsampling, and the NL activation function is used for nonlinear transformation. Then, the SE module is introduced as needed, which is used to enhance the interactivity between channels. The SE module includes global average pooling (Pool), a fully connected layer (FC), and an activation function (Relu or hard-α (hard-α is the part of the hard-swish (i.e., h-swish) activation function that linearly transforms the input value). Finally, a 1x1 convolution layer is used to adjust the number of channels of the feature map.

[0075] like Figure 7 As shown, it is a CA attention mechanism structure diagram provided according to an embodiment of the present application, which shows the structure of the CA attention mechanism used by the CA attention mechanism module mentioned in the embodiment of the present application, and the input feature map (i.e., the feature map, represented by the cube in the figure) is pooled in the horizontal direction (X Avg Pool) and the vertical direction (Y Avg Pool), respectively, to generate feature maps in two directions (C×1×W and C×H×1, i.e., the first horizontal feature vector corresponding to the horizontal direction and the first vertical feature vector corresponding to the vertical direction), and the spatial position information is retained. Concat+Conv1×1+Sigmoid means that after the two feature maps are connected (Concat) (i.e., the first horizontal feature vector and the first vertical feature vector are concatenated to obtain the concatenated vector), the features are extracted through the convolution (Conv) operation and nonlinearly activated through the Sigmoid function (same as the sigmoid function), after the split (split) operation (i.e., the fused feature vector is segmented along the spatial dimension to obtain the second horizontal feature vector and the second vertical feature vector), the attention weights in the horizontal and vertical directions are further generated. 1×1 convolution (Conv1×1) is used to perform dimensionality increase operation and convolve with the input feature map after sigmoid activation to achieve direction-aware and position-sensitive attention enhancement (that is, the second horizontal feature vector and the second vertical feature vector are activated respectively based on the second preset activation function to obtain the horizontal attention weight and the vertical attention weight).

[0076] like Figure 8 As shown, it is a Bottleneck structure diagram in a CA-MobileNetV3 model provided according to an embodiment of the present application; Figure 6The improvement of the process shows the structure of Bottleneck in the CA-MobileNetV3 model of this application. First, the feature map is subjected to 1×1 PW convolution (i.e. Conv1×1 in the figure) and 3×3 Depthwise (DW) convolution (i.e. Dwise3×3 in the figure) to form a Depthwise-Separable convolution operation, which fuses the information of different channels to obtain a feature map, so that the model can learn a more complex feature representation. 2. The feature map is pooled in the horizontal direction (X Avg Pool) and the vertical direction (YAvg Pool) to generate feature maps in two directions (C×1×W and C×H×1, i.e. the first horizontal feature vector corresponding to the horizontal direction and the first vertical feature vector corresponding to the vertical direction), and the spatial position information is retained. Concat+Conv2d means that after the two feature maps are connected (concat), the features are extracted through the convolution (conv2d) operation. BN+Activate means that the input of each layer of the neural network is normalized through the batch normalization layer (BN), so that the distribution of the input data is more stable. Through nonlinear activation (activate), after the split operation, the attention weights in the horizontal and vertical directions (i.e. the horizontal attention weight and the vertical attention weight mentioned above) are further generated. Specifically: the convolution (conv2d) operation extracts features, and the second horizontal feature vector and the second vertical feature vector are obtained after the split operation, and the Sigmoid (same as sigmoid) activation function in the nonlinear activation (activate) is used for activation to obtain the horizontal attention weight and the vertical attention weight. Finally, the 1×1 convolution is used for dimensionality increase operation and convolved with the input feature map to achieve direction-aware and position-sensitive attention enhancement.

[0077] The crop disease and pest recognition model used in the above steps is obtained by training an initial crop disease and pest recognition model with historical crop disease and pest images in the following manner: obtaining historical crop disease and pest images, wherein the historical crop disease and pest images are divided into different categories based on the types of crop disease and pests; determining a training set based on the historical crop disease and pest images, wherein the training set is used to train the initial crop disease and pest recognition model; training the initial crop disease and pest recognition model based on the training set until a preset number of iterations is reached to obtain a crop disease and pest recognition model.

[0078] The following are specific embodiments:

[0079] Taking 32 different plant disease images such as apple scab leaves, apple black rot leaves, peach bacterial spot, corn gray spot, etc. as the research objects, historical crop disease and insect pest images are obtained. The historical crop disease and insect pest images are divided into different categories based on the types of crop diseases and insect pests (i.e. the above-mentioned apple scab leaves, apple black rot leaves, peach bacterial spot, corn gray spot, etc., a total of 32 different plant diseases). Each type is stored in a dedicated folder. The name of each folder is the label of the category image. All image labels are annotated by agricultural and forestry experts. The above-mentioned historical crop disease and insect pest images are divided into training set, validation set and test set, with a division ratio of 8:1:1. For example, there are 38,000 historical crop disease and insect pest images, then the training set data is 30,400, the validation set data is 3,800, and the test set data is 3,800. The training set is a data set used to train the model and determine the model weights; the validation set is a sample set set aside during the model training process. The results of the validation set can be used to adjust the model's hyperparameters, determine the network structure, and make a preliminary evaluation of the model's effects; the test set is used to test the generalization ability of the final model. Before training the initial crop pest and disease recognition model based on the training set, it is necessary to preprocess the size of the input image corresponding to the training set, use bilinear interpolation to adjust the size of the training set to the size required by the initial crop pest and disease recognition model, and then train the initial crop pest and disease recognition model based on the training set until the preset number of iterations is reached to obtain the crop pest and disease recognition model. During training, the learning rate is a very important parameter. If the learning rate is set too large, the learning speed will be accelerated, but the loss function will oscillate or even deviate from the minimum point. If the learning rate is set too small, the network convergence will be very slow. In order to balance the convergence rate and accuracy of the model, the learning rate is set to a dynamic learning rate. The initial learning rate is a preset value (e.g., 0.001). Every preset number of iterations (e.g., 10 epochs, where epoch is the number of rounds of operation, i.e., the preset number of iterations mentioned above, the larger the epoch value, the more training times, which is set to 50 times. An epoch represents a complete data set passing through the neural network and returning, i.e., all training sets are input into the initial crop pest and disease recognition model once), the learning rate is reduced to 1 / 10 of the original value. This ensures that the learning rate is large at the beginning of the training, so that the model can find a smaller point in the loss function quickly; the learning rate is small in the later stage of training to avoid oscillation, and the learning rate is dynamically adjusted to achieve the best prediction results. The batch size is set to the batch preset value (e.g., 10). Specifically, the data set will be divided into multiple batches, and the number of samples contained in each batch is the batch size. The model will perform forward propagation and backward propagation on each batch, calculate the gradient of the loss function, and then update the model parameters based on these gradients.The larger the batch size value, the more memory is required on the computing platform. If the batch size is too small, the model will not converge. The model optimizer selects the Adaptive Moment Estimation optimizer (Adam optimizer for short). The Adam optimizer takes into account the first moment estimation (First Moment Estimation, i.e. the mean of the gradient) and the second moment estimation (Second Moment Estimation, i.e. the uncentered variance of the gradient) to calculate the update step size.

[0080] The present application also provides a schematic diagram of a device for identifying crop pests and diseases. Fig. 9 As shown, including:

[0081] The acquisition module 902 is used to acquire the image to be recognized.

[0082] The recognition module 904 is used to input the image to be recognized into the crop disease and pest recognition model to obtain a recognition result, wherein the crop disease and pest recognition model is obtained by training an initial crop disease and pest recognition model with historical crop disease and pest images, and the initial crop disease and pest recognition model is a deep learning model provided with an attention module, wherein the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be recognized.

[0083] The recognition module 904 is also used to input the image to be recognized into the input layer of the crop disease and pest recognition model; perform feature extraction on the image to be recognized through the convolution layer of the crop disease and pest recognition model to obtain a feature map; process the feature map based on the attention module to determine the attention weight, wherein the attention weight includes the horizontal attention weight and the vertical attention weight; recognize the image to be recognized based on the fully connected layer and the attention weight of the crop disease and pest recognition model to obtain a recognition result.

[0084] The recognition module 904 is also used to perform global average pooling on the feature map in the horizontal and vertical directions of the coordinate axis, respectively, to obtain a first horizontal feature vector corresponding to the horizontal direction and a first vertical feature vector corresponding to the vertical direction, wherein the horizontal direction corresponds to the width of the feature map and the vertical direction corresponds to the height of the feature map; to fuse the first horizontal feature vector and the first vertical feature vector to obtain a fused feature vector; and to process the fused feature vector to obtain an attention weight.

[0085] The recognition module 904 is also used to splice the first horizontal feature vector and the first vertical feature vector to obtain a spliced ​​vector; perform convolution dimensionality reduction on the spliced ​​vector based on the convolution kernel, and activate the spliced ​​vector after convolution dimensionality reduction based on the first preset activation function to obtain a fused feature vector.

[0086] The recognition module 904 is also used to segment the fused feature vector along the spatial dimension to obtain a second horizontal feature vector and a second vertical feature vector; based on a second preset activation function, the second horizontal feature vector and the second vertical feature vector are activated respectively to obtain a horizontal attention weight and a vertical attention weight.

[0087] The recognition module 904 is also used to weight each channel in the feature map using the attention weight to obtain a weighted feature map; and predict the weighted feature map based on the fully connected layer to obtain a recognition result.

[0088] The recognition module 904 is also used to use the fully connected layer to perform classification prediction on the weighted feature map to obtain an output vector, wherein each element of the output vector corresponds to a predicted value of a preset type of crop disease and pest; the output vector is processed by a third preset activation function to obtain a probability value corresponding to each element; and the preset type of crop disease and pest corresponding to the element with the highest probability value is determined as the recognition result.

[0089] It should be noted that Fig. 9 The crop pest identification device shown is used to perform Figure 2 The identification method of crop pests and diseases is shown, so Figure 2 The relevant explanations in the method for identifying crop diseases and insect pests are also applicable to the device for identifying crop diseases and insect pests, and will not be repeated here.

[0090] It should be noted that the various modules in the above-mentioned crop pest and disease identification device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0091] The embodiment of the present application also provides a non-volatile storage medium, the non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above crop pest identification method. For example, an image to be identified is obtained; the image to be identified is input into a crop pest identification model to obtain an identification result, wherein the crop pest identification model is obtained by training an initial crop pest identification model with historical crop pest images, and the initial crop pest identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weights of each channel in the feature map corresponding to the image to be identified.

[0092] The embodiment of the present application also provides an electronic device, the electronic device includes a processor, the processor is used to run a program, wherein the above crop pest identification method is executed when the program is running. For example, an image to be identified is obtained; the image to be identified is input into a crop pest identification model to obtain an identification result, wherein the crop pest identification model is obtained by training an initial crop pest identification model with historical crop pest images, and the initial crop pest identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be identified.

[0093] According to another aspect of the embodiment of the present application, a computer program product is also provided, including a computer program, which implements the above crop pest identification method when executed by a processor. For example, an image to be identified is obtained; the image to be identified is input into a crop pest identification model to obtain an identification result, wherein the crop pest identification model is obtained by training an initial crop pest identification model with historical crop pest images, and the initial crop pest identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be identified.

[0094] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0095] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0096] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0097] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or all or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), mobile hard disk, disk or optical disk, etc. Various media that can store program codes.

[0099] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for identifying crop pests and diseases, characterized in that: include: Obtain an image to be recognized; The image to be identified is input into a crop disease and pest identification model to obtain an identification result, wherein the crop disease and pest identification model is obtained by training an initial crop disease and pest identification model with historical crop disease and pest images, and the initial crop disease and pest identification model is a deep learning model provided with an attention module, and the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be identified.

2. The method according to claim 1, characterized in that The step of inputting the image to be identified into a crop pest identification model to obtain an identification result includes: Inputting the image to be identified into the input layer of the crop pest and disease identification model; Extracting features of the image to be identified through the convolutional layer of the crop pest and disease identification model to obtain a feature map; Processing the feature map based on the attention module to determine an attention weight, wherein the attention weight includes a horizontal attention weight and a vertical attention weight; The image to be identified is identified based on the fully connected layer of the crop disease and pest identification model and the attention weight to obtain the identification result.

3. The method according to claim 2, characterized in that The processing of the feature map based on the attention module to obtain an attention weight includes: Perform global average pooling on the feature map according to the horizontal direction and the vertical direction of the coordinate axis, respectively, to obtain a first horizontal feature vector corresponding to the horizontal direction and a first vertical feature vector corresponding to the vertical direction, wherein the horizontal direction corresponds to the width of the feature map, and the vertical direction corresponds to the height of the feature map; fusing the first horizontal feature vector and the first vertical feature vector to obtain a fused feature vector; The fused feature vector is processed to obtain the attention weight.

4. The method according to claim 3, characterized in that The fusing the horizontal feature vector and the vertical feature vector to obtain a fused feature vector includes: splicing the first horizontal feature vector and the first vertical feature vector to obtain a spliced ​​vector; The convolutional dimension reduction is performed on the concatenated vector based on the convolution kernel, and the concatenated vector after the convolutional dimension reduction is activated based on a first preset activation function to obtain the fused feature vector.

5. The method according to claim 3, characterized in that: The processing of the fused feature vector to obtain the attention weight includes: Slicing the fused feature vector along the spatial dimension to obtain a second horizontal feature vector and a second vertical feature vector; Based on a second preset activation function, the second horizontal feature vector and the second vertical feature vector are activated respectively to obtain a horizontal attention weight and a vertical attention weight.

6. The method according to claim 2, characterized in that The fully connected layer based on the crop pest and disease recognition model and the attention weights recognize the image to be recognized to obtain the recognition result, including: Using the attention weights to weight each channel in the feature map to obtain a weighted feature map; The weighted feature map is predicted based on the fully connected layer to obtain the recognition result.

7. The method according to claim 6, characterized in that The predicting the weighted feature map based on the fully connected layer to obtain the recognition result includes: Using the fully connected layer to perform classification prediction on the weighted feature map to obtain an output vector, wherein each element of the output vector corresponds to a predicted value of a preset crop pest type; Processing the output vector by a third preset activation function to obtain a probability value corresponding to each element; The preset crop disease and insect pest type corresponding to the element with the highest probability value is determined as the identification result.

8. The method according to claim 1, characterized in that The crop pest and disease recognition model is obtained by training the initial crop pest and disease recognition model with historical crop pest and disease images in the following manner: Acquire the historical crop disease and insect pest images, wherein the historical crop disease and insect pest images are divided into different categories based on the types of crop disease and insect pests; Determining a training set based on the historical crop pest and disease images, wherein the training set is used to train the initial crop pest and disease recognition model; The initial crop disease and pest identification model is trained based on the training set until a preset number of iterations is reached to obtain the crop disease and pest identification model.

9. The method according to claim 1, characterized in that: The method also includes: sending identification information associated with the recognition result to a user terminal, wherein the identification information at least includes the image to be identified, the recognition result details, the characteristics of the pest and disease and the control measures, and the recognition result details at least include the image of the pest and disease part and the corresponding pest and disease type.

10. A device for identifying crop pests and diseases, characterized in that: include: An acquisition module, used for acquiring an image to be recognized; The recognition module is used to input the image to be recognized into a crop disease and pest recognition model to obtain a recognition result, wherein the crop disease and pest recognition model is obtained by training an initial crop disease and pest recognition model through historical crop disease and pest images, and the initial crop disease and pest recognition model is a deep learning model provided with an attention module, wherein the attention module is used to determine the weight of each channel in the feature map corresponding to the image to be recognized.

11. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is executed, the device where the non-volatile storage medium is located is controlled to execute the method for identifying crop diseases and insect pests according to any one of claims 1 to 9.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program, when running, executes the method for identifying crop diseases and insect pests according to any one of claims 1 to 9.

13. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method for identifying crop diseases and insect pests according to any one of claims 1 to 9 is implemented.