Method, system and equipment for identifying fine granularity of pests in plant field based on deep learning and storage medium

By using the MaxViT model developed through deep learning and a deformable self-attention mechanism, the problem of identifying field pests with similar morphologies has been solved, achieving high-precision fine-grained classification, which is suitable for monitoring farmland pests and conducting ecological surveys.

CN120913167AActive Publication Date: 2025-11-07ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510992455.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-11-07
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify similar-looking agricultural pests in the field, resulting in low accuracy in pest monitoring and forecasting, and failing to meet the need for fine-grained classification of pests in the field.

Method used

A deep learning-based approach is adopted, which utilizes the MaxViT model combined with a deformable self-attention mechanism and a contrastive learning loss function to train a plant field pest identification model. End-to-end fine-grained classification is performed through convolutional modules, MaxViT-DAT attention modules, global average pooling modules, and fully connected layer modules.

Benefits of technology

It enables accurate identification of various morphologically similar pests with an accuracy rate of over 90%, and is applicable to farmland pest surveys and insect ecology surveys, improving the efficiency and accuracy of pest monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913167A_ABST
    Figure CN120913167A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crop pest intelligent identification systems, in particular to a plant field pest fine granularity identification method, system and device based on deep learning and a storage medium. The identification method provided by the invention is specially used for high-precision identification of a real field scene, and can provide technical support for important work such as development of a field patrol robot and an automatic identification and monitoring system for field pests in the future. Except for pest monitoring, field biological safety tests of transgenic plants are carried out step by step at present, by means of the identification method, dynamic transformation of farmland insect communities can be rapidly and accurately identified and predicted, and the efficiency and precision of ecological investigation are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent recognition system of crop pests, and particularly relates to a plant field pest fine-grained recognition method, system, device and storage medium based on deep learning. BACKGROUND

[0002] Crop pests and diseases are one of the major threats to global agriculture and food security, which has greatly affected global agricultural production. Plants, as one of the main global food crops, have a loss of more than 20% of the total yield affected by pests every year. In order to build sustainable agriculture and maintain global food production safety, it is necessary to accurately and timely predict and identify plant pests. In actual production, plant field pests have very high morphological similarity. For example, three common planthoppers in rice fields, brown planthopper (Nilaparvata lugens Stal) Nilaparvata lugens ), gray planthopper (Laodelphax striatellus Fabricius) Laodelphax striatellus ), and white-backed planthopper (Sogatella furcifera Horváth) Sogatella furcifera ), have different rice planthopper conditions in different regions. Accurate classification of rice planthopper species in the field helps to study the occurrence rules of migratory populations and provides data for local pest prediction, prevention and control. Traditional identification methods rely on manual accurate species classification of pests, which is time-consuming, labor-intensive and low in accuracy. Agricultural workers often misidentify due to lack of experience, environmental changes, and species similarity, which affects the progress of pest monitoring and prediction. Therefore, it is very meaningful to study a pest fine-grained recognition system that meets the needs of plant field work.

[0003] The existing detection systems mainly use remote sensing technology and machine learning technology, which have certain effect in large-scale and intensive agriculture. The high-spectral images of crops taken by remote sensing, high-speed cameras and other devices can monitor the damage of crops in the field to a certain extent, but it is difficult to analyze the specific pest species and conditions of crops. Using trapping and trapping devices to collect field insects and using machine learning methods to identify the trapped insects can monitor the occurrence of field pests to a certain extent, but it is difficult to achieve fine classification of similar morphological pests, and the recognition accuracy is not high.

[0004] In recent years, deep learning technology has developed rapidly and has achieved good results and applications in many fields. Many scholars have also tried to use deep learning methods to classify and identify insects, especially agricultural pests. Compared with traditional methods, using deep learning technology can extract features, identify and classify pests, significantly reducing labor costs and improving the accuracy of pest identification. Currently, there are two main types of intelligent identification technology applied in the agricultural field. One is coarse-grained pest classification based on classification tasks, which divides field pests into major pest groups such as planthoppers, thrips, and moths according to image features. The other is precise detection and identification of a single or small number of insect species, such as detecting and identifying Spodoptera frugiperda in insects captured in a field trap station to monitor the occurrence of Spodoptera frugiperda in the field.

[0005] However, current pest prediction and prevention work requires comprehensive monitoring and investigation of the overall occurrence of field pests. On the one hand, the method needs to cover the main pest groups in the field as much as possible. On the other hand, the method needs to have high recognition accuracy and accurately distinguish between small groups with small inter-species differences to achieve fine-grained identification. Currently, a large number of studies focus on the general classification or single detection of insects, and there are few fine-grained classification techniques for specific crop pest groups. SUMMARY

[0006] The present application provides a plant field pest fine-grained identification method, system, device and storage medium based on deep learning. The identification method can realize fine-grained classification and identification of multiple morphologically similar plant pests in the field, with an identification accuracy of more than 90%.

[0007] To achieve the above purpose, the present application provides the following technical solutions: The present application provides a plant field pest fine-grained identification method based on deep learning, comprising: S1. Data acquisition: collect plant field pest image data, pre-process and normalize to obtain image data, and classify the processed image data so that each image data has its corresponding class label to form a plant field pest image data set; S2. Model training: introduce a deformable self-attention mechanism and a contrastive learning loss function into the MaxViT model, train the MaxViT model using the plant field pest image data set, and obtain a trained plant field pest identification model; S3. Pest identification: acquire plant field pest image data to be identified, pre-process and normalize, and input into the trained plant field pest identification model for fine-grained classification, and output the class result.

[0008] The identification method provided by the application is specialized in high-precision identification of real field scenes, and can provide technical support for important work such as future field patrol robot development and automatic identification and monitoring system of field pests. In addition to pest monitoring, field biological safety tests of genetically modified plants are gradually carried out, and the identification method provided by the application can quickly and accurately identify and predict the dynamic changes of insect communities in farmland, greatly improving the efficiency and accuracy of ecological investigation.

[0009] In summary, the application has more adaptive application performance and wider application scenarios in the fields of pest integrated management and biological ecology investigation.

[0010] Preferably, in S1, the plant field pest image data includes pest images of multiple families corresponding to a certain plant, the inter-family differences are large, and the intra-family differences are small, and the morphological characteristics of various pests in each family are similar.

[0011] According to specific plants, relevant pest images can be selected to obtain a data set suitable for different plants. For example, the rice field pest image data can select the pests commonly found in rice fields.

[0012] Preferably, the plant is rice.

[0013] Preferably, in S1 and S3, The pre-processing method is: 1) uniformly adjusting the resolution of the images in the plant field pest image data; 2) after adjusting the resolution, the image is further subjected to data enhancement processing to obtain pre-processed image data; The enhancement processing is multiple of random flipping, random rotation, random cropping, random brightness, random contrast, random saturation, random hue, simulated light, random size, Gaussian blur, and Gaussian noise. The normalization processing method is: the pre-processed image data is converted into a three-dimensional tensor in RGB format; the three-dimensional tensor is normalized to obtain pre-processed and normalized three-dimensional tensor data, i.e. image data.

[0014] Since the resolutions of the plant field pest images in the public data set and the self-built data set are different, in order to better train the model, the size of all images is uniformly adjusted to 224x224 pixels through upsampling and downsampling, which is consistent with the input requirements of the MaxViT model.

[0015] In order to improve the robustness of the plant field pest identification model, the online data enhancement method is used to expand the plant field pest images in the public data set and the self-built data set during the training process of the plant field pest identification model.

[0016] Specifically, for example: randomly flipping the image horizontally and vertically with a 50% probability; randomly adjusting the brightness, contrast, and saturation of the image within a range of ±20%, and randomly adjusting the hue of the image within a range of ±10%; randomly adjusting the size of the image within a range of 80% to 100% in area and 90% to 110% in aspect ratio; randomly applying Gaussian blur to the image within a range of 0.1 to 2.0 in standard deviation; and randomly adding Gaussian noise to the image with a standard deviation of 0.005.

[0017] Preferably, in S2, the plant field pest recognition model comprises a convolution module, a MaxViT-DAT attention module, a global average pooling module, a fully connected layer module, and a classification module. The convolution module is configured to process image data in the plant field pest image data set, so that the resolution of each image data is halved and the number of channels is increased, and a preliminary feature spectrum is extracted. The MaxViT-DAT attention module is configured to process the preliminary feature spectrum and sequentially extract feature spectra through layers of MaxViT multi-axis attention networks to obtain multiple groups of feature spectra with resolutions decreasing by a factor of 2. The global average pooling module is configured to perform global average pooling processing on the feature spectra extracted by the MaxViT-DAT attention module to obtain a feature vector with a fixed dimension. The fully connected layer module is configured to learn the mapping relationship between the feature vector and the categories, calculate the score of each category, and output an unnormalized score vector. The classification module is configured to normalize the unnormalized score vector using a function to convert it into a probability value, which is the confidence of the plant field pest category corresponding to each plant field pest image in the plant field pest image data set input into the MaxViT model, to obtain a prediction result for each plant field pest category.

[0018] The plant field pest recognition model works cooperatively through the hierarchical modules to realize end-to-end processing from the input of raw images to the output of fine-grained classification results of plant field pests. The convolution module first processes the input pest image data: through a convolution operation with a step size of 2, the image resolution is halved (e.g., 224x224 pixels → 112x112 pixels), while the number of channels is expanded (e.g., RGB three channels → 64 channels), generating a preliminary feature spectrum containing basic texture and edge features.

[0019] After receiving the preliminary feature map, the MaxViT-DAT attention module performs deep feature extraction through a 4-layer multi-axis attention network: the resolution of each layer of the network decreases by 2 times (112x112 pixels → 56x56 pixels → 28x28 pixels → 14x14 pixels → 7x7 pixels), forming a multi-scale feature pyramid.

[0020] The global average pooling module compresses the spatial dimensions of the final 7x7 feature map: the average of all feature values in each channel is taken to generate a fixed-dimension feature vector (such as 512 dimensions), preserving high-order semantic information.

[0021] The fully connected layer module learns the mapping relationship between the feature vector and the pest class: it outputs an unnormalized score vector, with each element corresponding to the original score of a specific pest class.

[0022] The classification module normalizes the score vector through the Softmax function: it converts the score vector into a probability distribution, where the probability value represents the model's confidence in each class, and finally takes the maximum value index to map to the pest fine-grained classification unit name.

[0023] Preferably, in the MaxViT-DAT attention module: each layer of the MaxViT multi-axis attention network includes a mobile inverted convolution submodule, a local attention submodule, a global attention submodule, a feedforward network submodule, and a residual connection submodule, and a deformable self-attention submodule embedded in the local attention submodule and the global attention submodule; The mobile inverted convolution submodule is used to enhance the features of the input preliminary feature map, and outputs an enhanced feature map with the same number of channels and resolution as the input; The local attention submodule is used to divide the enhanced feature map into multiple non-overlapping and equal-sized windows, perform self-attention calculation in each window, and integrate the deformable self-attention submodule to dynamically adjust the sampling position of feature information, capturing the feature information of the local region of the enhanced feature map, and outputting a local feature map with the same number of channels and resolution as the input; The global attention submodule is used to divide the local feature map into uniformly distributed and equal-sized global grids, perform self-attention calculation in each grid, and integrate the deformable self-attention submodule to realize cross-region context feature fusion, outputting a global feature map with the same number of channels and resolution as the input; The deformable self-attention submodule dynamically adjusts the sampling position of feature information through vector reorganization, reference point introduction, offset learning, and attention aggregation with offset sampling, enabling the plant field pest recognition model to adaptively focus on key regions and enhancing the fine-grained classification ability of the plant field pest recognition model; The feedforward network submodule is used for performing nonlinear transformation on the global feature map, improving the discrimination degree of subtle feature information, and outputting a feature map with the same number of output channels and resolution as the input; The residual connection submodule is used for constructing an information protection mechanism in each submodule of the MaxViT multi-axis attention network; the features input into the submodule are normalized through layer normalization, the normalized input features are processed through the submodule to obtain output features, and the output features of the submodule are added to the input features before normalization to form a residual connection.

[0024] In the MaxViT-DAT attention module, each layer of the multi-axis attention network processes the feature map through the precise collaborative submodule to realize fine-grained feature extraction of plant pests.

[0025] The local attention captures detailed features such as the texture of the pest wing and the shape of the back plate; the global attention establishes context association such as the interaction between the pest and the environment; the dynamic offset sampling adaptively focuses on the key area such as the dense wing spot of the brown planthopper; and the residual connection protects the original morphological information from being lost in each operation.

[0026] Preferably, the feature vector obtained through the global average pooling processing is calculated by a contrast learning loss function to obtain a contrast learning loss; at the same time, the feature vector obtained through the global average pooling processing is calculated by a fully connected layer module and a classification module to obtain a classification loss; the contrast learning loss and the classification loss are substituted into formula (1) to obtain a total loss; the total loss is used for back propagation to optimize the parameters of the plant field pest recognition model, and the total loss minimization optimizes the feature space distribution to improve the accuracy of the prediction result; In formula (1), the weight ranges from 0.9 to 1.1; (1); In the formula, Ltotal is the total loss; Lcontrast is the contrast learning loss; Lclass is the classification loss; is the weight. The value of the weight

[0027] is continuously improved, and the total loss value after calculation reaches the minimum. Specifically, the total loss is used for back propagation to optimize the model parameters, and the model parameters are adjusted by the optimizer Adam to realize the minimum total loss.

[0028] Preferably, in S1, the category result includes: the category name of the fine-grained species of the plant field pest in the plant field pest image data to be identified.

[0029] The present application provides a plant pest fine-grained recognition system, comprising: ​The data acquisition module is configured to collect plant field pest image data, perform preprocessing and normalization processing, and obtain standardized image data. The model processing module includes a plant field pest recognition model based on a MaxViT model architecture, which integrates a deformable self-attention mechanism and a contrastive learning loss function. The pest recognition module is configured to receive plant field pest image data to be recognized, perform preprocessing and normalization processing, and perform fine-grained classification through the plant field pest recognition model to output a category result.

[0030] The application provides a plant pest recognition device, which includes a memory, a processor, and a plant pest recognition program stored on the memory and executable on the processor.

[0031] The application provides a storage medium, which stores a computer program executable on a processor to implement the fine-grained plant field pest recognition method based on deep learning.

[0032] Therefore, the application has the following advantages: (1) The recognition method provided by the application realizes fine recognition of various rice field pests, can better distinguish species with similar morphologies, can accurately recognize pests in complex fields, and is conducive to timely disease control in actual agricultural production.

[0033] (2) The application trains a plant field pest recognition model by embedding a deformable self-attention submodule in the local attention submodule and the global attention submodule of the MaxViT-DAT attention module based on the MaxViT model architecture. The deformable self-attention mechanism in the deformable self-attention submodule makes the feature extraction of the plant field pest recognition model more accurate, and improves the fine-grained classification ability of the model.

[0034] (3) The application introduces a contrastive learning loss function to calculate a contrastive learning loss while calculating a classification loss through a fully connected layer module and a classification module based on the feature vector obtained through global average pooling processing, introduces a classification loss weight to calculate a total loss, and obtains a total loss that optimizes the model to minimize the final total loss. Minimizing the total loss optimizes the feature space distribution and improves the accuracy of the prediction result. The contrastive learning optimizes the feature space of the same samples to gather and the different samples to separate, enhances the feature discriminability, and forces the plant field pest recognition model to learn the subtle difference features, so that the plant field pest recognition model is more accurate in recognizing the subtle difference features.

[0035] (4) The plant field pest recognition model trained by the method has higher recognition accuracy, and an accuracy of more than 90% is achieved on a test data set composed of various similar plant pests; and the recognition ability is stronger when facing insects with similar morphologies.

[0036] (5) The plant field pest recognition model trained by the method performs well in the morphologically similar pest classification task, and is suitable for the technical needs of future farmland pest investigation and insect ecological investigation. BRIEF DESCRIPTION OF DRAWINGS

[0037] Figure 1 FIG. 1 is a flowchart of a plant field pest fine-grained recognition method based on deep learning; Figure 2 FIG. 2 is a structural diagram of a plant field pest recognition model; Figure 3 FIG. 3 is a model structure diagram of a deformable self-attention sub-module; Figure 4 FIG. 4 is a structural diagram of a plant pest fine-grained recognition system; Figure 5 FIG. 5 is a structural diagram of a plant pest recognition device; Figure 6 FIG. 6 is another structural diagram of a plant pest recognition device. DETAILED DESCRIPTION

[0038] The present application will be further described below in conjunction with specific embodiments. Those skilled in the art will be able to implement the present application based on these descriptions. In addition, the embodiments of the present application involved in the following description are generally only embodiments of a part of the present application, not all embodiments. Therefore, based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor shall fall within the scope of protection of the present application.

[0039] Embodiment 1: Recognition method The present embodiment provides a plant field pest fine-grained recognition method based on deep learning, S1. Data acquisition: the plant is rice; S1.1. Collection: collect rice field pest image data, i.e. self-built data set.

[0040] The self-built dataset is a fine-grained classification dataset of rice pests, obtained by combining images of common rice paddy pests photographed in rice fields in Hangzhou, Zhejiang Province, between 2024 and 2025 with related images from the internet and after expert classification and identification. The dataset contains 15,982 training images and 3,985 validation images, including image data of 20 species of rice-related pests from 8 families. The morphological characteristics of various pests within the same family are similar. The related pests are: rice green stink bug, rice edged stink bug, rice heterocercus stink bug, brown planthopper, gray planthopper, white-backed planthopper, rice leaf roller, rice stem borer, rice leafcutter, rice red-spotted black-spotted leafhopper, rice straight-striped skipper, rice evening-eye butterfly, large green leafhopper, black-tailed leafhopper, fall armyworm, rice golden-winged noctuid moth, yellow cutworm, gray-winged noctuid moth, beet armyworm, and armyworm.

[0041] When creating the self-built dataset, images of common pests found in rice paddies were selected for collection, with unprocessed natural backgrounds. Images of various crop pests were chosen that showed significant interfamilial differences but minimal intrafamilial differences.

[0042] S1.2. Classification: Images from the rice field pest image data are stored in corresponding folders according to their label categories. The folder name is the category label name corresponding to the image. Then, the training images and validation images are merged, and the training set, validation set and test set are re-divided in a stratified sampling method at a ratio of 7:2:1.

[0043] S1.3. Data Processing: Preprocess and normalize all images after dividing them into training, validation and test sets.

[0044] (1) Use the Resize method to uniformly adjust the resolution of all images to 224×224 pixels.

[0045] (2) Perform data augmentation processing on the adjusted image, and perform random horizontal flip (probability 0.5), vertical flip (probability 0.5), ±30° random rotation, color jitter (brightness / contrast / saturation ±0.2, hue ±0.1), random cropping (scale 0.8~1.0, scale 0.9~1.1), and Gaussian blur (kernel size 3, sigma 0.1~2.0).

[0046] (3) Convert the enhanced images into three-dimensional tensors in RGB format. Use ImageNet standard normalization: normalize with mean vector [0.485, 0.456, 0.406] and standard deviation vector [0.229, 0.224, 0.225]. Finally, form a rice field pest image dataset, which consists of three subsets: training set, validation set, and test set in a 7:2:1 ratio.

[0047] S2. Model training: training the MaxViT model using the rice field pest image dataset obtained in S1 to obtain a plant field pest recognition model 1.

[0048] The key parameters of the MaxViT model are configured as follows: the number of classification categories is 20, the feature dimension is 64, the network depth adopts a hierarchical design of (2, 2, 2), the attention head dimension is 32, the convolution stem dimension is 64, the window size is set to 7, the expansion rate of the mobile inverted convolution submodule 121 is 4, the shrinkage rate is 0.25, and the global recycling rate is 0.1 to reduce overfitting.

[0049] The plant field pest recognition model 1 includes a convolution module 110, a MaxViT-DAT attention module 120, a global average pooling module 130, a fully connected layer module 140, and a classification module 150. The optimization training of the plant field pest recognition model 1 is completed in each module. Specifically as follows: The convolution module 110 is composed of two convolution layers. The first convolution layer has a convolution kernel size of 3x3, a step size of 2, a padding of 1, and an output channel number of 64. The second convolution layer has a convolution kernel size of 3x3, a step size of 1, a padding of 1, and an output channel of 64. After passing through the convolution module 110, the image resolution is halved and the channel number is increased to obtain a preliminary feature map.

[0050] The preliminary feature map extracted by the convolution module 110 is input into the 4-layer MaxViT multi-axis hierarchical attention network in the MaxViT-DAT attention module 120 for processing. After inputting the preliminary feature map, the module obtains feature maps with resolutions of 56x56 pixels, 28x28 pixels, 14x14 pixels, and 7x7 pixels in turn. Each layer of the MaxViT multi-axis attention network includes a mobile inverted convolution submodule 121, a local attention submodule 122, a global attention submodule 123, a feedforward network submodule 125, and a residual connection submodule 126, as well as a deformable self-attention submodule 124 embedded in the local attention submodule 122 and the global attention submodule 123.

[0051] The preliminary feature map first enters the mobile inverted convolution submodule 121, which uses a depth separable convolution and a compression-excitation network module to process the features. The enhanced feature map has the same channel number and resolution as the input.

[0052] The enhanced feature map is input into the local attention submodule 122, which divides the feature map into multiple non-overlapping windows with a size of PxP. Self-attention calculation is performed within each window, and the deformable self-attention submodule 124 dynamically adjusts the sampling position of the feature information to capture the local region feature information of the enhanced feature map. The output local feature map has the same channel number and resolution as the input.

[0053] The local feature map processed by the local attention sub-module 122 is input into the global attention sub-module 123, the global attention sub-module 123 divides the local feature map into a global grid with a size of GxG, performs self-attention calculation on each grid, and integrates the deformable self-attention sub-module 124 to realize cross-region context feature fusion and global information interaction, and outputs a global feature map with the same channel number and resolution as the input.

[0054] The deformable self-attention sub-module 124, the input preliminary feature map is first subjected to a vector reorganization operation and converted into an input feature map, and a feature vector and a reference point are introduced. The input feature map is then subjected to a linear transformation to prepare for subsequent processing. The feature after the linear transformation enters the offset module. The offset module has multiple offset learning heads (offset head 1, offset head 2, offset head 3, etc.) inside. Each head independently learns an offset, and based on these offsets, the feature values in the feature map are sampled with an offset, allowing the feature sampling to better meet actual needs. The feature processed by the offset module enters the attention module. In the deformable self-attention sub-module 124, through vector reorganization, reference point introduction, offset learning, and attention aggregation with offset sampling, the sampling position of the feature information is dynamically adjusted, allowing the plant field pest recognition model 1 to adaptively focus on key areas and enhance the fine-grained classification ability of the plant field pest recognition model 1.

[0055] The global feature map processed by the global attention sub-module 123 is input into the feedforward network sub-module 125, which uses the GELU activation function for further nonlinear transformation to improve the discrimination of subtle feature information, and outputs a feature map with the same channel number and resolution as the input.

[0056] In the 4-layer MaxViT multi-axis attention network, normalization processing is used to realize the residual connection of all sub-modules (except the residual connection sub-module 126). The residual connection sub-module 126 normalizes the input features of the sub-module through layer normalization, and the output features of the sub-module are obtained by processing the normalized input features. The output features of the sub-module are added to the input features before normalization to form a residual connection. Taking the local attention sub-module 122 as an example, the features input into the local attention sub-module 122 are first normalized by applying a layer normalization function before entering the module; then, the features are processed by the module to obtain the output; finally, the output of the module is added to the input features after layer normalization before entering the module and the connection features are output, thereby realizing the residual connection and allowing the information to be preserved and transmitted before and after the module processing.

[0057] After multi-layer extraction in the MaxViT-DAT attention module 120, the feature spectrum is input into the global average pooling module 130 for global average pooling processing. The global average pooling processing performs a feature value averaging operation on each channel of the feature spectrum along the spatial dimension to generate a fixed-dimension feature vector.

[0058] The obtained feature vector is input into the full connection layer module 140, and the full connection layer learns the mapping relationship between the feature vector and the category, calculates the score of each category, and outputs an unnormalized score vector.

[0059] In the classification module 150, the unnormalized score vector is converted into a probability distribution using a softmax activation function. Specifically, the softmax function converts the score vector into a probability value between 0 and 1, which is the confidence of the rice field pest image corresponding to each rice field pest image in the rice field pest image data set input into the MaxViT model, and obtains the prediction result of each rice field pest category.

[0060] The feature vector obtained by the global average pooling processing is used to calculate the contrastive learning loss by a contrastive learning loss function; at the same time, the feature vector obtained by the global average pooling processing is used to calculate the classification loss by the full connection layer module 140 and the classification module 150; the contrastive learning loss and the classification loss are substituted into formula (1) to calculate the total loss. The total loss is used for back propagation to optimize the model parameters, and the model parameters are adjusted by the optimizer Adam to minimize the total loss, so as to repeatedly optimize the accuracy of the plant field pest recognition model 1.

[0061] The specific parameters of the Adam optimizer are: the initial learning rate is set to 1e -4 , the batch size during training is 16, and the maximum number of iterations is 200. The learning rate adjustment uses a stepwise learning rate scheduler with a step size of 50 and a decay factor of 0.5, i.e., the learning rate is automatically multiplied by 0.5 for decay every 50 iterations.

[0062] In formula (1), the weight ranges from 0.9 to 1.1; (1); In the formula, is the total loss; is the contrastive learning loss; is the classification loss; is the weight.

[0063] S3. Pest identification: Obtain the rice field pest image data to be identified. Perform preprocessing and normalization processing: (1) Use the Resize method to adjust the resolution of the image to 224x224 pixels.

[0064] (2) Perform data augmentation processing on the adjusted image, sequentially performing random horizontal flipping (probability 0.5), vertical flipping (probability 0.5), ±30° random rotation, color jittering (brightness / contrast / saturation ±0.2, hue ±0.1), random cropping (scale 0.8~1.0, ratio 0.9~1.1), and Gaussian blur (kernel size 3, sigma 0.1~2.0).

[0065] (3) Convert the enhanced image to a three-dimensional tensor in RGB format, and perform ImageNet standard normalization: normalize with the mean vector [0.485, 0.456, 0.406] and the standard deviation vector [0.229, 0.224, 0.225].

[0066] Input into the trained plant field pest recognition model 1 for fine-grained classification, and output the category result.

[0067] When the weight is 1.0, the accuracy of the final model obtained by the recognition method of Application Example 1 is 0.90103.

[0068] Example 2 Plant field pest recognition model As shown in Figure 2 , the plant field pest recognition model 1 includes a convolution module 110, a MaxViT-DAT attention module 120, a global average pooling module 130, a fully connected layer module 140, and a classification module 150.

[0069] The convolution module 110 is used to process image data in the plant field pest image data set, so that the resolution of each image data is halved and the number of channels is increased, and a preliminary feature map is extracted.

[0070] The MaxViT-DAT attention module 120 is used to process the preliminary feature map, sequentially extracting feature maps through 4 layers of MaxViT multi-axis attention networks 120a to obtain 4 groups of feature maps with resolution decreasing by a factor of 2. Each layer of MaxViT multi-axis attention network 120a includes a mobile inverted convolution submodule 121, a local attention submodule 122, a global attention submodule 123, a feedforward network submodule 125, and a residual connection submodule 126, and a deformable self-attention submodule 124 embedded in the local attention submodule 122 and the global attention submodule 123.

[0071] The mobile inverted convolution submodule 121 is used to enhance the features of the input preliminary feature map, and outputs an enhanced feature map with the same number of channels and resolution as the input.

[0072] The local attention sub-module 122 is configured to divide the enhanced feature map into a plurality of non-overlapping and equal-size windows, perform self-attention calculation in each window, and integrate the deformable self-attention sub-module 124 to dynamically adjust the sampling position of the feature information, capture the feature information of the local area of the enhanced feature map, and output a local feature map with the same number of channels and resolution as the input.

[0073] The global attention sub-module 123 is configured to divide the local feature map into a uniform distribution and equal-size global grid, perform self-attention calculation in each grid, and integrate the deformable self-attention sub-module 124 to realize cross-area context feature fusion, and output a global feature map with the same number of channels and resolution as the input.

[0074] The deformable self-attention sub-module 124 dynamically adjusts the sampling position of the feature information by vector reorganization, reference point introduction, offset learning, and attention aggregation with offset sampling, so that the plant field pest recognition model can adaptively focus on the key area and enhance the fine-grained classification ability of the plant field pest recognition model.

[0075] The feedforward network sub-module 125 is configured to perform nonlinear transformation on the global feature map to improve the discrimination of subtle feature information, and output a feature map with the same number of channels and resolution as the input.

[0076] The residual connection sub-module 126 is configured to construct an information protection mechanism in each sub-module of the MaxViT multi-axis attention network; the input features of the sub-module are normalized by layer normalization, the normalized input features are processed by the sub-module to obtain output features, and the output features of the sub-module are added to the input features before normalization to form a residual connection.

[0077] The global average pooling module 130 is configured to perform global average pooling on the feature spectrum obtained by the MaxViT-DAT attention module to obtain a feature vector with a fixed dimension.

[0078] The fully connected layer module 140 is configured to learn the mapping relationship between the feature vector and the category, calculate the score of each category, and output an unnormalized score vector.

[0079] The classification module 150 is configured to use a function to normalize the unnormalized score vector, and convert the normalized value to a probability value between 0 and 1, which is the confidence of the plant field pest category corresponding to each plant field pest image in the plant field pest image data set input into the MaxViT model, to obtain the prediction result of each plant field pest category.

[0080] Example 3 AsFigure 4 As shown, a fine-grained plant pest identification system 2 includes: Data acquisition module 210 is used to collect image data of plant field pests, perform preprocessing and normalization processing to obtain standardized image data; The model processing module 220 includes a plant field pest identification model based on the MaxViT model architecture, which integrates a deformable self-attention mechanism and a contrastive learning loss function. The pest identification module 230 is used to receive image data of plant field pests to be identified, perform preprocessing and normalization, perform fine-grained classification through the plant field pest identification model, and output the category results.

[0081] Example 4 This invention provides a plant pest identification device, comprising: a memory, a processor, and a plant pest identification program stored in the memory and executable on the processor. When the plant pest identification program is executed by the processor, it implements a fine-grained identification method for plant field pests based on deep learning.

[0082] Specific plant pest identification equipment 4 can be like Figure 5 As shown, the system includes: a memory 410, a processor 420, and a plant pest identification program 430 stored in the memory 410 and capable of running on the processor 420. The memory 410 stores an operating system 440, a user port module 450, a central processing unit 460, and the plant pest identification program 430. Specifically, the user port module 450 may include input / output devices such as a keyboard, mouse, monitor, and camera, for inputting and outputting user signals, images, and other information; the central processing unit 460 is used to process information between modules and handle input / output; the plant pest identification program 430 is an interactive program built around a deep learning-based fine-grained identification method for field pests.

[0083] Another specific plant pest identification device 5 can be like Figure 6 As shown, the system includes a client 500 and a cloud 600. The cloud 600 includes a memory 610 and a processor 620. The client 500 includes an operating system 510 and a user port module 520. The memory 610 stores a network communication module 611, an input / output module 612, a central processing unit 613, and a plant pest identification program 614. Specifically, the client 500 can be a mobile app, application, or webpage; the operating system 510 can be a mobile app operating system or a webpage operating system. Users can input pest images into the client 600, which are transmitted to the cloud 600 via the network communication module 611. The cloud runs the plant pest identification program 614 and outputs the results back to the user, achieving fine-grained identification of plant pests in the field.

[0084] Example 5 A storage medium, comprising: a computer program stored on a computer readable storage medium, the computer program being executed by a processor to implement a plant field pest fine-grained recognition method based on deep learning.

[0085] Comparative Example 1 The recognition method used in this comparative example is basically the same as that of Example 1, except that the contrast learning loss function is not introduced. The accuracy of the final obtained model is 0.87663.

[0086] Comparative Example 2 The recognition method used in this comparative example is basically the same as that of Example 1, except that the deformable self-attention sub-module is not embedded in the local attention sub-module and the global attention sub-module. The accuracy of the final obtained model is 0.85495.

[0087] Comparative Example 3 The recognition method used in this comparative example is basically the same as that of Example 1, except that the rice field pest image data used is a public data set. The public data set is a data set composed of images of 30 rice-related pests selected from the “IP102: Large-scale Benchmark Dataset for Pest Recognition” published by Wuxiaoping et al. of Nankai University in 2019. The data set has a total of 16869 training images and 4826 validation images, including image data of 14 families and 30 rice-related pests. The related pests are: red spider, aphid, wheat two aphid, brown planthopper, small brown planthopper, white-backed planthopper, rice leafhopper, rice armyworm, rice large caterpillar, rice borer, rice borer, rice leaf roller, rice borer, small tiger, yellow tiger, white border moth, Asian corn borer, peach borer, rice stem borer, rice gall midge, rice stem fly, oriental mole cricket, Chinese rice locust, rice weevil (larvae), rice water weevil, grubs, rice thrips, rice green bugs, rice rim bugs.

[0088] The accuracy of the final obtained model is 0.73224.

[0089] Comparative Example 4: Selection of Model Architecture The recognition method used in this comparative example is basically the same as that of Example 1, except that the deformable self-attention sub-module is not embedded in the local attention sub-module and the global attention sub-module; the contrast learning loss function is not introduced; the MaxViT model is replaced by other models, and the specific content of the replacement is shown in Table 1.

[0090] According to the results in Table 1, the best model selected is the MaxViT model, with an accuracy of 0.84071.

[0091] Table 1: Replacement models and their corresponding accuracies

[0092] Determination of the weight of the total loss The identification method used in this comparative example is basically the same as that in Example 1, except that the weight in the contrast loss learning function is changed , and the total loss value corresponding to the change is shown in Table 2. Different weights are substituted until the minimum total loss value is obtained. It is determined that the optimal weight is 1.0, and within this range, the total loss value is the smallest and the accuracy is the highest; a small total loss value can help improve the identification accuracy of the rice field pest identification model.

[0093] (1) ; wherein, is the total loss, unit; is the contrast learning loss, unit; is the classification loss, unit; is the weight.

[0094] Table 2 Different weights and their corresponding training accuracy

Claims

1. A method for identifying plant field pests in a fine-grained manner based on deep learning, characterized in that, The application relates to a plant field pest image recognition method and device. S1. Data acquisition: plant field pest image data is collected, preprocessed and normalized to obtain image data, and the processed image data is classified so that each image data has a corresponding category label, thereby forming a plant field pest image data set; S2. Model training: a deformable self-attention mechanism and a contrastive learning loss function are introduced into a MaxViT model, the MaxViT model is trained using the plant field pest image data set, and a trained plant field pest recognition model is obtained; S3. Pest recognition: plant field pest image data to be recognized is preprocessed and normalized, and then input into the trained plant field pest recognition model for fine-grained classification, and a category result is output.

2. The identification method of claim 1, wherein, In S1, the plant field pest image data includes pest images of multiple families corresponding to a certain plant, and the inter-family differences are large while the intra-family differences are small, and the morphological characteristics of various pests in each family are similar; the plant is rice.

3. The identification method of claim 1, wherein, In S1 and S3, The pre-processing mode is as follows: 1) uniform resolution of the images in the plant field pest image data; 2) after adjusting the resolution, the images are subjected to data enhancement processing to obtain pre-processed image data; The enhancement processing is multiple kinds of random flipping, random rotation, random cropping, random brightness, random contrast, random saturation, random hue, simulated illumination, random size, Gaussian blur and Gaussian noise; The normalization processing mode is as follows: the pre-processed image data is converted into a three-dimensional tensor in RGB format; the three-dimensional tensor is normalized to obtain pre-processed and normalized three-dimensional tensor data, that is, image data.

4. The identification method of claim 1, wherein, In S2, the plant field pest recognition model comprises a convolution module, a MaxViT-DAT attention module, a global average pooling module, a full connection layer module and a classification module; The convolution module is used for processing the image data in the plant field pest image data set, so that the resolution of each image data is halved and the channel number is increased, and a preliminary feature spectrum is extracted; The MaxViT-DAT attention module is used for processing the preliminary feature spectrum, and the feature spectrum is sequentially extracted through each layer of MaxViT multi-axis attention network to obtain multiple groups of feature spectra with resolution decreasing by a factor of 2; The global average pooling module is used for performing global average pooling processing on the feature spectrum extracted by the MaxViT-DAT attention module to obtain a fixed-dimension feature vector; The full connection layer module is used for learning the mapping relationship between the feature vector and the category, calculating the score of each category, and outputting an unnormalized score vector; The classification module is used for normalizing the unnormalized score vector using a function to convert it into a probability value, that is, the confidence of the plant field pest category corresponding to each plant field pest image in the plant field pest image data set input into the MaxViT model, to obtain the prediction result of each plant field pest category.

5. The identification method of claim 4, wherein, In the MaxViT-DAT attention module: each layer of the MaxViT multi-axis attention network comprises a mobile inverted convolution submodule, a local attention submodule, a global attention submodule, a feedforward network submodule, and a residual connection submodule, and a deformable self-attention submodule embedded in the local attention submodule and the global attention submodule; The mobile inverted convolution submodule is configured to perform feature enhancement on the input preliminary feature map, and output an enhanced feature map with consistent channel number and resolution as the input; The local attention submodule is configured to divide the enhanced feature map into a plurality of non-overlapping and equal-size windows, perform self-attention calculation in each window, and integrate the deformable self-attention submodule to dynamically adjust the sampling position of the feature information, capture the feature information of the local area of the enhanced feature map, and output a local feature map with consistent channel number and resolution as the input; The global attention submodule is configured to divide the local feature map into a plurality of global grids with uniform distribution and equal size, perform self-attention calculation in each grid, and integrate the deformable self-attention submodule to realize cross-area context feature fusion, and output a global feature map with consistent channel number and resolution as the input; The deformable self-attention submodule dynamically adjusts the sampling position of the feature information by vector reorganization, reference point introduction, offset learning, and attention aggregation with offset sampling, so that the plant field pest recognition model can adaptively focus on the key area and enhance the fine-grained classification ability of the plant field pest recognition model; The feedforward network submodule is configured to perform nonlinear transformation on the global feature map to improve the discrimination of subtle feature information, and output a feature map with consistent channel number and resolution as the input; The residual connection submodule is configured to build an information protection mechanism in each submodule of the MaxViT multi-axis attention network; the features of the input submodule are normalized by layer normalization, the output features are obtained by processing the normalized input features by the submodule, and the output features of the submodule are added to the input features before normalization to form a residual connection.

6. The identification method of claim 4 or 5, wherein The feature vector obtained by the global average pooling processing is used to calculate a contrast learning loss by a contrast learning loss function; meanwhile, the feature vector obtained by the global average pooling processing is used to calculate a classification loss by a fully connected layer module and a classification module; the contrast learning loss and the classification loss are substituted into formula (1) to obtain a total loss; the total loss is used for back propagation to optimize the parameters of the plant field pest recognition model, and the total loss minimization optimizes the feature space distribution and improves the accuracy of the prediction result; In formula (1), the weight ranges from 0.9 to 1.1; (1); wherein is the total loss; is the contrastive learning loss; is the classification loss; is the weight.

7. The identification method of claim 1, wherein, In S1, the category result comprises a category name of a fine-grained species of the plant field pest in the plant field pest image data to be identified.

8. A plant pest fine particle discrimination system characterized by, The data acquisition module is configured to collect plant field pest image data, perform preprocessing and normalization processing, and obtain standardized image data; ​ The model processing module comprises a plant field pest identification model based on a MaxViT model architecture, wherein the plant field pest identification model integrates a deformable self-attention mechanism and a contrastive learning loss function. The pest identification module is configured to receive plant field pest image data to be identified, perform preprocessing and normalization processing, perform fine-grained classification on the plant field pest image data by using the plant field pest identification model, and output a category result.

9. A plant pest identification apparatus, characterized by, The plant pest identification program comprises: The plant pest identification program is executed by the processor to implement the fine-grained plant field pest identification method based on deep learning according to any one of claims 1-7.

10. A storage medium, characterized by The computer program is stored on the computer readable storage medium and is executed by the processor to implement the fine-grained plant field pest identification method based on deep learning according to any one of claims 1-7.

Citation Information

Patent Citations

  • Unsupervised echocardiogram section identification method

    CN115578589A

  • Crop pest fine-grained identification method, device and equipment and storage medium

    CN117496222A