Crop insect pest detection method and device, detection model training method and device, equipment and medium
By downsampling and upsampling crop images, combined with feature extraction and classification networks, the receptive field is increased, solving the problems of insufficient timeliness and accuracy in existing crop pest detection, and achieving higher accuracy in pest area detection.
Patent Information
- Application Number
- CN202511136092.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-10-31
AI Technical Summary
Current methods for detecting crop pests rely on human judgment or simple image recognition, which suffer from insufficient timeliness and accuracy.
By downsampling and upsampling crop images, combined with feature extraction and classification networks, the receptive field is increased, improving the accuracy of pest detection.
It improves the accuracy and timeliness of crop pest detection, enhances the receptive field for small target pest areas, and increases the detection capability of effective image information.
Smart Images

Figure CN120876891A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image detection technology, and in particular to a method, apparatus, equipment and medium for detecting crop pests and training detection models. Background Technology
[0002] As crops serve as an important basis for credit granting for agricultural enterprises in agricultural and supply chain credit products offered by the Agricultural Bank of China, the disease status of crops should be considered an important factor in credit risk assessment.
[0003] Currently, crop disease status is mainly assessed through human judgment or relatively simple image recognition, and there is room for improvement in the timeliness and accuracy of disease detection. Summary of the Invention
[0004] This invention provides a method, apparatus, equipment, and medium for detecting crop pests and training detection models, in order to improve the accuracy of crop pest detection.
[0005] According to one aspect of the present invention, an embodiment of the present invention provides a method for detecting crop pests, the method comprising:
[0006] The crop image is downsampled to obtain the first extracted features;
[0007] The first extracted feature is subjected to feature extraction to obtain the second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0008] The first extracted feature is upsampled to obtain the third extracted feature;
[0009] The target features are determined based on the third extracted features and the second extracted features;
[0010] The target features are classified to obtain the pest-infested areas in the crop image.
[0011] According to another aspect of the present invention, an embodiment of the present invention provides a detection model training method, the method comprising:
[0012] The sample image is first downsampled using the downsampling network in the detection model to obtain the first extracted features.
[0013] The first extracted feature is extracted using the feature extraction network in the detection model to obtain the second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0014] The first extracted feature is upsampled using the upsampling network in the detection model to obtain the third extracted feature.
[0015] The target features are determined using the upsampling network based on the third extracted features and the second extracted features.
[0016] The target features are classified using the classification network in the detection model to obtain the pest-infested areas in the sample image.
[0017] The detection model is trained based on the differences between the pest-affected areas and the standard target areas in the sample images.
[0018] According to another aspect of the present invention, embodiments of the present invention also provide a crop pest detection device, the device comprising:
[0019] The downsampling network in the detection model is used to perform the first downsampling on the crop image to obtain the first extracted features;
[0020] The feature extraction network in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0021] The upsampling network in the detection model is used to upsample the first extracted features to obtain the third extracted features;
[0022] The upsampling network in the detection model is used to determine the target features based on the third extracted features and the second extracted features;
[0023] The classification network in the detection model is used to classify the target features to obtain the pest-infested areas in the crop image.
[0024] According to another aspect of the present invention, embodiments of the present invention also provide a detection model training apparatus, the apparatus comprising:
[0025] In the detection model, the downsampling network is used to perform the first downsampling on the sample image to obtain the first extracted features;
[0026] The feature extraction network in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0027] The upsampling network in the detection model is used to upsample the first extracted features to obtain the third extracted features;
[0028] The upsampling network is used to determine the target feature based on the third extracted feature and the second extracted feature;
[0029] The classification network in the detection model is used to classify the target features to obtain the pest-infested areas in the sample image;
[0030] The model training module is used to train the detection model based on the differences between the pest-affected area and the standard target area in the sample image.
[0031] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0032] At least one processor; and
[0033] A memory that is communicatively connected to at least one processor; wherein,
[0034] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the crop pest detection method according to any embodiment of the present invention.
[0035] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the crop pest detection method of any embodiment of the present invention.
[0036] The technical solution of this invention involves downsampling a crop image to obtain a first extracted feature, then increasing the receptive field of the first extracted feature to obtain a second extracted feature, and simultaneously upsampling the first extracted feature to obtain a third extracted feature. Based on the second and third extracted features, a target feature is obtained. This method can increase the receptive field of the extracted features, thereby increasing the receptive field of small target pest areas in crops, increasing effective image information in crops, and improving the detection accuracy of pest areas in crop images.
[0037] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 This is a flowchart of a crop pest detection method provided in Embodiment 1 of the present invention;
[0040] Figure 2 This is a flowchart of a detection model training method provided in Embodiment 2 of the present invention;
[0041] Figure 3 This is a scene diagram of a crop pest detection method provided by an embodiment of the present invention;
[0042] Figure 4 This is an application scenario diagram of a detection model provided according to an embodiment of the present invention;
[0043] Figure 5 This is an application scenario diagram of a detection model provided according to an embodiment of the present invention;
[0044] Figure 6 This is a structural diagram of a crop pest detection device provided in Embodiment 3 of the present invention;
[0045] Figure 7 This is a structural diagram of a detection model training device provided in Embodiment 4 of the present invention;
[0046] Figure 8 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation
[0047] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0048] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0049] The acquisition, storage, and application of surveillance videos and other related technologies in the technical solutions of this invention comply with relevant laws and regulations and do not violate public order and good morals.
[0050] Example 1
[0051] Figure 1 This is a flowchart illustrating a method for detecting crop pests according to Embodiment 1 of the present invention. This embodiment of the invention is applicable to crop pest detection. The method can be executed by a crop pest detection device, which can be implemented in hardware and / or software and can be configured in an electronic device.
[0052] See Figure 1 The methods for detecting crop pests shown include:
[0053] S101. Downsample the crop image to obtain the first extracted feature.
[0054] The crop image can include the crop itself. Image acquisition can be performed on crop areas to obtain crop images. The pest-infested area can refer to the image region of the crop affected by pests. The pest-infested area can include at least one of the following: areas with signs of grazing, specific objects (gelatinous crop secretions and insect secretions), discoloration, and deformation. Downsampling is used to extract features from the crop and generate a feature map. Downsampling is specifically the process of sampling a sequence of samples at intervals to obtain a new sequence. In an image, downsampling is a sampling process that reduces the image size and decreases the number of sampling points in the matrix. The first extracted feature can refer to the feature map generated through convolution. Typically, the size of the first extracted feature is smaller than the size of the crop image. In some embodiments, downsampling can be performed n times serially on the crop image to obtain n features of different sizes, which are then used as the first extracted feature. n is an integer greater than or equal to 1. Optionally, n is 5.
[0055] In some embodiments, n first downsampling layers can be used in series, each first downsampling layer outputting a feature map, and the n feature maps are determined as the first extracted features.
[0056] S102. Perform feature extraction on the first extracted feature to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0057] The second extracted feature can refer to a feature map obtained by processing the first extracted feature, and the size of the second extracted feature is the same as that of the first extracted feature. The number of feature maps included in the first extracted feature is the same as the number of feature maps included in the second extracted feature. The second extracted feature can be obtained by further extracting features from the first extracted feature by increasing the receptive field. Methods for increasing the receptive field include: stacked convolution, downsampling, using large-size convolution kernels, dilated convolution, and multi-size fusion. The receptive field can refer to the size of the region in the original image onto which the extracted feature is mapped. The receptive field of the second extracted feature is larger than that of the first extracted feature; it is essentially processing the first extracted feature by increasing its receptive field. In fact, the value of each element in the feature map is calculated from the pixels in the original input image, and the size of the region corresponding to an element in the original input image is the receptive field of that element.
[0058] In some embodiments, the first extracted feature includes n feature maps, and the second extracted feature also includes n feature maps. Among the feature maps of the same size, the receptive field of the feature maps of the same size included in the first extracted feature map is smaller than that of the feature maps of the same size included in the second extracted feature map.
[0059] S103. Upsample the first extracted feature to obtain the third extracted feature.
[0060] Upsampling can be image interpolation, where new elements are inserted between pixels using a suitable interpolation algorithm, thus enlarging the image. Upsampling the first extracted feature can refer to upsampling the smallest feature map among the first extracted features to obtain a single feature map, which becomes the third extracted feature. The smallest feature map is also the last feature map obtained from the first extracted features.
[0061] S104. Determine the target feature based on the third extracted feature and the second extracted feature.
[0062] The target feature can be the feature map extracted from the crop image that best represents the effective image information. The size of the target feature is the same as the size of the crop image. The third extracted feature and the second extracted feature are upsampled and fused to obtain the target feature with the same size as the crop image.
[0063] In some embodiments, the third extracted feature is stitched together with the feature map of the same size in the second extracted feature. The stitched result is upsampled to obtain an intermediate feature map. The intermediate feature map is stitched together with the feature map of the same size in the second extracted feature. The stitched result is upsampled again until a feature map of the same size as the crop image is obtained, and this feature map is used as the target feature.
[0064] S105. Classify the target features to obtain the pest areas in the crop image.
[0065] This process involves classifying target features to obtain the locations of pixels classified as pest-affected areas in the crop image, and then marking these pixels as pest-affected areas within the crop image, thereby enabling the detection of pest-affected areas in the crop image. Classifying the target features outputs the class probability of each pixel in the crop image, and the connected components formed by pixels whose class probability belongs to pest-affected areas are considered as a single pest-affected region.
[0066] The technical solution of this invention involves downsampling a crop image to obtain a first extracted feature, then increasing the receptive field of the first extracted feature to obtain a second extracted feature, and simultaneously upsampling the first extracted feature to obtain a third extracted feature. Based on the second and third extracted features, a target feature is obtained. This method can increase the receptive field of the extracted features, thereby increasing the receptive field of small target pest areas in crops, increasing effective image information in crops, and improving the detection accuracy of pest areas in crop images.
[0067] In an optional embodiment, the step of extracting features from the first extracted features to obtain the second extracted features includes: downsampling, stacking convolution, and upsampling the first extracted features to obtain the second extracted features.
[0068] The first extracted features can be convolved using a feature extraction network to obtain the second extracted features. The feature extraction network comprises n feature extraction units, the number of which is the same as the number of feature maps included in the first extracted features. Each feature extraction unit includes at least an upsampling layer. A feature extraction unit may also include downsampling layers, the number of which is one less than the number of upsampling layers. The feature extraction unit may also include stacked convolutional layers. For example, the feature extraction unit may include U-Net (U-Network). U-Net is a symmetrical U-shaped network. One side performs a series of continuous downsampling operations to extract high-level semantic features, while the other side performs corresponding upsampling operations to restore the image resolution. Simultaneously, the U-Net network applies skip connections at the same level to concatenate the fine-grained surface information from the lower layers with the higher-level semantic information, thereby fusing features from different receptive fields. These fused features are then input into the next layer. The skip connection structure compensates for the information loss caused by downsampling to some extent, ultimately resulting in a segmented image with clear edges.
[0069] In some embodiments, the feature extraction unit includes at least one feature extraction layer, at least one stacked convolutional layer, and at least one upsampling layer connected in series. The feature extraction unit may also include at least one downsampling layer. The feature extraction layer includes at least one convolutionally normalized linear unit (CNN), where a CNN includes a convolutional unit, a batch normalization (BN) unit, and a rectified linear unit (ReLU). For example, the feature extraction layer includes one CNN. The downsampling layer includes a half-downsampling unit and one CNN. The stacked convolutional layer may include at least one CNN, and the stacked convolution may include two CNNs. The upsampling layer includes one CNN and a half-upsampling unit. Downsampling and upsampling layers at the same level or depth are connected to achieve skip connections to concatenate high-resolution features within the same layer of the encoder, forming multi-scale perceptual capabilities and compensating for lost detail information.
[0070] In one example, the number of downsampling layers is one less than the number of upsampling layers, and the number of upsampling layers is the same, corresponding to the size (depth) of the feature map received by the feature extraction unit.
[0071] As can be seen, downsampling achieves exponential expansion of the receptive field, stacked convolutions achieve progressive feature aggregation, and upsampling restores the output features to the same size as the first feature extraction layer, so as to facilitate correct fusion in the subsequent stages.
[0072] In an optional embodiment, the crop pest detection method further includes: extracting edge features from the crop image to obtain a fourth extracted feature; and determining the target feature based on the third extracted feature and the second extracted feature, which includes: determining the target feature based on the fourth extracted feature, the third extracted feature, and the second extracted feature.
[0073] The edge features can refer to the edges of at least one object in the crop image. These edge features can be not limited to the edges of pest-infested areas, but can also include the edges of crops and background objects (such as non-crop plants, people, vehicles, or drones). The fourth extracted feature can be the feature map extracted from the edges. The size of the fourth extracted feature is the same as the size of the first extracted feature. The number of feature maps included in the fourth extracted feature can be the same as the number of feature maps included in the first extracted feature.
[0074] In some embodiments, the crop image can be downsampled n times serially to obtain features of n sizes, which are then used as the fourth extracted feature. n is an integer greater than or equal to 1. Optionally, n is 5. An edge detection network can be used to extract edge features from the crop image. This edge detection network includes a downsampling network and an upsampling network. The edge detection network may also include an edge classification network. The downsampling network in the edge detection network can have the same structure as the downsampling network used to extract the first extracted feature, but with different parameters. Similarly, the upsampling network in the edge detection network can have the same structure as the upsampling network used to extract the third extracted feature, but with different parameters.
[0075] The fourth, third, and second extracted features are upsampled and fused to obtain target features with the same size as the crop image.
[0076] In some embodiments, feature maps of the same size in the third extracted features, the fourth extracted features, and the second extracted features are stitched together. The stitched result is upsampled to obtain an intermediate feature map. The intermediate feature map, the fourth extracted features, and the second extracted features are stitched together. The stitched result is upsampled again until a feature map of the same size as the crop image is obtained, which is then used as the target feature.
[0077] It is evident that by extracting edge features and jointly determining target features, edge information in the target features can be increased. The extracted image edge information can then be used to assist in pest region detection, compensating for the loss of edge information caused by downsampling, enriching the content of the target features, and ultimately achieving satisfactory semantic segmentation results. On the other hand, the features after each downsampling are processed through a simple neural network to enhance the receptive field, enabling the shallow network to enhance semantic representation capabilities without losing detailed information as much as possible, thereby improving the accuracy of semantic segmentation and achieving better semantic segmentation results, thus improving the accuracy of pest region detection.
[0078] In an optional embodiment, determining the target feature based on the fourth extracted feature, the third extracted feature, and the second extracted feature includes: using the third extracted feature as a fusion feature; performing channel stitching on the fusion feature, features of the same size in the second extracted feature, and features of the same size in the fourth extracted feature, upsampling the stitching result, and updating the fusion feature; and determining the stitching result with the same size as the crop image as the target feature.
[0079] Specifically, the feature maps of the same size from the third and fourth extracted features, as well as the feature map of the same size from the second extracted features, are concatenated. The concatenated result is then upsampled to obtain the updated fused feature. This process is repeated until a feature map of the same size as the crop image is obtained, which is then used as the target feature.
[0080] Here, splicing can refer to splicing multiple feature maps with the same size on the channel side. The number of channels in the spliced result is the sum of the number of channels of the spliced feature maps, and the size of the spliced result is the same as the size of each spliced feature map.
[0081] In some embodiments, the crop image is downsampled five times to obtain a first extracted feature, which includes five feature maps. Feature extraction is then performed on the first extracted feature to obtain a second extracted feature, which also includes five feature maps. These feature maps are arranged in descending order of size. Figure 1 ,feature Figure 2 ,feature Figure 3 ,feature Figure 4 and characteristics Figure 5 The crop image was downsampled five times to obtain the fourth extracted feature, which consists of five feature maps arranged in descending order of size. These feature maps are: [feature maps would be inserted here]. Figure 6 ,feature Figure 7 ,feature Figure 8 Feature maps 9 and 10. The smallest feature map obtained from the first extracted features is upsampled to obtain the third extracted features. Among them, feature... Figure 1 and characteristics Figure 6 Same size, features Figure 2 and characteristics Figure 7 Same size, features Figure 3 and characteristics Figure 8 Same size, features Figure 4 The same size as feature image 9, feature Figure 5 The feature map 10 and the third extracted feature have the same size.
[0082] The third extracted feature is used as the fusion feature, along with the feature... Figure 5 Channel concatenation is performed with feature map 10 to obtain the concatenated result. The concatenated result is then upsampled to update the fused features. The fused features are then combined with the feature map 10. Figure 4Channel concatenation is performed with feature map 9 to obtain the concatenated result. The concatenated result is then upsampled to update the fused features. The fused features are then combined with the feature map 9. Figure 3 and characteristics Figure 8 Channel splicing is performed to obtain the spliced result. The spliced result is then upsampled to update the fused features. The fused features are then compared with the features... Figure 2 and characteristics Figure 7 Channel splicing is performed to obtain the spliced result. The spliced result is then upsampled to update the fused features. The fused features are then compared with the features... Figure 1 and characteristics Figure 6 Perform channel splicing to obtain the spliced result.
[0083] The spliced result can be identified as the target feature, or the spliced result can be integrated and convolved to obtain the target feature.
[0084] As can be seen, by performing channel splicing on features of the same size in the third and second extraction features, as well as features of the same size in the fourth extraction features, the splicing result is upsampled and the fused features are updated. This achieves the fusion of feature maps of various sizes to obtain target features. Effective detailed information can be added to the target features, increasing the detail and representativeness of the target features, thereby improving the detection accuracy of pest areas.
[0085] Example 2
[0086] Figure 2 This is a flowchart illustrating a detection model training method provided in Embodiment 2 of the present invention. This embodiment of the invention is applicable to the training of detection models for crop pest detection. The method can be executed by a detection model training device, which can be implemented in hardware and / or software and can be configured within a detection model training equipment.
[0087] See Figure 2 The detection model training method shown includes:
[0088] S201. The sample image is downsampled first through the downsampling network in the detection model to obtain the first extracted feature.
[0089] The downsampling network is used to downsample the sample image multiple times; for example, the downsampling network may include five downsampling units. The sample images can include images of crops. A large number of crop images can be collected as sample images, and the areas infested with pests within the sample images can be labeled.
[0090] S202. Through the feature extraction network in the detection model, the first extracted feature is extracted to obtain the second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0091] The feature extraction network is used to increase the receptive field of the first extracted features.
[0092] S203. The first extracted feature is upsampled through the upsampling network in the detection model to obtain the third extracted feature.
[0093] The upsampling network is used to fuse and upsample the third extracted feature and the second extracted feature to obtain the target feature.
[0094] S204. Using the upsampling network, determine the target features based on the third extracted features and the second extracted features.
[0095] S205. The target features are classified using the classification network in the detection model to obtain the pest area in the sample image.
[0096] The classification network can include a classification function, such as softmax. The classification network can output the class probability of each pixel in the sample image. Pixels with a pest class probability greater than a preset probability threshold are identified as pest pixels. Pixels with a pest class probability less than or equal to the probability threshold are identified as not pest pixels. Pest pixels that form a connected component are identified as a pest region.
[0097] S206. Train the detection model based on the difference between the pest-affected area and the standard target area in the sample image.
[0098] The difference can be calculated by counting the number of differing pixels between the infested area and the standard target area. Alternatively, the difference can be calculated by measuring the intersection-union ratio (IUU) between the infested area and the corresponding standard target area. The loss value is then calculated based on this difference, for example, using the cross-entropy loss function. The detection model can be trained with the goal of minimizing the loss value. Alternatively, a large number of collected sample images can be divided into training and validation sets. The detection model can be trained using the sample images from the training set, and its correctness can be verified using the sample images from the validation set. If the accuracy rate is greater than a preset accuracy threshold, the detection model is considered to have completed training; if the accuracy rate is less than or equal to the accuracy threshold, the detection model is considered to have not completed training, and training continues.
[0099] In this embodiment of the invention, the detection model downsamples the sample image to obtain a first extracted feature, and then performs receptive field increase processing on the first extracted feature to obtain a second extracted feature. At the same time, the first extracted feature is upsampled to obtain a third extracted feature. Based on the second and third extracted features, the target feature is obtained. This can increase the receptive field of the extracted features, thereby increasing the receptive field of small target pest areas in crops, increasing the effective image information in crops, and effectively improving the detection accuracy of the detection model for pest areas.
[0100] In an optional embodiment, the detection model training method further includes: extracting edge features from the crop image using an edge extraction network to obtain a fourth extracted feature; determining a target feature using the upsampling network based on the fourth extracted feature, the third extracted feature, and the second extracted feature; processing the fourth extracted feature using an edge detection network to obtain an edge image; and training the detection model based on the difference between the pest area in the sample image and the standard target area in the sample image includes: training the detection model based on the difference between the pest area in the sample image and the standard target area in the sample image, and the difference between the edge image and the standard image corresponding to the sample image.
[0101] The detection model may further include an edge detection network, which is used to detect the edges of at least one object in the sample image. The sample image is labeled with the edges of the objects to generate a standard image. The standard image can be any of the sample images with correctly labeled edges.
[0102] The difference can be calculated by using the number of pixels differing between the statistical edge image and the standard image. The detection loss value is calculated based on the difference between the infested area and the standard target area, while the edge loss value is calculated based on the difference between the edge image and the standard image; for example, the edge loss value can be calculated using the cross-entropy loss function. The detection loss value and the edge loss value can be summed to obtain the loss value of the detection model. Alternatively, a weighted sum of the detection loss value and the edge loss value can be used to obtain the loss value of the detection model. The detection model can be trained with the goal of minimizing its loss value.
[0103] The parameters of the downsampling network, feature extraction network, upsampling network, edge extraction network, and edge detection network can be adjusted based on the loss value of the detection model. When applying the detection model, the fourth extracted feature does not need to be processed to obtain the edge image. The edge detection network can be removed from the trained detection model, or the trained detection model can be run without the edge detection network. The edge detection network is only used for training and not for applying pest area detection.
[0104] It is evident that by configuring the detection model to include an edge extraction network and an edge detection network, edge features are extracted from sample images through the edge extraction network, and the edge detection network processes the extracted features to obtain an edge image. Based on the difference between the edge image and the standard image corresponding to the sample image, and combined with the difference between the pest area and the standard target area, the detection model is trained together. This can increase the edge detection capability of the detection model and improve the edge detection accuracy. Thus, the extracted, more accurate and richer edge features are integrated into the target features, which can increase the edge information in the target features. The extracted image edge information can then be used to assist in the detection of pest areas, compensate for the loss of edge information caused by downsampling, and improve the accuracy of pest area detection by the detection model.
[0105] In one scenario, when a user initiates a loan request for agricultural products, the bank system needs to review it. During the loan review process or offline approval, the Agricultural Bank of China's remote sensing satellite technology for agricultural products is used to acquire crop images. The acquired crop images are then used to detect diseases, and users with a large number of crop diseases are marked and given early warnings, thereby assisting the Agricultural Bank of China's credit personnel in controlling credit risks.
[0106] The overall process can be as follows Figure 3 As shown:
[0107] S301. Obtain the crop information involved in the user request.
[0108] The user request may contain information such as the type and location of the crops. Optionally, the user request could be a financing or credit request from a banking system. Alternatively, the user request could be a request to detect and identify crop images, such as a request initiated by a user who grows crops.
[0109] S302. Based on the user's request, instruct the remote sensing satellite to photograph crops within the planting area and obtain at least one crop image.
[0110] The system can determine the planting area to be photographed based on the crop information requested by the user. Specifically, this includes the extent and location of the planting area.
[0111] S303. Identify pest-infested areas in images of various crops.
[0112] The detection model of the crop pest detection method implemented in the embodiments of the present invention is used to identify the pest areas in each crop image.
[0113] The training process of the detection model may include: acquiring n crop images for model training using satellite remote sensing technology, marking diseased areas in crop leaf images to obtain a sample set of labeled crop leaf images. The crop image sample set is divided into a training set and a validation set. The training set is imported into the detection model for training, and the validation set is used to test the trained detection model. The recognition results of each sample image are obtained, and the recognition accuracy of each sample image is determined based on the recognition results. When the accuracy is greater than a preset accuracy threshold, the detection model training is considered complete; when the accuracy is less than or equal to the accuracy threshold, the detection model training is considered incomplete, and training continues. Once training is complete, the current detection model is used as a detection model for assessing crop disease conditions.
[0114] Existing detection models can only fuse multi-scale features from lower layers through cross-layer connections during downsampling, and the number of fusions depends on the number of downsampling iterations. Because the receptive field of the network's features is too small during the initial downsampling, its ability to identify targets is weak. Furthermore, to reduce computational overhead, the network typically requires 2-3 downsampling iterations before cross-layer connections are made. Additionally, the detection models acquire insufficient multi-scale features, resulting in poor segmentation accuracy for small targets.
[0115] In this embodiment of the invention, the detection model can be an Image Edge Supervised Same Downsample Frequency Semantic Segmentation Network (IEDF4SNet). The specific process of the detection model is as follows: Figure 4 and Figure 5 As shown:
[0116] The crop image is processed by a residual convolutional network to extract target features. This residual convolutional network can be a 50-layer network (ResNet50). It includes a downsampling network, an upsampling network, and a classification network. In the detection model, a feature extraction network is added after the downsampling network within the residual convolutional network. This ensures that the low-level information after each downsampling passes through a feature extraction network (such as a miniature U-Net) to increase the receptive field and obtain multi-scale features. The downsampling network in the residual convolutional network includes 5 downsampling units, and the upsampling network includes 5 upsampling units. The crop image undergoes five downsampling passes, and the smallest feature map obtained is input into the upsampling network. After five upsampling passes, the feature maps from the five feature extraction units are fused together, ultimately yielding the target features. These target features are then processed by a classification network to identify the pest region. Figure 5As shown, the crop image includes crop leaves. Specifically, the first downsampling unit outputs a feature map of a first size; the second downsampling unit outputs a feature map of a second size; the third downsampling unit outputs a feature map of a third size; the fourth downsampling unit outputs a feature map of a fourth size; and the fifth downsampling unit outputs a feature map of a fifth size. The first size is larger than the second size, the second size is larger than the third size, the third size is larger than the fourth size, and the fourth size is larger than the fifth size.
[0117] Meanwhile, the crop image extracts the edge contour information of the crop through an edge extraction network. The edge extraction network may include a downsampling network with 5 downsampling layers and an upsampling network with 5 upsampling layers. For example, the edge extraction network may include a simple U-Net network with five downsampling and five upsampling layers.
[0118] Finally, during each upsampling, the residual convolutional network concatenates the processed low-level information, the crop edge contour information captured by the edge extraction network, and the high-level semantic information extracted from each downsampling layer in the residual convolutional network through cross-layer connections. This concatenated information is then fused through a convolutional layer to ultimately output the semantic segmentation result, i.e., the pest-affected area. Figure 5 As shown, the pest-infested areas output by the detection model are the red areas in the crop images marked with pests.
[0119] like Figure 5 As shown, the feature map of size 5 output by the fifth downsampling unit is processed by the first upsampling unit to output a feature map of size 4. This feature map is then concatenated with the feature map output by the feature extraction unit in the same row as the fifth downsampling unit, and the feature map output by the first upsampling in the edge extraction network, to obtain the concatenated result. The size of the concatenated result is 4. The concatenated result is then processed by the second upsampling unit to output a feature map of size 3. The number of upsampling units included in the feature extraction unit in the same row as the fifth downsampling unit is the number of downsampling units plus 1 (i.e., one more upsampling step, making the size of the feature map of size 5 become size 4). This ensures that the feature map output by the feature extraction unit and the feature map output by the first upsampling unit are both 4.
[0120] Similarly, the stitched result, after being processed by the second upsampling unit, outputs a feature map of the third size. This feature map is then combined with the feature map output by the feature extraction unit in the same row as the fourth downsampling unit, and the feature map output by the second upsampling in the edge extraction network. The resulting stitched result has the third size. After being processed by the third upsampling unit, it outputs a feature map of the second size.
[0121] The feature map of the second size output by the third upsampling unit, the feature map output by the feature extraction unit in the same row as the third downsampling unit, and the feature map output by the third upsampling in the edge extraction network are concatenated to obtain the concatenated result. The size of the concatenated result is the second size. The concatenated result is then processed by the fourth upsampling unit to output a feature map of the first size.
[0122] The feature map of the first size output by the fourth upsampling unit, the feature map output by the feature extraction unit in the same row as the second downsampling unit, and the feature map output by the fourth upsampling in the edge extraction network are concatenated to obtain the concatenated result. The size of the concatenated result is the first size. The concatenated result is then processed by the fifth upsampling unit to output a feature map of the original image size.
[0123] The feature map of the original image size output from the fifth upsampling unit, the feature map output from the feature extraction unit in the same row as the first downsampling unit, and the feature map output from the fifth upsampling in the edge extraction network are concatenated to obtain the concatenated result. The concatenated result is then processed by convolutional fusion and a classification function in the classification network to obtain the pest-affected area. For example... Figure 5 As shown, the red areas in the crop image represent pest-infested areas, corresponding to the brown areas in the original crop image, thus achieving accurate identification of pest-infested areas.
[0124] After the above processing, the detection model can compensate for the loss of edge information caused by continuous downsampling, and can also improve the receptive field of shallow features and obtain multi-scale features, thereby improving the accuracy of crop identification.
[0125] During training, to ensure the accuracy of the features extracted by the edge extraction network, an edge detection network is added after the edge extraction network to generate an edge image based on the extracted fourth feature, such as... Figure 5 As shown, the edges of the leaves are marked in the edge image. The parameters of the edge extraction network are adjusted based on the edge image to improve the accuracy of the extracted edge features.
[0126] S304. Based on the recognition results of each crop image, determine the risk probability of the user's request.
[0127] It can statistically analyze the ratio of pest-infested areas in crop images and / or count the number of crops in pest-infested areas, calculating the ratio of these pest-infested areas to the total number of crops. Based on the ratio of pest areas and / or the ratio of the number of crops, it can calculate the probability of risk for a user's request. It can also calculate the risk probability by weighting the ratio of pest areas and the ratio of the number of crops. A pre-set warning probability threshold can be set, and a warning can be issued to the user who initiated the request if the risk probability is greater than or equal to the threshold.
[0128] Example 3
[0129] Figure 6 This is a schematic diagram of a crop pest detection device provided in Embodiment 3 of the present invention. This embodiment of the invention is applicable to crop pest detection, and the device can perform crop pest detection methods. The device can be implemented in hardware and / or software, and can be configured in an electronic device.
[0130] See Figure 6 The crop pest detection device shown includes:
[0131] In the detection model, the downsampling network 601 is used to perform the first downsampling on the crop image to obtain the first extracted features;
[0132] The feature extraction network 602 in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0133] The upsampling network 603 in the detection model is used to upsample the first extracted feature to obtain the third extracted feature;
[0134] The upsampling network 603 in the detection model is used to determine the target features based on the third extracted features and the second extracted features;
[0135] The classification network 604 in the detection model is used to classify the target features to obtain the pest areas in the crop image.
[0136] The technical solution of this invention involves downsampling a crop image to obtain a first extracted feature, then increasing the receptive field of the first extracted feature to obtain a second extracted feature, and simultaneously upsampling the first extracted feature to obtain a third extracted feature. Based on the second and third extracted features, a target feature is obtained. This method can increase the receptive field of the extracted features, thereby increasing the receptive field of small target pest areas in crops, increasing effective image information in crops, and improving the detection accuracy of pest areas in crop images.
[0137] Optional, feature extraction network 602, specifically used for:
[0138] The first extracted feature is subjected to convolution processing to obtain the second extracted feature; the convolution processing includes at least one of the following: downsampling, stacked convolution and upsampling.
[0139] Optionally, the crop pest detection device also includes:
[0140] An edge extraction network is used to extract edge features from the crop image to obtain a fourth extracted feature;
[0141] The upsampling network 603 is specifically used to determine the target feature based on the fourth extracted feature, the third extracted feature, and the second extracted feature.
[0142] Optional, the upsampling network 603 is specifically used for:
[0143] The third extracted feature is used as the fusion feature;
[0144] The fused feature, features of the same size in the second extracted feature, and features of the same size in the fourth extracted feature are concatenated by channels, and the concatenation result is upsampled to update the fused feature;
[0145] The stitched result with the same size as the crop image is identified as the target feature.
[0146] The crop pest detection device provided in the embodiments of the present invention can execute the crop pest detection method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the crop pest detection method.
[0147] Example 4
[0148] Figure 7 This is a schematic diagram of a detection model training device provided in Embodiment 4 of the present invention. This embodiment of the present invention is applicable to the training of detection models. This device can execute a detection model training method and can be implemented in hardware and / or software. This device can be configured in an electronic device.
[0149] See Figure 7 The detection model training device shown includes:
[0150] In the detection model, the downsampling network 701 is used to perform the first downsampling on the sample image to obtain the first extracted features;
[0151] The feature extraction network 702 in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature.
[0152] The upsampling network 703 in the detection model is used to upsample the first extracted features to obtain the third extracted features;
[0153] The upsampling network 703 is used to determine the target feature based on the third extracted feature and the second extracted feature;
[0154] The classification network 704 in the detection model is used to classify the target features to obtain the pest area in the sample image;
[0155] The model training module 705 is used to train the detection model based on the difference between the pest area and the standard target area in the sample image.
[0156] In this embodiment of the invention, the detection model downsamples the sample image to obtain a first extracted feature, and then performs receptive field increase processing on the first extracted feature to obtain a second extracted feature. At the same time, the first extracted feature is upsampled to obtain a third extracted feature. Based on the second and third extracted features, the target feature is obtained. This can increase the receptive field of the extracted features, thereby increasing the receptive field of small target pest areas in crops, increasing the effective image information in crops, and effectively improving the detection accuracy of the detection model for pest areas.
[0157] Optionally, the detection model training device also includes:
[0158] An edge extraction network is used to extract edge features from the crop image to obtain a fourth extracted feature;
[0159] The upsampling network 703 is specifically used to determine the target feature based on the fourth extracted feature, the third extracted feature, and the second extracted feature;
[0160] An edge detection network is used to process the fourth extracted features to obtain an edge image;
[0161] Model training module 705 is specifically used for:
[0162] The detection model is trained based on the differences between the pest-affected area and the standard target area in the sample image, as well as the differences between the edge image and the standard image corresponding to the sample image.
[0163] The detection model training device provided in this embodiment of the invention can execute the detection model training method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects for executing the detection model training method.
[0164] Example 5
[0165] Figure 8 A schematic diagram of the structure of an electronic device 800 that can be used to implement an embodiment of the present invention is shown.
[0166] like Figure 8As shown, the electronic device 800 includes at least one processor 801 and a memory, such as a read-only memory (ROM) 802 and a random access memory (RAM) 803, communicatively connected to the at least one processor 801. The memory stores computer programs executable by the at least one processor. The processor 801 can perform various appropriate actions and processes based on the computer program stored in the ROM 802 or loaded into the RAM 803 from storage unit 808. The RAM 803 can also store various programs and data required for the operation of the electronic device 800. The processor 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0167] Multiple components in electronic device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of displays, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows electronic device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0168] Processor 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 801 performs the various methods and processes described above, such as crop pest detection methods.
[0169] In some embodiments, the crop pest detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded into and / or installed on electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by processor 801, one or more steps of the crop pest detection method described above may be performed. Alternatively, in other embodiments, processor 801 may be configured to perform the crop pest detection method by any other suitable means (e.g., by means of firmware).
[0170] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0171] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0172] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0174] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0175] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0176] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0177] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for detecting crop pests, characterized in that, The method includes: The crop image is downsampled to obtain the first extracted feature; The first extracted feature is subjected to feature extraction to obtain the second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature. The first extracted feature is upsampled to obtain the third extracted feature; The target features are determined based on the third extracted features and the second extracted features; The target features are classified to obtain the pest-infested areas in the crop image.
2. The method according to claim 1, characterized in that, The step of extracting features from the first extracted features to obtain the second extracted features includes: The first extracted feature is subjected to convolution processing to obtain the second extracted feature; the convolution processing includes at least one of the following: downsampling, stacked convolution and upsampling.
3. The method according to claim 1, characterized in that, Also includes: Edge features are extracted from the crop image to obtain the fourth extracted feature; The step of determining the target feature based on the third extracted feature and the second extracted feature includes: The target feature is determined based on the fourth extracted feature, the third extracted feature, and the second extracted feature.
4. The method according to claim 3, characterized in that, The step of determining the target feature based on the fourth extracted feature, the third extracted feature, and the second extracted feature includes: The third extracted feature is used as the fusion feature; The fused feature, features of the same size in the second extracted feature, and features of the same size in the fourth extracted feature are concatenated by channels, and the concatenation result is upsampled to update the fused feature; The stitched result with the same size as the crop image is identified as the target feature.
5. A method for training a detection model, characterized in that, The method includes: The sample image is first downsampled using the downsampling network in the detection model to obtain the first extracted features. The first extracted feature is extracted using the feature extraction network in the detection model to obtain the second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature. The first extracted feature is upsampled using the upsampling network in the detection model to obtain the third extracted feature. The target features are determined using the upsampling network based on the third extracted features and the second extracted features. The target features are classified using the classification network in the detection model to obtain the pest-infested areas in the sample image. The detection model is trained based on the differences between the pest-affected areas and the standard target areas in the sample images.
6. The method according to claim 5, characterized in that, Also includes: Edge features are extracted from the crop image using an edge extraction network to obtain the fourth extracted feature; The target feature is determined using the upsampling network based on the fourth extracted feature, the third extracted feature, and the second extracted feature. The edge image is obtained by processing the fourth extracted features through an edge detection network. The step of training the detection model based on the difference between the pest-infested area in the sample image and the standard target area in the sample image includes: The detection model is trained based on the differences between the pest-affected area and the standard target area in the sample image, as well as the differences between the edge image and the standard image corresponding to the sample image.
7. A crop pest detection device, characterized in that, The device includes: The downsampling network in the detection model is used to perform the first downsampling on the crop image to obtain the first extracted features; The feature extraction network in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature. The upsampling network in the detection model is used to upsample the first extracted features to obtain the third extracted features; The upsampling network in the detection model is used to determine the target features based on the third extracted features and the second extracted features; The classification network in the detection model is used to classify the target features to obtain the pest-infested areas in the crop image.
8. A detection model training device, characterized in that, The device includes: In the detection model, the downsampling network is used to perform the first downsampling on the sample image to obtain the first extracted features; The feature extraction network in the detection model is used to extract features from the first extracted features to obtain a second extracted feature; the receptive field of the second extracted feature is larger than that of the first extracted feature. The upsampling network in the detection model is used to upsample the first extracted features to obtain the third extracted features; The upsampling network is used to determine the target feature based on the third extracted feature and the second extracted feature; The classification network in the detection model is used to classify the target features to obtain the pest-infested areas in the sample image; The model training module is used to train the detection model based on the differences between the pest-affected area and the standard target area in the sample image.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the crop pest detection method according to any one of claims 1-4, or the detection model training method according to any one of claims 5-6.
10. A computer-readable medium, characterized in that, The computer-readable medium stores computer instructions that cause a processor to execute the crop pest detection method of any one of claims 1-4, or the detection model training method of any one of claims 5-6.
Citation Information
Cited By
Crop disease and pest identification and classification method and system, storage medium and equipment
CN121937886A