Fruit tree pest recognition method and system based on convolutional neural network

CN122416219BActive Publication Date: 2026-09-11YANAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202610889636.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-18
Publication Date
2026-09-11
Estimated Expiration
2046-06-18

AI Technical Summary

Technical Problem

[0003]然而,上述方法主要面向自然光照或实验室可控环境下的病虫害图像,未能充分考虑设施大棚特有的高湿、弱光及光学干扰问题

Benefits of technology

[0016] Compared with the prior art, the beneficial effects of the present invention are: the present invention constructs and quantifies... Lens condensation interference index and The leaf water film interference index is used, and the interference level is uniformly divided into 1-5 levels, which can achieve objective classification of the interference level in greenhouse images. This invention is based on... Dynamically adjust the dark channel window and transmittance correction factor, based on By dynamically adjusting the reflection threshold and adapting the low-light enhancement amplitude based on the image grayscale mean, it can specifically eliminate three typical interferences: lens condensation and fogging, local reflection of water film on leaves, and low contrast in low light. It directly restores the edge and texture details of tiny pests, significantly improving the usability of image features and providing a high-quality image foundation for subsequent feature extraction and pest detection, avoiding feature failure and false positives or false negatives due to interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122416219B_ABST
    Figure CN122416219B_ABST
Patent Text Reader

Abstract

The present application provides a kind of fruit tree pest identification method and system based on convolutional neural network, it is related to agricultural intelligent detection technical field, the present application is aimed at the image interference problem caused by condensation, water film and illumination change in facility greenhouse, first, construct lens condensation interference index and leaf water film interference index, divide interference grade and collect pest image dataset under multiple interference conditions;Subsequently, global defogging, local reflection repair and weak light contrast adaptive enhancement are carried out in turn, to obtain high-quality restored image;Then, multi-scale features are extracted using a lightweight backbone network, and the features are enhanced using a cross-layer dense fusion and spatial-channel collaborative attention mechanism;Finally, a specialized detection head and a multi-task joint loss function are used to train the model, and the detection results are post-processed and evaluated for reliability. This method effectively improves the recognition accuracy and robustness of small pests in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural intelligent detection technology, specifically to a method and system for identifying fruit tree pests based on convolutional neural networks. Background Technology

[0002] In facility agriculture environments, intelligent identification of fruit tree pests is crucial for ensuring fruit yield and quality. In recent years, with the development of deep learning technology, pest identification methods based on convolutional neural networks (CNNs) have gradually become a research hotspot. Existing technologies typically achieve automatic pest identification by constructing a convolutional neural network model and combining image preprocessing, feature extraction, and classification detection steps. For example, the existing technology with publication number CN109344883A proposes a fruit tree pest identification method based on dilated convolution. It extracts image features through multi-scale convolution kernels and uses a Softmax classifier to determine the disease category, thus improving the identification efficiency to some extent in complex backgrounds. The existing technology with publication number CN113468984A uses a deep residual network combined with embedded image preprocessing technology to extract color and texture features for pest leaf identification, improving the model's accuracy under standard lighting conditions. Furthermore, the existing technology with publication number CN110188635A introduces an attention mechanism and a multi-level convolutional feature fusion strategy to suppress interference from complex backgrounds in natural scenes and enhance the feature representation ability of the target region. The existing technology with publication number CN109409170A further proposes a pest identification method based on confidence adjustment, which improves the reliability of the identification results by calculating initial probability values ​​and making dynamic corrections. The existing technology with publication number CN114445785A combines the Internet of Things and generative adversarial networks to construct a litchi pest identification and hierarchical early warning system, realizing closed-loop management from identification to prevention and control.

[0003] However, the aforementioned methods primarily target pest images under natural light or controlled laboratory environments, failing to adequately consider the unique challenges of high humidity, low light, and optical interference found in greenhouses. Specifically, in high humidity environments, condensation easily forms on lens surfaces, resulting in global fogging and blurring of the image; simultaneously, water films adhering to leaf surfaces cause localized strong reflections, severely obscuring details of tiny insects; coupled with significantly reduced image contrast under low light conditions, all three factors contribute to severe degradation of the input image quality, rendering subsequent feature extraction ineffective. Existing technologies generally lack quantitative modeling and adaptive restoration mechanisms for these interfering factors. Image preprocessing often employs general denoising or enhancement algorithms, failing to specifically restore pest-affected areas affected by condensation, water films, and low illumination. Furthermore, target pests such as spider mites (0.2–0.5 mm) and aphids often account for less than 1% of the image, representing a typical challenge in small-target detection. While some methods attempt to introduce multi-scale feature fusion, their fusion strategies often neglect the preservation of shallow, high-resolution features, resulting in the loss of texture and edge information of minute pests in deep networks and persistently high false negative rates. On the other hand, lightweight models (such as the MobileNet series) have been attempted to meet edge deployment requirements, but their representational capabilities are limited by the reduced parameter compression, resulting in insufficient sensitivity to small targets and an accuracy rate that falls short of the 85% threshold required for agricultural production. More critically, existing systems generally only output category and location coordinates, lacking a mechanism to dynamically assess the reliability of detection results by incorporating real-time environmental interference intensity (such as dew level and water film coverage). This makes it impossible to quantify the risk of misjudgment or false negatives, hindering automated decision-making processes such as precision pesticide application. Therefore, a comprehensive technical solution is urgently needed that can synergistically address high-humidity, low-light image restoration, minute pest feature enhancement, lightweight model optimization, and result reliability assessment.

[0004] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for identifying fruit tree pests based on convolutional neural networks, so as to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: A method for identifying fruit tree pests based on convolutional neural networks, comprising the following steps: S1: Construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level. Collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and label target pest individuals. Divide the dataset into training set, validation set and test set. S2: Perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image. S3: Based on the convolutional neural network backbone network, multi-level feature extraction is performed on the obtained high-quality restored image to output multi-scale feature maps with different downsampling factors; cross-layer dense fusion is performed on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; feature processing is performed on the obtained high-resolution fused feature maps to output enhanced high-resolution feature maps. S4: Construct a detection head specifically for small pests based on the enhanced high-resolution feature map. Based on the annotation of individual target pests, generate anchor boxes adapted to the size of the target pests through clustering. Output the pest classification and localization results of the corresponding anchor boxes through the detection head. Construct a multi-task joint loss function. Train and optimize the convolutional neural network model based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. S5: After processing the fruit tree images in the greenhouse to be detected through steps S2 and S3, input them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. Perform post-processing steps on the results and output the final identification results.

[0007] Furthermore, the lens condensation interference index The leaf water film interference index is constructed by combining the image grayscale mean and contrast. Constructed by combining average image reflectance and edge sharpness; based on the above and The threshold range for the levels is set, and the levels of condensation interference and water film interference are divided into 5 levels from 1 to 5, and the level values ​​are directly used in subsequent calculations; at the same time, the light intensity is divided into three categories: weak light range, normal light range, and strong light range.

[0008] Furthermore, the lens condensation interference index The calculation formula is: in, The average gray level of the image; Image contrast; To preset the maximum contrast value, Image grayscale weights, Image contrast weights, satisfying , ; The value range is 0 to 1; The blade water film interference index The calculation formula is: in, The average reflectance of the image. This represents the maximum reflectivity. For image edge sharpness, This represents the maximum edge sharpness. Weights for average image reflectance. As the image edge sharpness weight, satisfying , , The value range is 0 to 1.

[0009] Furthermore, the global fog removal process employs an improved dark channel prior algorithm, incorporating the lens condensation interference index. Dynamically adjusting algorithm parameters specifically includes: calculating the dark channel image of the input fruit tree pest image, and the size of the dark channel calculation window. Depend on Dynamic adjustment, the formula is: in, This is a rounding function; the top 0.1% of bright pixels are selected from the dark channel image, and the maximum grayscale value of these pixels in the original image is taken as the atmospheric light value; combined with... The transmittance is calculated by dynamically adjusting the correction factor, using the following formula: in, coordinates Transmittance at that location For the dark channel image in coordinates The dark channel value at that location, To dynamically adjust the correction coefficient, The atmospheric light value is used; a dehazed image is generated based on the dark channel prior model, using the following formula: in, To dehaze image pixels in coordinates grayscale value at that location Fruit tree pest images in coordinates grayscale value at that location This is the minimum transmittance threshold.

[0010] Furthermore, the detection and repair of the local reflective area specifically includes: based on the leaf water film interference index... The reflectivity threshold can be dynamically set using the following formula: in, The reflection threshold is used for dehazed images. Perform grayscale conversion to obtain a grayscale image. Calculate the difference between the gray value of each pixel and the mean of its surrounding 3×3 neighborhood. ,when When this happens, the pixel is determined to be a reflective pixel, and all reflective pixels constitute a reflective area. The texture filling method was used to repair the reflective areas, and the reflective areas were selected. The surrounding 5×5 non-reflective area is used as a texture sample library. The normalized cross-correlation coefficient between each texture block in the sample library and the block to be repaired is calculated. The texture block with the largest coefficient is selected to replace the block to be repaired until all reflective areas are repaired, and the de-reflective image is obtained.

[0011] Furthermore, the adaptive enhancement of low-light contrast specifically includes: calculating the grayscale mean of the dereflected image as the illumination intensity coefficient. The formula is: in, Image width, Image height; based on illumination intensity coefficient Dynamically adjust the enhancement level to generate an enhanced image. The formula is: in, To enhance the coefficients, the pixel values ​​of the enhanced image are normalized to 0~255 to obtain a high-quality restored image.

[0012] Further, in step S3, the convolutional neural network backbone network is MobileNetV3-small, and the network configuration is as follows: the input layer receives a 224×224×3 RGB image and normalizes it to [-1,1]; the feature extraction layer contains 16 convolutional blocks, of which the first 13 are depth-separable convolutional blocks and the last 3 are standard convolutional blocks, and a BatchNorm2d normalization layer and a Swish activation function are added after each convolutional block; the output layer outputs four multi-scale feature maps with different downsampling factors, namely C1, C2, C3, and C4; the cross-layer dense fusion specifically includes: upsampling the deep feature maps C4, C3, and C2 to a size of 56×56 using bilinear interpolation; and concatenating the upsampled C4, C3, and C2 with the shallow feature map C1 through channels to obtain a concatenated feature map. The number of channels is 176; 1×1 convolution is used to... The number of channels was adjusted to 64, and then an SE channel attention mechanism was introduced to weight the channel features, resulting in a high-resolution fused feature map. The aforementioned feature processing of the obtained high-resolution fused feature map specifically involves feature processing to suppress interference and enhance the target, employing a spatial-channel dual-branch collaborative attention mechanism, specifically including: ... Global average pooling is performed, followed by two fully connected layers and an activation function to obtain the channel attention weights. ;right Perform 1×1 convolution to reduce dimensionality to 1 channel, and then obtain spatial attention weights using the Sigmoid activation function. ; Using a broadcast mechanism to and Multiplication yields the collaborative attention weights. ,Will and Pixel-by-pixel multiplication yields an enhanced high-resolution feature map.

[0013] Furthermore, the formula for the multi-task joint loss function mentioned in step S4 is: in, The total value of the joint loss across multiple tasks. For classifying losses, To pinpoint the loss, Weighted loss for small goals These are the weighting coefficients for each loss. In step S4, the hyperparameters for model training are set as follows: batch size 32, initial learning rate 0.001, cosine annealing learning rate adjustment strategy, weight decay coefficient 0.0001, 200 iterations, early stopping strategy and patience=20; during training, a gradient clipping strategy is introduced to limit the gradient norm to within 5.0, and a Dropout layer is added between the feature extraction layer and the detection head with dropout rate=0.3.

[0014] Further, the post-processing step described in S5 specifically includes: setting a classification confidence threshold and removing anchor boxes with a classification confidence less than the threshold; using a non-maximum suppression algorithm to remove overlapping anchor boxes and setting an NMS threshold; converting the anchor box offsets output by the model into image pixel coordinates; and outputting the final recognition result, which also includes a confidence evaluation step: calculating the confidence level based on the interference level and classification confidence, using the following formula: in, For credibility, For classification confidence, For condensation interference level, The level of water film interference is categorized; the credibility is divided into three levels: high credibility. ≥90%, Medium confidence level ≤70% <90%, low credibility <70%, and output the corresponding confidence level.

[0015] The present invention also provides a fruit tree pest identification system based on a convolutional neural network, wherein the fruit tree pest identification system based on a convolutional neural network is used to execute the above-described fruit tree pest identification method based on a convolutional neural network, comprising: Multi-interference dataset construction module: used to construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level, collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and perform target pest individual labeling, and divide the dataset into training set, validation set and test set. Image restoration module: used to perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image; Feature extraction and enhancement module: Based on the convolutional neural network backbone network, it performs multi-level feature extraction on the obtained high-quality restored image, and outputs multi-scale feature maps with different downsampling factors; it performs cross-layer dense fusion on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; it performs feature processing on the obtained high-resolution fused feature maps and outputs enhanced high-resolution feature maps. Model training and optimization module: This module is used to construct a detection head specifically for small pests based on the enhanced high-resolution feature maps. Based on the annotation of individual target pests, it generates anchor boxes adapted to the size of the target pests through clustering. The detection head outputs the pest classification and localization results corresponding to the anchor boxes. A multi-task joint loss function is constructed, and the convolutional neural network model is trained and optimized based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. Pest and disease identification output module: After processing the images of fruit trees in the greenhouse to be detected in steps S2 and S3, the module inputs them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. The module then performs post-processing steps to output the final identification result.

[0016] Compared with the prior art, the beneficial effects of the present invention are: the present invention constructs and quantifies... Lens condensation interference index and The leaf water film interference index is used, and the interference level is uniformly divided into 1-5 levels, which can achieve objective classification of the interference level in greenhouse images. This invention is based on... Dynamically adjust the dark channel window and transmittance correction factor, based on By dynamically adjusting the reflection threshold and adapting the low-light enhancement amplitude based on the image grayscale mean, it can specifically eliminate three typical interferences: lens condensation and fogging, local reflection of water film on leaves, and low contrast in low light. It directly restores the edge and texture details of tiny pests, significantly improving the usability of image features and providing a high-quality image foundation for subsequent feature extraction and pest detection, avoiding feature failure and false positives or false negatives due to interference.

[0017] Compared with existing general detection heads and single loss functions, this invention generates anchor boxes adapted to the size of target pests through clustering and constructs a multi-task joint loss function that includes classification loss, localization loss, and small target weighted loss. This enables the model training process to be optimized in the direction of small pest detection, directly improving the localization accuracy and classification accuracy of small pests, and solving the technical problems of insufficient weight for small targets, inaccurate regression, and easy obscuring of small targets by traditional models.

[0018] This invention combines classification confidence level and interference level to calculate and output confidence level, which can provide a quantifiable reliability evaluation of the detection results. This allows end users to take corresponding prevention and control measures based on the confidence level, avoid misjudgment of low confidence results caused by interference, and improve the security and availability of the system in actual production applications.

[0019] This invention establishes a complete collaborative link by integrating interference quantization index, adaptive image restoration, lightweight multi-scale feature enhancement, small target-specific loss, and credibility-level post-processing. This maintains high recognition accuracy even under strong interference levels, while achieving a balance between lightweight design and high accuracy. It overcomes the long-standing technical contradictions in the field of "difficulty in achieving both anti-interference and model lightweighting" and "difficulty in distinguishing between small targets and complex backgrounds." Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the overall method flow of the present invention; Figure 2 This is a graph showing the pest identification results of an embodiment of the present invention; Figure 3 This is a block diagram of the system module structure of the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0022] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0023] Example: Please see Figures 1-2 The present invention provides a technical solution: A method for identifying fruit tree pests based on convolutional neural networks, comprising the following steps: S1: Construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level. Collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and label target pest individuals. Divide the dataset into training set, validation set and test set. In this embodiment, the lens condensation interference index The leaf water film interference index is constructed by combining the image grayscale mean and contrast. Constructed by combining average image reflectance and edge sharpness; based on the above and A threshold range was set, dividing both condensation interference and water film interference into five levels, from 1 to 5. The interference levels of lens condensation and leaf water film represent a continuous change from "no interference" to "completely unrecognizable." This five-level division can fully cover all possible interference scenarios without excessively increasing data annotation and algorithm complexity. If there are fewer than five levels, such as three, the difference in interference levels within each level will be too large, making it impossible to achieve accurate adaptive adjustment of the subsequent image restoration algorithm. If there are more than five levels, such as seven, the number of samples for each level will be too small, resulting in insufficient model training and a significant decrease in the marginal benefit of algorithm parameter adjustment, with the level values ​​directly participating in subsequent calculations. Simultaneously, light intensity was divided into three categories: low light range, normal light range, and high light range.

[0024] Furthermore, the lens condensation interference index The calculation formula is: in, The average gray level of the image; Image contrast; To preset the maximum contrast value, Image grayscale weights, Image contrast weights, satisfying , , The value ranges from 0 to 1, and the output range is also from 0 to 1. A larger value indicates more severe condensation interference. In this embodiment... The impact of condensation on images is mainly reflected in two aspects: increased overall brightness and blurred details, corresponding to an increase in the average grayscale value and a decrease in contrast; the weight allocation in this embodiment... Based on engineering statistics: condensation has a greater impact on image brightness than on contrast, when At that time, the image appears white overall, and the contrast is almost completely lost.

[0025] The blade water film interference index The calculation formula is: in, The average reflectance of the image. This represents the maximum reflectivity. For image edge sharpness, This represents the maximum edge sharpness. Weights for average image reflectance. As the image edge sharpness weight, satisfying , In this embodiment , The value range is 0~1. The effect of the water film on the image is mainly reflected in: local reflection and edge blurring, corresponding to increased reflectivity and decreased edge sharpness. In this embodiment... This is because the reflectivity and edge blurring effect of the water film are equally important, as both jointly determine the severity of water film interference; and both are normalized to ensure that... and Having the same numerical range facilitates unified processing later.

[0026] The dataset was constructed using a combination of real-world greenhouse data collection and laboratory semi-physical simulation. Images were taken during the high humidity period in the early morning, after irrigation, and under different natural light conditions at various times. For extreme interference conditions, supplementary data collection was conducted in a temperature- and humidity-controlled laboratory chamber by adjusting condensation levels and quantitatively spraying deionized water. At least 2000 images were collected, and the dataset was divided into training, validation, and test sets in a 7:1:2 ratio. The dataset included bounding boxes for pests such as aphids and spider mites, with cross-over coverage of different interference levels. Professional open-source annotation tools were used for annotation, and the results were cross-validated by two researchers.

[0027] In this embodiment, the dataset is constructed collaboratively in two ways: data collection from actual facility greenhouses and supplementation by laboratory semi-physical simulation. To ensure adequate coverage of interference conditions during actual greenhouse data collection, the following measures were taken to address lens condensation interference: photos were taken in the early morning when the temperature difference between inside and outside the greenhouse was large and the humidity inside was high; to address leaf water film interference: photos were taken within 30 minutes after irrigation and during periods when the leaves were not completely dry after rain; some samples were simulated by uniformly spraying deionized water onto the leaf surface using a micro-sprayer; and to address light variations: photos were taken naturally during three typical time periods: noon on a sunny day, dusk on a cloudy day, and morning on a sunny day. Laboratory-based semi-physical simulations supplement data collection under extreme interference conditions, such as condensation levels 4-5 and water film levels 4-5. These simulations address issues that are difficult to reproduce naturally at high frequencies in greenhouses, by constructing a simulated environment in the laboratory for additional data acquisition. Condensation simulation: Freshly collected insect-infested leaf samples were placed in a temperature and humidity controlled chamber, and the temperature and humidity inside the chamber were adjusted to induce condensation of varying degrees on the lens surface. Water film simulation: Deionized water is quantitatively sprayed onto the blade surface using a micro-injection pump, and the water film thickness is adjusted by controlling the spraying amount.

[0028] The annotation was performed in Pascal VOC format by annotators with a background in plant protection. Individual pests that could be identified by the naked eye in the image were annotated one by one. If the individuals were too dense and the boundaries were indistinguishable (such as a community formed by aphids), the entire community was marked as a separate annotation unit. The annotation box was closely attached to the outer edge of the outline of the insect / community without leaving any blank margins. If the insect was more than 50% covered by leaves, it was not annotated. When the coverage was ≤50%, the estimated complete outline was added to the visible part. Each image is independently annotated by two annotators, and then a third person performs a consistency comparison after completion. The dataset constructed in this embodiment has a total size of 8,000 images, of which 70% are used as the training set, 10% as the validation set, and 20% as the test set. This step involves building... and This approach transforms environmental disturbances, which were previously only qualitatively describable, into numerical levels of 1-5, making the degree of disturbance measurable and comparable. Existing technologies generally employ indiscriminate data collection and general augmentation to construct datasets, failing to quantitatively characterize the lens condensation and leaf water film disturbances unique to greenhouse environments. This results in a sharp drop in model accuracy of 30%-50% in actual greenhouse environments. This step, for the first time, proposes two quantifiable environmental disturbance indices, achieving numerical grading of the disturbance degree rather than subjective qualitative description. This ensures that the dataset covers all typical disturbance combinations. Furthermore, it uses individual pest labeling instead of leaf-level labeling, solving the problem of missed detections caused by insufficient granularity in labeling of minute pests.

[0029] S2: Perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image. In this embodiment, the global fog effect removal process employs an improved dark channel prior algorithm, incorporating the lens condensation interference index. Dynamically adjusting algorithm parameters specifically includes: calculating the dark channel image of the input fruit tree pest image, and the size of the dark channel calculation window. Depend on Dynamic adjustment, the formula is: in, This is a rounding down function; it selects the top 0.1% of bright pixels from the dark channel image, and the maximum gray value of these pixels in the original image is taken as the atmospheric light value. Increase: Increasing the window size enhances defogging intensity; traditional dark channel prior algorithms use a fixed window size, but in cases of severe condensation, a small window cannot effectively capture the global fog effect, resulting in insufficient defogging; this solution adjusts the window size according to the desired size. Linear increase: The more severe the condensation, the larger the window is needed to estimate atmospheric light and transmittance; setting the step size to 5 is an empirical value that ensures significant differences in defogging effects between different levels without causing a sharp increase in computational load; the floor function ensures that the window size can only be an integer, which meets the requirements of convolution operations.

[0030] Combination The transmittance is calculated by dynamically adjusting the correction factor, using the following formula: in, coordinates Transmittance at that location For the dark channel image in coordinates The dark channel value at that location, To dynamically adjust the correction coefficient, This refers to atmospheric light values; in traditional dark channel prior algorithms... Fixed value, but when condensation is severe, fixed This can lead to an underestimation of transmittance and an overly dark image; in this solution Follow Dynamic adjustment: hour , hour This ensures effective dehazing even with slight condensation while avoiding excessively dark images with severe condensation. The dehazed image is generated based on a dark channel prior model, using the following formula: in, To dehaze image pixels in coordinates grayscale value at that location Fruit tree pest images in coordinates grayscale value at that location The minimum transmittance threshold is used; images with different levels of condensation interference are randomly selected from the dataset constructed in S1 to form... Optimize dedicated datasets and set them separately. Four candidate values ​​(0.05, 0.1, 0.15, and 0.2) were used. An objective index combined with subjective evaluation was employed to assess the dehazing effect. An improved dark channel prior dehazing algorithm was applied to all images in the dataset. Experimental results show that when… At that time, the average PSNR reached 30.42 dB across all interference levels, the average SSIM reached 0.845, and the subjective evaluation score was the highest. If This can lead to severe over-dehazing and color distortion; if If this is not done, incomplete defogging will occur. Therefore, it is necessary to determine the minimum transmittance threshold. To prevent the denominator from being too small, which would amplify noise, i.e., cause the image to be overexposed due to low transmittance.

[0031] In this embodiment, the detection and repair of local reflective areas specifically includes: based on the leaf water film interference index... The reflectivity threshold can be dynamically set using the following formula: in, This is the reflectivity threshold. The higher the grade, the lower the threshold, and the more sensitive the detection. Traditional reflective detection uses a fixed threshold, which cannot adapt to the dynamic changes in the reflective intensity of water film, and is prone to missed detection or false detection. The more severe the water film, the greater the reflective intensity, and a higher threshold is needed to distinguish reflective pixels from normal pixels. The basic threshold of 20 is an engineering statistical value: when there is no water film, the difference between the pixel gray value of a normal leaf and the mean of the neighborhood is usually less than 20.

[0032] For dehazed images Perform grayscale conversion to obtain a grayscale image. Calculate the difference between the gray value of each pixel and the mean of its surrounding 3×3 neighborhood. ,when When this happens, the pixel is determined to be a reflective pixel, and all reflective pixels constitute a reflective area. The texture filling method was used to repair the reflective areas, and the reflective areas were selected. The surrounding 5×5 non-reflective area serves as a texture sample library, while the reflective area... The image is divided into 5×5 blocks of the same size as the texture blocks in the sample library. The normalized cross-correlation coefficient between each texture block in the sample library and the corresponding block to be repaired is calculated. The texture block with the largest coefficient is selected to replace the block to be repaired until all reflective areas are repaired, and the de-reflective image is obtained.

[0033] In this embodiment, the adaptive enhancement of low-light contrast specifically includes: calculating the de-reflected image. The grayscale mean is used as the illumination intensity coefficient. The formula is: in, Image width, Image height; based on illumination intensity coefficient Dynamically adjust the enhancement level to generate an enhanced image. The formula is: in, To enhance the image, the pixel values ​​of the enhanced image are normalized to 0-255 to obtain a high-quality restored image. This formula uses a linear enhancement method, where the enhancement magnitude is inversely proportional to the light intensity: the weaker the light, the greater the enhancement magnitude. In this embodiment, if k is set too large, the image will be overexposed, and details in highlight areas will be lost; if k is set too small, the low-light image will not be sufficiently enhanced, and details of insect damage will remain blurry. The enhancement coefficient in this embodiment... It is an engineering experience value: it can effectively improve the contrast of low-light images without causing overexposure.

[0034] This step dynamically binds the interference index to the restoration algorithm parameters, achieving adaptive processing where the more severe the interference, the greater the restoration strength. It employs a three-level restoration process of global dehazing, local de-reflection, and low-light enhancement, specifically addressing the three main image quality problems in greenhouses. It eliminates the impact of environmental interference on images, enabling the model to focus on the characteristics of the pests themselves, significantly improving the detection accuracy of small pests. Adaptive parameter adjustment avoids the tediousness of manual parameter tuning, allowing the system to automatically adapt to environmental changes in different greenhouses and at different times.

[0035] S3: Based on the convolutional neural network backbone network, multi-level feature extraction is performed on the obtained high-quality restored image to output multi-scale feature maps with different downsampling factors; cross-layer dense fusion is performed on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; feature processing is performed on the obtained high-resolution fused feature maps to output enhanced high-resolution feature maps. The convolutional neural network backbone network selected is the lightweight MobileNetV3-small, with only 1.5M parameters, far lower than ResNet50's 25.6M, resulting in fast inference speed and suitability for deployment on large-scale mobile devices. It incorporates depthwise separable convolution and attention mechanisms, ensuring feature extraction capabilities while maintaining a lightweight design, achieving over 15% improvement in feature representation compared to MobileNetV2. It is adapted for small-sample training, and combined with the augmented dataset of S1, it effectively avoids overfitting. The network configuration is as follows: the input layer receives a 224×224×3 RGB image and normalizes it to [-1,1]; the feature extraction layer contains 16 convolutional blocks, of which the first 13 are depthwise separable convolutional blocks and the last 3 are standard convolutional blocks, with a BatchNorm2d normalization layer and a Swish activation function added after each convolutional block; the output layer outputs four multi-scale feature maps with different downsampling factors, namely C1, C2, C3, and C4. In this embodiment, C1 is 56×56×16, downsampled by 4 times, C2 is 56×56×16, C3 is 56×56×16, C4 ... 2: 28×28×24, downsampled 8 times; C3: 14×14×40, downsampled 16 times; C4: 7×7×96, downsampled 32 times. C1 corresponds to shallow features: pest edges and textures; C4 corresponds to deep features: pest semantics and category features. The cross-layer dense fusion specifically includes: upsampling the deep feature maps C4, C3, and C2 sequentially to a size of 56×56 using bilinear interpolation; and concatenating the upsampled C4, C3, and C2 with the shallow feature map C1 through channels to obtain a concatenated feature map. 176 channels; 1×1 convolution is used to... The number of channels was adjusted to 64, and then an SE channel attention mechanism was introduced to weight the channel features, resulting in a high-resolution fused feature map. The SE channel attention mechanism automatically assigns higher weights to important channels and suppresses interference from unimportant channels by learning the dependencies between channels. The feature processing of the obtained high-resolution fused feature map specifically involves feature processing to suppress interference and enhance the target, employing a spatial-channel dual-branch collaborative attention mechanism, specifically including: Global average pooling is performed, followed by two fully connected layers and an activation function to obtain the channel attention weights. ;right Perform 1×1 convolution to reduce dimensionality to 1 channel, and then obtain spatial attention weights using the Sigmoid activation function. ; Using a broadcast mechanism to and Multiplication yields the collaborative attention weights. ,Will and Pixel-by-pixel multiplication yields the enhanced high-resolution feature map. Specific operations: Feature upsampling: Deep feature maps C4, C3, and C2 are sequentially upsampled to the same size as the next layer feature map, using bilinear interpolation with upsampling factors of 4x (C4→56×56), 2x (C3→56×56), and 2x (C2→56×56), respectively, ensuring all feature maps are uniformly 56×56 in size.

[0036] Perform dense stitching: The upsampled C4, C3, and C2 are stitched together with the shallow feature map C1 through channels to obtain the stitched feature map. The number of channels is 16 + 24 + 40 + 96 = 176, and the dimensions are 56 × 56 × 176. The formula is: ) in, Indicates upsampling operation; channel adjustment and attention weighting: through 1×1 convolution... The number of channels was adjusted to 64 to reduce computation. Then, an SE channel attention mechanism was introduced to weight channel features, highlighting pest-related channels and suppressing irrelevant channels. The formula is as follows: in, The high-resolution fused feature map is obtained after channel adjustment and SE channel attention weighting. The calculation process of the SE attention mechanism is as follows: Global average pooling: We obtain a 64×1 channel feature vector; additionally, the fully connected layer mappings are as follows: , Reduce the number of channels from 64 to 32; , Increasing the number of channels from 32 to 64 yields the channel attention weights. (64×1); Weighted fusion: in, Using channel indexing, a high-resolution fused feature map is ultimately obtained that preserves minute details of insect damage. Finally, a dual-branch collaborative attention module was designed to enhance the pest target features from both spatial and channel dimensions, while suppressing background interference such as leaf texture, soil, and dew. Specifically, the channel branch attention calculation involves... Global average pooling is performed to obtain channel feature vectors (64×1). Channel attention weights are then obtained by passing them through two fully connected layers and an activation function. The formula is as follows: The value ranges from 0 to 1; the larger the value, the more important the pest characteristics of the corresponding channel. Spatial branch attention calculation: for Perform 1×1 convolution dimensionality reduction to reduce the number of channels from 64 to 1, and then pass the sigmoid activation function to obtain the spatial attention weights. The formula is: The value ranges from 0 to 1. A larger value indicates a more prominent pest target feature at the corresponding location and weaker background interference. Collaborative enhancement: A broadcast mechanism is used to distribute channel attention weights. Spatial attention weights Multiply to obtain the collaborative attention weights. The formula is: Combine collaborative attention weights with Pixel-by-pixel multiplication yields the enhanced high-resolution feature map. The formula is: After completing multi-level feature extraction and dual-branch attention enhancement, we obtain This feature map not only preserves the detailed features of minor insect damage but also effectively suppresses background interference, providing accurate feature support for subsequent insect damage detection.

[0037] S4: Construct a detection head specifically for small pests based on the enhanced high-resolution feature map. Based on the annotation of individual target pests, generate anchor boxes adapted to the size of the target pests through clustering. Output the pest classification and localization results of the corresponding anchor boxes through the detection head. Construct a multi-task joint loss function. Train and optimize the convolutional neural network model based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. In this embodiment, anchor boxes adapted to the size of the target pests are generated through clustering. Specifically, the K-means++ clustering algorithm is used to generate anchor boxes adapted to the size of the target pests, and the elbow rule is used in combination with the size distribution characteristics of agricultural pests to determine the optimal number of clusters. Compared with the traditional K-means algorithm, K-means++ can effectively avoid the clustering results from getting trapped in local optima by optimizing the selection of the initial cluster centers, making the generated anchor boxes more in line with the size distribution of real pests and the clustering stability is higher.

[0038] In this embodiment, the formula for the multi-task joint loss function is: in, The total value of the joint loss across multiple tasks. For classifying losses, To pinpoint the loss, Weighted loss for small goals These are the weighting coefficients for each loss; in this embodiment, the classification loss weights... Classification is the core foundation of fruit tree pest identification, and accurate pest classification is a prerequisite for developing targeted control measures. The classification loss weight is set to 1.0 as the baseline weight to ensure the model prioritizes learning the category features of different pests, avoiding category confusion. The small target weighted loss has the highest weight, prioritizing the detection accuracy of minor pests. The localization loss weight... Fruit tree pests are mostly tiny targets, and accurate location is not only a necessary condition for pest identification, but also the foundation for subsequent pest counting and damage assessment. Because the boundaries of tiny targets are blurred and their features are not obvious, their location is significantly more difficult than classification. Therefore, a higher weight is given to the location loss than the classification loss to guide the model to pay more attention to the boundary features of pest targets and improve location accuracy; the weight of the small target weighted loss is adjusted accordingly. Tiny insect pests (such as aphids, spider mites, and thrips) are the most damaging and difficult-to-detect pests in agricultural production. They reproduce rapidly, are highly concealed, and can easily cause outbreaks in a short period of time. Existing detection models generally suffer from low accuracy in detecting tiny insect pests. Therefore, we assign the highest weight to small targets, forcing the model to prioritize learning the features of tiny insect pests, which significantly improves the recall and accuracy of tiny insect pest detection.

[0039] in, The classification loss, using cross-entropy loss, is employed to optimize the prediction accuracy of pest categories. The formula is as follows: in, The number of positive samples. The number of anchor frames ≥ 0.5; The number of negative samples (0.1≤ (Number of anchor frames < 0.5); For the first The true labels for each positive sample are: 1 represents the corresponding pest category, and 0 represents the background. For the first The predicted class probability of a positive sample; For the first The true label of each negative sample is fixed at 0, i.e., the background; For the first The predicted class probability of each negative sample; this loss function can effectively optimize the accuracy of class prediction and avoid the model bias towards the background class prediction caused by the imbalance of positive and negative samples. The positioning loss uses GIoU loss, which, compared to traditional IoU loss, effectively solves the problem of zero loss and inability to optimize when the anchor box and annotation box do not overlap. It is suitable for the precise positioning of tiny insect pests. The formula is: in: To predict the coordinates of the anchor frame ( , , , ); The coordinates of the actual annotation box ( , , , ); The calculation process is as follows: ,in To predict the anchor frame, This is the actual annotation box. For inclusion and The smallest bounding rectangle, |·| represents the area of ​​the region; The value range is [-1, 1]. The closer it is to 1, the more accurate the positioning of the anchor box and the label box.

[0040] The small-target weighted loss is specifically designed for tiny pests, with an annotation box area < 100 pixels². The formula is as follows: in, Number of samples of minute insect pests; , These are the predicted anchor box and the actual labeled box for the kth minute insect pest, respectively; The weighting coefficients for smaller objectives are given by the following formula: , For the first The area (pixels²) of each tiny pest label box. The smaller, The larger the value, the higher the weight given to smaller pests, ensuring the accuracy of locating and classifying minute pests.

[0041] The hyperparameters for model training were set as follows: batch size 32, initial learning rate 0.001, cosine annealing learning rate adjustment strategy, weight decay coefficient 0.0001, 200 iterations, early stopping strategy with patience=20; during training, a gradient clipping strategy was introduced to limit the gradient norm to within 5.0, and a Dropout layer was added between the feature extraction layer and the detection head with a dropout rate=0.3.

[0042] S5: After processing the fruit tree images in the greenhouse to be detected through steps S2 and S3, input them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. Perform post-processing steps on the results and output the final identification results.

[0043] The post-processing steps described in S5 are as follows: setting a classification confidence threshold and removing anchor boxes with a classification confidence score less than the threshold; using a non-maximum suppression algorithm to remove overlapping anchor boxes and setting an NMS threshold; converting the anchor box offsets output by the model into image pixel coordinates; and outputting the final recognition result, which also includes a confidence evaluation step. In this embodiment, the classification confidence threshold is set to 0.7, and anchor boxes with a classification confidence < 0.7 are removed; non-maximum suppression (NMS) algorithm is used to remove overlapping anchor boxes, and the NMS threshold is set to 0.3; the anchor box offsets output by the model are converted into image pixel coordinates; the embodiment also includes a step for evaluating the confidence of the detection results: the confidence is calculated based on the interference level and the classification confidence, using the following formula: in, For credibility, For classification confidence, For condensation interference level, The water film interference level is used to determine the reliability of the prediction results. The reliability depends primarily on two factors: the model's classification confidence and the degree of environmental interference. Higher model classification confidence leads to higher reliability, while more severe environmental interference results in lower reliability. The weights of 0.6 and 0.4 are assigned because model classification confidence has a greater impact on reliability than environmental interference. 0.2 is the interference penalty coefficient: for every level increase in interference, reliability decreases by 10%. Reliability is divided into three levels: high reliability... ≥90%: Level 0~1 interference, classification confidence ≥0.8, detection results can be directly used for pest control decision-making; Medium confidence ≤70% <90%: Level 2 interference, classification confidence level 0.7~0.8, on-site verification by staff is recommended; low confidence level. <70%: Level 3-4 interference, classification confidence level ≥0.7, it is recommended to re-acquire images for detection, or use manual detection for confirmation; output the corresponding confidence level.

[0044] This embodiment selects 31 pest detection box samples under different condensation interference levels (1-5), different water film interference levels (1-5), and different lighting conditions (weak light, normal light, and strong light combinations), covering both correct and false detection results. The reliability level of each detection result is evaluated according to the reliability calculation formula of this scheme. Each image undergoes adaptive image restoration, multi-scale feature extraction and enhancement, inference using a micro-pest detection model, and post-processing to obtain the category classification confidence score for each detection box. Based on the confidence calculation formula in step five of this scheme, and combined with the LDI and LWI levels corresponding to the image, the confidence level Cr of each detection box is calculated and divided into three levels: high, medium, and low, according to the threshold. The original experimental data are shown in Table 1 below: Table 1: Reliability Evaluation Table for Multi-Interference Detection Please see Figure 2 This table reflects the quantitative relationship between environmental interference level, model classification confidence, and detection result reliability. The data trends show that under conditions of no interference or slight interference (LDI+LWI≤2), the model classification confidence is generally higher than 0.90, and the reliability exceeds 90%, falling into the "high reliability" range. As the interference level increases, both classification confidence and reliability decrease simultaneously. False positives (Img15-bbox_01, Img16-bbox_01, Img31-bbox_01) all appear under high interference conditions, and their reliability is at the "low" or "medium" level, indicating that false positives under high interference can be effectively identified through reliability levels. Furthermore, under the same interference conditions, the classification confidence of spider mites is usually slightly lower than that of aphids, consistent with the reality that spider mites are smaller and their features are more easily obscured. The data in this table validates the rationality of the model probability weight of 0.6 and the environmental interference weight of 0.4 in the reliability formula, as well as the effectiveness of the environmental penalty coefficient of 0.1.

[0045] Please see Figure 3 The present invention also provides a fruit tree pest identification system based on a convolutional neural network, used to execute the above-mentioned fruit tree pest identification method based on a convolutional neural network, comprising: Multi-interference dataset construction module: used to construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level, collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and perform target pest individual labeling, and divide the dataset into training set, validation set and test set. Image restoration module: used to perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image; Feature extraction and enhancement module: Based on the convolutional neural network backbone network, it performs multi-level feature extraction on the obtained high-quality restored image, and outputs multi-scale feature maps with different downsampling factors; it performs cross-layer dense fusion on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; it performs feature processing on the obtained high-resolution fused feature maps and outputs enhanced high-resolution feature maps. Model training and optimization module: This module is used to construct a detection head specifically for small pests based on the enhanced high-resolution feature maps. Based on the annotation of individual target pests, it generates anchor boxes adapted to the size of the target pests through clustering. The detection head outputs the pest classification and localization results corresponding to the anchor boxes. A multi-task joint loss function is constructed, and the convolutional neural network model is trained and optimized based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. Pest and disease identification output module: After processing the images of fruit trees in the greenhouse to be detected in steps S2 and S3, the module inputs them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. The module then performs post-processing steps to output the final identification result.

[0046] The above formulas are all dimensionless calculations. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0047] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0048] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0049] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for identifying fruit tree pests based on convolutional neural networks, characterized in that, The specific steps include: S1: Construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level. Collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and label target pest individuals. Divide the dataset into training set, validation set and test set. S2: Perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image. S3: Based on the convolutional neural network backbone network, multi-level feature extraction is performed on the obtained high-quality restored image to output multi-scale feature maps with different downsampling factors; cross-layer dense fusion is performed on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; feature processing is performed on the obtained high-resolution fused feature maps to output enhanced high-resolution feature maps. S4: Construct a detection head specifically for small pests based on the enhanced high-resolution feature map. Based on the annotation of individual target pests, generate anchor boxes adapted to the size of the target pests through clustering. Output the pest classification and localization results of the corresponding anchor boxes through the detection head. Construct a multi-task joint loss function. Train and optimize the convolutional neural network model based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. S5: After processing the fruit tree images in the greenhouse to be detected through steps S2 and S3, input them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. Perform post-processing steps on the results and output the final identification results. The lens condensation interference index The leaf water film interference index is constructed by combining the image grayscale mean and contrast. Constructed by combining average image reflectance and edge sharpness; based on the above and The threshold range for the levels is set, and the levels of condensation interference and water film interference are divided into 5 levels from 1 to 5, and the level values ​​are directly used in subsequent calculations; at the same time, the light intensity is divided into three categories: weak light range, normal light range, and strong light range.

2. The method for identifying fruit tree pests based on convolutional neural networks according to claim 1, characterized in that: The lens condensation interference index The calculation formula is: in, The average gray level of the image; Image contrast; To preset the maximum contrast value, Image grayscale weights, Image contrast weights, satisfying , ; The value range is 0 to 1; The blade water film interference index The calculation formula is: in, The average reflectance of the image. This represents the maximum reflectivity. For image edge sharpness, This represents the maximum edge sharpness. Weights for average image reflectance. As the image edge sharpness weight, satisfying , , The value range is 0 to 1.

3. The method for identifying fruit tree pests based on convolutional neural networks according to claim 1, characterized in that: The global fog removal process employs an improved dark channel prior algorithm, incorporating the lens condensation interference index. The algorithm parameters are dynamically adjusted, specifically including: calculating the dark channel image of the input fruit tree pest image, and the size of the dark channel calculation window. Depend on Dynamic adjustment, the formula is: in, This is a floor function; the top 0.1% of bright pixels are selected from the dark channel image, and the maximum grayscale value of these pixels in the original image is used as the atmospheric light value; combined with... The transmittance is calculated by dynamically adjusting the correction factor, using the following formula: in, coordinates Transmittance at that location For the dark channel image in coordinates The dark channel value at that location, To dynamically adjust the correction coefficient, The atmospheric light value is used; a dehazed image is generated based on the dark channel prior model, using the following formula: in, To dehaze image pixels in coordinates grayscale value at that location Representing fruit tree pest images in coordinates grayscale value at that location This is the minimum transmittance threshold.

4. The method for identifying fruit tree pests based on convolutional neural networks according to claim 3, characterized in that: The localized reflective area detection and repair process specifically includes: based on the leaf water film interference index... The reflectivity threshold can be dynamically set using the following formula: in, The reflection threshold is used for dehazed images. Perform grayscale conversion to obtain a grayscale image. Calculate the difference between the gray value of each pixel and the mean of its surrounding 3×3 neighborhood. ,when When this happens, the pixel is determined to be a reflective pixel, and all reflective pixels constitute a reflective area. The texture filling method was used to repair the reflective areas, and the reflective areas were selected. The surrounding 5×5 non-reflective area is used as a texture sample library. The normalized cross-correlation coefficient between each texture block in the sample library and the block to be repaired is calculated. The texture block with the largest coefficient is selected to replace the block to be repaired until all reflective areas are repaired, and the de-reflective image is obtained.

5. The method for identifying fruit tree pests based on convolutional neural networks according to claim 4, characterized in that: The adaptive enhancement of low-light contrast specifically includes: calculating the grayscale mean of the dereflected image as the illumination intensity coefficient. The formula is: in, Image width, Image height; based on illumination intensity coefficient Dynamically adjust the enhancement level to generate an enhanced image. The formula is: in, To enhance the coefficients, the pixel values ​​of the enhanced image are normalized to 0~255 to obtain a high-quality restored image.

6. The method for identifying fruit tree pests based on convolutional neural networks according to claim 1, characterized in that: In step S3, the convolutional neural network backbone network is MobileNetV3-small, and the network configuration is as follows: the input layer receives a 224×224×3 RGB image and normalizes it to [-1,1]; the feature extraction layer contains 16 convolutional blocks, of which the first 13 are depth-separable convolutional blocks and the last 3 are standard convolutional blocks, and a BatchNorm2d normalization layer and a Swish activation function are added after each convolutional block; the output layer outputs four multi-scale feature maps with different downsampling factors, namely C1, C2, C3, and C4; the cross-layer dense fusion specifically includes: upsampling the deep feature maps C4, C3, and C2 to a size of 56×56 using bilinear interpolation; and concatenating the upsampled C4, C3, and C2 with the shallow feature map C1 through channels to obtain a concatenated feature map. The number of channels is 176; 1×1 convolution is used to... The number of channels was adjusted to 64, and then an SE channel attention mechanism was introduced to weight the channel features, resulting in a high-resolution fused feature map. The feature processing of the obtained high-resolution fused feature map employs a spatial-channel dual-branch collaborative attention mechanism, specifically including: […]. Global average pooling is performed, followed by two fully connected layers and an activation function to obtain the channel attention weights. ;right Perform 1×1 convolution to reduce dimensionality to 1 channel, and then obtain spatial attention weights using the Sigmoid activation function. ; Using a broadcast mechanism to and Multiplication yields the collaborative attention weights. ,Will and Pixel-by-pixel multiplication yields an enhanced high-resolution feature map.

7. The method for identifying fruit tree pests based on convolutional neural networks according to claim 6, characterized in that: The formula for the multi-task joint loss function mentioned in step S4 is: in, The total value of the joint loss across multiple tasks. For classifying losses, To pinpoint the loss, Weighted loss for small goals These are the weighting coefficients for each loss. In step S4, the hyperparameters for model training are set as follows: batch size 32, initial learning rate 0.001, cosine annealing learning rate adjustment strategy, weight decay coefficient 0.0001, 200 iterations, early stopping strategy and patience=20; during training, a gradient clipping strategy is introduced to limit the gradient norm to within 5.0, and a Dropout layer is added between the feature extraction layer and the detection head with dropout rate=0.

3.

8. The method for identifying fruit tree pests based on convolutional neural networks according to claim 1, characterized in that: The post-processing steps described in S5 are as follows: Setting a classification confidence threshold, removing anchor boxes with a classification confidence level less than the threshold; using a non-maximum suppression algorithm to remove overlapping anchor boxes and setting an NMS threshold; converting the anchor box offsets output by the model into image pixel coordinates; and outputting the final recognition result, which also includes a confidence evaluation step: calculating the confidence level based on the interference level and classification confidence level, using the following formula: in, For credibility, For classification confidence, For condensation interference level, The level of water film interference is categorized; the credibility is divided into three levels: high credibility. ≥90%, Medium confidence level ≤70% <90%, low credibility <70%, and output the corresponding confidence level.

9. A fruit tree pest identification system based on convolutional neural networks, characterized in that: The fruit tree pest identification system based on convolutional neural networks is used to execute the fruit tree pest identification method based on convolutional neural networks as described in any one of claims 1-8, including: Multi-interference dataset construction module: used to construct lens condensation interference index and leaf water film interference index, and classify condensation interference level and water film interference level, collect fruit tree pest image dataset containing different condensation interference level, water film interference level and light intensity, and perform target pest individual labeling, and divide the dataset into training set, validation set and test set. Image restoration module: used to perform global fog removal processing on the fruit tree pest image obtained in step S1 to obtain a defogging image; then perform local reflective area detection and repair processing on the defogging image to obtain a de-reflective image; perform weak light contrast adaptive enhancement on the de-reflective image to obtain a high-quality restored image; Feature extraction and enhancement module: Based on the convolutional neural network backbone network, it performs multi-level feature extraction on the obtained high-quality restored image, and outputs multi-scale feature maps with different downsampling factors; it performs cross-layer dense fusion on the obtained multi-scale feature maps to obtain high-resolution fused feature maps; it performs feature processing on the obtained high-resolution fused feature maps and outputs enhanced high-resolution feature maps. Model training and optimization module: This module is used to construct a detection head specifically for small pests based on the enhanced high-resolution feature maps. Based on the annotation of individual target pests, it generates anchor boxes adapted to the size of the target pests through clustering. The detection head outputs the pest classification and localization results corresponding to the anchor boxes. A multi-task joint loss function is constructed, and the convolutional neural network model is trained and optimized based on the dataset obtained in step S1 to obtain the trained fruit tree pest identification model. Pest and disease identification output module: After processing the images of fruit trees in the greenhouse to be detected in steps S2 and S3, the module inputs them into the trained fruit tree pest identification model to obtain the detection results containing the pest classification confidence and anchor frame offset of the corresponding anchor frame. The module then performs post-processing steps to output the final identification result.

Citation Information

Patent Citations

  • A method for identifying fruit tree diseases and pests under complex background based on cavity convolution

    CN109344883A

  • Crop pest identification method and device

    CN109409170A

  • Plant disease and pest identification method based on attention mechanism and multi-level convolution characteristics

    CN110188635A

  • Crop disease and insect pest leaf identification system, identification method and disease and insect pest prevention method

    CN113468984A

  • Litchi insect pest monitoring and early warning method and system based on Internet of Things, and storage medium

    CN114445785A