Lemon orchard pest and disease damage detection method, device, equipment and medium in complex environment

By improving the YOLOv11n model, introducing the RepViT network and RFAConv module, and optimizing the data augmentation strategy, the problem of pest and disease detection in complex orchard environments was solved, achieving efficient and accurate pest and disease identification, reducing equipment costs, and improving agricultural production efficiency and sustainability.

CN120953674APending Publication Date: 2025-11-14SOUTH CHINA AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511062794.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

In complex orchard environments, existing technologies are hampered by background noise and varying light intensity, making it difficult for models to fully capture the complex manifestations of pests and diseases. Furthermore, the multi-scale feature fusion capability is insufficient, and the equipment is expensive, making it difficult for farmers to use directly.

Method used

An improved YOLOv11n model was adopted, with the RepViT network introduced as the backbone network. The training process was optimized by combining the RFAConv receptive field attention module and the DySample upsampling module. Combined with data augmentation strategies of reinforcement learning, a method for detecting pests and diseases in lemon orchards under complex environments was constructed.

Benefits of technology

It significantly improves the model's ability to detect pests and diseases in complex environments, enhances the comprehensiveness and accuracy of detection, reduces computational overhead and equipment costs, simplifies farmer operations, and improves agricultural production efficiency and environmental sustainability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953674A_ABST
    Figure CN120953674A_ABST
Patent Text Reader

Abstract

The invention relates to a lemon orchard disease and pest detection method, device and equipment in a complex environment and a medium. The method comprises the following steps: acquiring a to-be-detected lemon orchard image containing lemon diseases and pests; the method comprises the following steps: updating a backbone network of a preset first disease and insect pest detection model into a RepViT network, introducing an RFAConv receptive field attention module into a C3k2 module on a neck network of the first disease and insect pest detection model to construct a C3k2RFA module, and updating Upsample up-sampling modules of original feature pyramid layers P2, P3 and P4 in the neck network into DySample modules to construct a second disease and insect pest detection model; and inputting the to-be-detected lemon orchard image into a second pest and disease damage detection model trained to a convergence state to detect lemon pest and disease damage in the to-be-detected lemon orchard image so as to complete lemon orchard pest and disease damage detection in a complex environment. According to the method, the detection effect of the model under the complex background can be greatly improved, the calculation overhead is reduced, and the energy consumption of equipment can also be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition, and in particular to a method, apparatus, electronic device and computer-readable storage medium for detecting pests and diseases in lemon orchards under complex environments. Background Technology

[0002] As one of the world's most important economic crops, lemons face numerous pest and disease threats during cultivation, severely impacting fruit quality and yield. Traditional pest and disease identification methods rely primarily on the experience of agricultural experts, which suffers from limitations such as low efficiency, high cost, and strong subjectivity. With the development of technologies like computer vision and deep learning, intelligent pest and disease identification technology based on image analysis has gradually become a research hotspot.

[0003] However, existing object detection technologies based on image processing and deep learning still have the following limitations, including:

[0004] Firstly, in complex orchard environments, background noise, light intensity, and other background and environmental conditions are not uniform, which may interfere with the identification of pests and diseases, resulting in insufficient feature extraction of specific detection objects and failure to fully extract the key features of pests and diseases.

[0005] Secondly, the diseases and pests affecting lemon trees are complex and diverse. For example, the large-scale disease characteristics, such as the pathological conditions of the plant at different stages and the growth of the plant after being affected by pests, are quite complex. When dealing with diseases and pests with different scale characteristics, such as small target pests like whiteflies and citrus psyllids, there are often problems of missed detection and false detection. Sooty mold and yellow spot disease sometimes have similar black spots, making it difficult for the model to fully capture the complex manifestations of disease and pest targets. The model's ability to fuse multi-scale features is insufficient, which affects the accuracy of the detection results.

[0006] Third, the deployment of models is mostly applied to edge devices, such as drones, intelligent robot platforms, and embedded intelligent cameras, which often requires a certain cost, placing a certain economic burden on farmers. At the same time, it requires a certain level of professional expertise, making it inconvenient for farmers to use directly.

[0007] In summary, in order to address the problems in existing technologies, such as background noise and light intensity interfering with the identification of pests and diseases in complex orchard environments, the inability of existing models to fully capture the complex characteristics of pests and diseases, and the insufficient multi-scale feature fusion capabilities, the applicant has made corresponding explorations to solve these problems. Summary of the Invention

[0008] The purpose of this application is to solve the above-mentioned problems by providing a method, device, electronic equipment and computer-readable storage medium for detecting pests and diseases in lemon orchards under complex environments.

[0009] To achieve the various objectives of this application, the following technical solution is adopted:

[0010] A method for detecting pests and diseases in lemon orchards under complex environments, proposed to meet one of the purposes of this application, includes:

[0011] In response to an instruction to detect diseases and pests in a lemon orchard under complex conditions, the system acquires an image of the lemon orchard to be detected containing lemon diseases and pests, wherein the lemon diseases and pests include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly, and citrus psyllid.

[0012] The backbone network of the preset first pest and disease detection model is updated to a RepViT network. The RFAConv receptive field attention module is introduced into the C3k2 module of the neck network of the first pest and disease detection model to construct the C3k2_RFA module. The Upsampling module of the original feature pyramid P2, P3, P4 layers in the neck network is updated to a DySample module to construct the second pest and disease detection model.

[0013] The image of the lemon orchard to be detected is input into the second pest and disease detection model that has been trained to convergence, so as to detect lemon pests and diseases in the image of the lemon orchard to complete the detection of lemon pests and diseases in complex environments.

[0014] Optionally, in the ViTBlock module of the RepViT network, a 3×3 depth-separable convolutional layer is provided after a 1×1 extended convolutional layer for spatial information fusion. The depth-separable convolutional layer is optimized using structural reparameterization technology, a multi-branch topology is introduced through structural reorganization, and the expansion ratio in the convolutional layer is adjusted. The expansion ratio is set to 2 in the channel mixer of all stages.

[0015] Optionally, in the macroscopic design optimization, the RepViT network uses two 3×3 convolutional layers with a stride of 2 as the starting module, wherein the number of filters in the first convolutional layer is set to 24 and the number of filters in the second convolutional layer is set to 48; a depthwise separable convolutional layer with a stride of 2 and a 1×1 point convolutional layer are used to perform spatial downsampling and channel dimension adjustment respectively; a RepViT module is added before the downsampling layer to further deepen the network depth, and a feedforward network module is placed after the 1×1 convolutional layer to memorize more potential information;

[0016] In the micro-design adjustments, all modules preferentially use 3×3 size convolution kernels, and the compressed excitation layers are alternately deployed in the odd number of blocks in each stage.

[0017] Optionally, the step of inputting the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the image of the lemon orchard to be detected includes:

[0018] In the C3k2_RFA module, the spatial features of each receptive field in the pest and disease feature map are extracted by the grouped convolution method. The pest and disease feature map includes pest and disease features, which characterize the texture of lesions or the morphology of insect bodies.

[0019] Average pooling is used to aggregate global information of spatial features of each receptive field, and the information is interacted through 1×1 convolution operations to learn the correlation between different channels;

[0020] The output of the 1×1 convolution is normalized using the softmax function to generate importance weights for the spatial features of each receptive field.

[0021] The importance weights are combined with the spatial features of the original receptive field extracted by grouped convolution to output an enhanced pest and disease feature map for subsequent pest and disease detection.

[0022] Optionally, the step of inputting the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the image of the lemon orchard to be detected includes:

[0023] The enhanced pest and disease feature map obtained after processing by the C3k2_RFA module is input into the DySample module. In the sampling point generator, the static range factor is used to generate the first offset by combining the linear layer and pixel shuffling technology with the fixed first range factor.

[0024] In the sampling point generator, a second range factor is first generated using a dynamic range factor, and then the first offset is adjusted using the second range factor to determine the second offset. The second offset is added to the original grid position to generate a sampling set. The second range factor is generated by the Sigmoid function.

[0025] Based on the generated sample set, bilinear interpolation is used to resample the enhanced pest and disease feature map to generate an upsampled feature map for subsequent pest and disease detection.

[0026] Optionally, the steps for training the second pest and disease detection model include:

[0027] Obtain a sample dataset, wherein the sample dataset includes multiple training samples and their corresponding supervision labels, the training samples represent a single lemon orchard image, and the supervision labels represent the annotation information of lemon diseases and pests in the lemon orchard image;

[0028] A data augmentation strategy combination is randomly generated, wherein the data augmentation strategy combination includes multiple sub-strategy combinations, including shear transformation, translation, rotation, contrast adjustment, histogram equalization, reducing the number of bits in each color channel of the image, brightness adjustment, sharpness adjustment, and simulating occlusion.

[0029] The sub-strategy combinations are used to augment the sample datasets to determine the augmented sample datasets.

[0030] The second pest and disease detection model is trained using the data-enhanced sample dataset until it reaches a convergence state. The second pest and disease detection model is then evaluated on the validation set to determine the reward value corresponding to the performance index, wherein the higher the performance index, the greater the reward value.

[0031] The reinforcement learning algorithm adjusts its own parameters based on the reward value corresponding to the performance metric to find the optimal combination of sub-policies.

[0032] Optionally, the basic network architecture of the first pest and disease detection model is the original YOLOv11n model, and the basic network architecture of the second pest and disease detection model is the improved YOLOv11n model.

[0033] A pest and disease detection device for lemon orchards in complex environments, provided for another purpose of this application, includes:

[0034] The image acquisition module is configured to respond to instructions for detecting diseases and pests in a lemon orchard under complex conditions, and acquire images of the lemon orchard to be detected containing lemon diseases and pests, wherein the lemon diseases and pests include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly and citrus psyllid.

[0035] The detection model construction module is configured to update the backbone network of the preset first pest and disease detection model to a RepViT network, introduce the RFAConv receptive field attention module into the C3k2 module of the neck network of the first pest and disease detection model to construct the C3k2_RFA module, and update the Upsampling module of the original feature pyramid P2, P3, P4 layers in the neck network to a DySample module to construct the second pest and disease detection model.

[0036] The pest and disease detection module is configured to input the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to convergence, so as to detect lemon pests and diseases in the image of the lemon orchard to be detected, thereby completing the detection of lemon orchard pests and diseases in complex environments.

[0037] An electronic device provided for another purpose of this application includes a central processing unit and a memory, the central processing unit being used to invoke and run a computer program stored in the memory to perform the steps of the lemon orchard pest and disease detection method under complex conditions described in this application.

[0038] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the method for detecting pests and diseases in lemon orchards under complex environments, which, when called by a computer, executes the steps included in the corresponding method.

[0039] Compared to existing technologies, this application addresses the problems that background noise and light intensity may interfere with the identification of pests and diseases in complex orchard environments, as well as the inability of existing models to fully capture the complex characteristics of pests and diseases and their insufficient multi-scale feature fusion capabilities. This application provides, but is not limited to, the following beneficial effects:

[0040] First, the improved YOLOv11n model in this application introduces the RepViT network as the backbone network, significantly enhancing the model's learning ability. By optimizing the training process, it eliminates the computational and memory overhead caused by skip connections, while ensuring efficient inference. During model inference, by reducing unnecessary computation, it achieves optimal accuracy with minimal latency increase, thereby improving detection performance, especially on resource-constrained devices where the performance improvement is more pronounced.

[0041] Secondly, the improved YOLOv11n model in this application effectively aggregates global information by introducing the RFAConv receptive field attention module, enabling the network to focus on the statistical characteristics within the entire receptive field and generate unique attention weights for each convolutional kernel. This approach allows the model to adaptively focus on key feature regions, significantly improving the detection capability of lemon pests and diseases, especially in complex backgrounds, particularly under background noise or occlusion conditions.

[0042] Third, this application updates the Upsample module of the original YOLOv11n model to the DySample model, resulting in a superior YOLOv11n model that performs better when handling targets with indistinct edges, irregular shapes, and multi-scale features. The DySample upsampling method improves the efficiency of feature extraction, especially in lemon pest and disease detection, helping the model to identify features at different scales. Simultaneously, it enhances the fusion capability of multi-scale features, enabling the model to better handle pest and disease targets of different sizes and shapes, thus improving the comprehensiveness and accuracy of detection.

[0043] Fourth, this application successfully achieves lightweight model deployment by adopting the RepViT network architecture, reducing the number of model parameters and weights. This enables the model to be quickly deployed and run efficiently in real-world scenarios, and is particularly valuable for resource-constrained terminal devices (such as mobile devices or monitoring equipment in farmland). The lightweight design not only reduces computational overhead but also reduces energy consumption and improves versatility.

[0044] Fifth, by introducing a data augmentation strategy based on reinforcement learning, the data augmentation process is automatically generated and optimized, reducing the workload of manually designing data augmentation strategies. This strategy significantly improves the model's generalization ability in complex contexts, enabling it to adapt to different environmental conditions and effectively cope with the complex characteristics of lemon pests and diseases and the changing environmental factors.

[0045] Sixth, the method for detecting lemon orchard pests and diseases in complex environments proposed in this application not only represents significant theoretical innovation and improvement, but also demonstrates outstanding effectiveness in practical applications. By combining a WeChat mini-program app with a server-side backend model deployment, rapid and accurate identification of lemon pests and diseases is achieved. Farmers do not need extensive technical backgrounds; they can obtain real-time pest and disease detection results through simple operations, greatly facilitating farmland management and improving agricultural production efficiency. This convenient tool not only helps farmers improve the yield and quality of lemon crops but also promotes the modernization of agriculture. By providing farmers with an efficient and accurate lemon pest and disease detection service, it not only enables timely identification of pests and diseases, improving the efficiency of pest and disease control, and reducing crop losses, but also reduces pesticide use and environmental pollution. This solution contributes to improving agricultural sustainability, promoting the development of green agriculture, and has significant social and economic value.

[0046] In summary, the lemon pest and disease detection method provided in this application incorporates several innovations in model design, particularly in optimizing the network structure, introducing a receptive field attention mechanism, and replacing the upsampling module. These innovations significantly improve the model's detection performance in complex environments. Furthermore, combined with intelligent data augmentation strategies and a lightweight deployment scheme, the model can be efficiently applied in real-world scenarios, bringing tangible convenience and economic benefits to farmers and promoting the modernization of agriculture. Attached Figure Description

[0047] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0048] Figure 1 This is a flowchart illustrating the method for detecting pests and diseases in lemon orchards under complex environments, as described in this application.

[0049] Figure 2This is a schematic diagram of the confusion matrix of the original YOLOv11n model in the embodiments of this application;

[0050] Figure 3 This is an exemplary network architecture of the RepViT network in the embodiments of this application;

[0051] Figure 4 This is an exemplary network architecture for the RFAConv receptive field attention module in the embodiments of this application;

[0052] Figure 5 This is an exemplary network architecture for the DySample module in the embodiments of this application;

[0053] Figure 6 This is an exemplary network architecture for the improved YOLOv11n model in the embodiments of this application;

[0054] Figure 7 This is a schematic diagram comparing various pests and diseases with mAP@0.5 values ​​in the original YOLOv11n model and the improved YOLOv11n model in the embodiments of this application.

[0055] Figure 8 This is a schematic diagram illustrating the effect of the WeChat mini program user interface in the embodiments of this application;

[0056] Figure 9 This is a schematic diagram of the lemon orchard pest and disease detection device under complex environment in the embodiments of this application;

[0057] Figure 10 This is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation

[0058] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.

[0059] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.

[0060] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0061] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.

[0062] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the Von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.

[0063] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.

[0064] Please see Figure 1 In one embodiment of the method for detecting pests and diseases in lemon orchards under complex environments, this application includes:

[0065] Step S10: In response to the instruction to detect diseases and pests in a lemon orchard under complex conditions, obtain an image of the lemon orchard to be detected containing lemon diseases and pests, wherein the lemon diseases and pests include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly and citrus psyllid.

[0066] The lemon orchard pest and disease detection system in the terminal device can respond to the instruction to detect pests and diseases in lemon orchards under complex environments and acquire images of lemon orchards to be detected containing lemon pests and diseases. The lemon pests and diseases include one or more of the following: sooty mold, anthracnose, grease spot, leaf miner, whitefly, and citrus psyllid.

[0067] In some embodiments, the acquired images are classified and labeled, including: labeling the images using LabelImg labeling software, constructing a pest and disease database containing data such as the pest and disease categories detected by the target, labeling information, and dates.

[0068] Specifically, the process of constructing an image database of lemon orchards containing images of lemon diseases and pests (constructing an image dataset of real-world scenes with complex backgrounds) is as follows:

[0069] Images of lemon orchards containing lemon pests and diseases were acquired using high-definition cameras in real-world settings. These pests and diseases included sooty mold, anthracnose, grease spot, leaf miners, whiteflies, and citrus psyllids. The acquired images were classified and labeled using Labelimg software, and the labeled data files were stored in an image database. The lemon pest dataset was divided into training, validation, and test sets in a 7:2:1 ratio for model training and validation.

[0070] In a further embodiment, the generalization ability of real-world applications is enhanced by comparing manual operation and reinforcement learning data augmentation. Data augmentation includes shear transformation, translation, rotation, contrast adjustment, histogram equalization, posterization (reducing the number of bits per color channel of the image), brightness adjustment, sharpness adjustment (making object edges clearer), and cutout (simulating occlusion phenomena to simulate real orchard scenes). Manual operation refers to researchers selecting data augmentation strategies and setting parameters based on experience.

[0071] The core idea of ​​reinforcement learning is to automatically discover optimal data augmentation strategies through search and learning, rather than relying on manual design. AutoAugment is a technique for automatically searching for optimal data augmentation strategies. It automatically explores different combinations of data augmentation operations within a predefined operation space to find strategies that maximize the model's performance on the validation set. This operation space includes various common data augmentation operations such as rotation, flipping, cropping, and color adjustment, each with corresponding parameter settings. AutoAugment represents data augmentation strategies as a combination of a series of sub-strategies, each containing specific data augmentation operations, their application probabilities, and parameter settings. This automates the data augmentation process, significantly reducing the time and effort required for manual design and improving the model's generalization ability on unseen data.

[0072] Specifically, to achieve the generalization ability of the model, data augmentation techniques are needed to expand the dataset. To simulate real complex scenarios, the generalization ability of real-world applications is enhanced by comparing manual operation and reinforcement learning data augmentation methods. Manual operation data augmentation methods include random Gaussian noise, random rotation, mirroring, brightness adjustment, cropping, and cutout (simulating occlusion).

[0073] The generalization ability of the model in real-world applications can be compared with the detection results of the trained model on the test set. The network model is trained on the lemon pest and disease dataset using the YOLOv11n model suitable for mobile devices. The data augmentation used in reinforcement learning includes: shear transformation, translation, rotation, contrast adjustment, histogram equalization, posterization (reducing the number of bits in each color channel of the image), brightness adjustment, sharpness adjustment (making the edges of objects clearer), and cutout (simulating occlusion).

[0074] The test results obtained by training the improved yolov11n model with test sets obtained through two different data augmentation methods, manual selection and reinforcement learning, are shown in Table 1.

[0075] Table 1 shows the test results obtained by training the improved YOLOv11n model using test sets obtained from two different data augmentation methods.

[0076]

[0077] As can be seen, the recall rate of each disease and pest was improved after adopting the reinforcement learning strategy. Therefore, in the subsequent training of the improved yolov11n model, reinforcement learning was given priority for data augmentation.

[0078] Step S20: Update the backbone network of the preset first pest and disease detection model to a RepViT network. Introduce the RFAConv receptive field attention module into the C3k2 module of the neck network of the first pest and disease detection model to construct the C3k2_RFA module. Update the Upsampling module of the original feature pyramid P2, P3, and P4 layers in the neck network to a DySample module to construct the second pest and disease detection model.

[0079] After acquiring images of lemon orchards containing lemon pests and diseases, the backbone network of the preset first pest and disease detection model is updated to a RepViT network. An RFAConv receptive field attention module is introduced into the C3k2 module of the neck network of the first pest and disease detection model to construct a C3k2_RFA module. The Upsampling module of the original feature pyramid P2, P3, and P4 layers in the neck network is updated to a DySample module to construct a second pest and disease detection model. The basic network architecture of the first pest and disease detection model is the original YOLOv11n model, and the basic network architecture of the second pest and disease detection model is an improved YOLOv11n model, which is named the YOLOv11-RRM model.

[0080] In some embodiments, please refer to Figure 2 The original YOLOv11n model is a version of the YOLOv11 series developed by the Ultralytics team. It is designed for efficient object detection, especially suitable for resource-constrained environments.

[0081] In some embodiments, the backbone network of the preset original YOLOv11n model is updated to a RepViT network. In the ViTBlock module of the RepViT network, a 1×1 extended convolutional layer is followed by a 3×3 depthwise separable convolutional layer for spatial information fusion. The depthwise separable convolutional layer is optimized using structural reparameterization technology. A multi-branch topology is introduced through structural reorganization, and the expansion ratio in the convolutional layer is adjusted. The expansion ratio is set to 2 in the channel mixer of all stages.

[0082] In the macroscopic design optimization, the RepViT network uses two 3×3 convolutional layers with a stride of 2 as the starting module. The number of filters in the first convolutional layer is set to 24, and the number of filters in the second convolutional layer is set to 48. A depthwise separable convolutional layer with a stride of 2 and a 1×1 point convolutional layer are used to perform spatial downsampling and channel dimension adjustment, respectively. A RepViT module is added before the downsampling layer to further deepen the network depth, and a feedforward network module is placed after the 1×1 convolutional layer to remember more potential information. In the microscopic design adjustment, all modules preferentially use 3×3 convolutional kernels, and the compressed excitation layer is alternately deployed in the odd-numbered blocks of each stage.

[0083] Specifically, in the ViTBlock module of the RepViT network, a 3×3 depthwise separable (DW) convolution is equipped after a 1×1 expanded convolution for spatial information fusion. At the same time, the structural reparameterization technique, which is widely used during model training, is used to optimize the DW layer to enhance the model's learning ability during training. Through structural reparameterization (SR), a multi-branch topology is introduced, and the expansion ratio in the convolutional layers is adjusted. The expansion ratio is set to 2 in the channel mixer of all stages to reduce parameter redundancy and latency. At the same time, the macro-architecture of the network is optimized, including the design of early convolutional layers, deeper downsampling layers, simplified classifiers, and adjustments to the overall stage ratio. The micro-design is also adjusted, including the selection of convolutional kernel size, and the optimal placement of 3×3 convolutions and compressed excitation (SE) layers is prioritized in all modules.

[0084] More specifically, in the backbone of the original YOLOv11n model, the backbone network is updated to a RepViT network. An exemplary network structure of the RepViT network is as follows: Figure 3 As shown; the main idea is to apply the design principles of the lightweight visual transformer (ViT) to the traditional lightweight convolutional neural network (CNN). The main improvement mechanisms include:

[0085] In the ViTBlock module, 1×1 expanded convolutions are followed by 3×3 depthwise separable (DW) convolutions for spatial information fusion. The SE layer in the RepViT module is optional. The DW layer is optimized using structural reparameterization (SR), a technique widely used during model training, to enhance the model's learning ability during training. This eliminates the computational and memory overhead of skip connections during inference, improving model performance. Structural reparameterization (SR) introduces a multi-branch topology to further improve training performance. The expansion ratio in the convolutional layers is adjusted, with a ratio of 2 set in the channel mixer at all stages to reduce parameter redundancy and latency. This achieves efficiency gains through a smaller expansion ratio. The number of channels is doubled after each stage, ultimately reaching 48, 96, 192, and 384 channels for each stage, increasing network width to enhance model performance. In the macro-design optimization, the macro-architecture of the network is optimized, including: the RepViT network uses two 3×3 convolutions with a stride of 2 as the starting module, with the first convolution layer having 24 filters and the second convolution layer having 48 filters. Spatial downsampling and channel dimension adjustment are performed using depthwise separable (DW) convolutional layers with a stride of 2 and 1×1 point convolutional layers, respectively. A RepViT module is added before the downsampling layer to further deepen the network, and a feedforward network (FFN) module is placed after the 1×1 convolution to remember more latent information. A simplified classifier (global average pooling layer + linear layer) and an overall stage ratio are also adjusted. In the micro-design adjustment, optimization is performed at the micro-architecture level, including the selection of convolutional kernel size, prioritizing the optimal placement of 3×3 convolutions and compression excitation (SE) layers in all modules, and designing a cross-block SE layer usage strategy—alternating the deployment of compression excitation layers in odd-numbered blocks (1, 3, 5, etc.) of each stage to achieve maximum accuracy gain with minimal latency increase.

[0086] In a further embodiment, an RFAConv receptive field attention module is introduced into the C3k2 module of the original YOLOv11n model's neck network to construct the C3k2_RFA module, wherein an exemplary network structure of the RFAConv receptive field attention module is as follows: Figure 4 As shown, its main idea is to combine spatial attention mechanism with convolution operation to improve the performance of convolutional neural network.

[0087] RFA defines a design for convolutional kernels. For each receptive field feature, it mainly extracts features using the Unfold method or the Group Conv method provided by PyTorch. The Unfold method is a parameter-free method, but it is relatively slow; the Group Conv method is a fast method for extracting spatial features of the receptive field.

[0088] This embodiment takes the Group Conv method as an example and does not constitute any limitation on this application. The spatial features of each receptive field in the pest and disease feature map are extracted by the Group Conv method, and the global information is aggregated by using average pooling (AvgPool). This helps to capture the spatial features within the entire receptive field while minimizing computational overhead and the number of parameters.

[0089] After average pooling, 1x1 group convolution operations are applied for further processing to interact with information and enhance network performance. Then, the output of the 1x1 group convolution is normalized by the Softmax function to generate importance weights for the spatial features of each receptive field, emphasizing the importance of each feature in the spatial features of the receptive field.

[0090] Finally, the attention weights generated by Softmax are combined with the spatial features of the receptive fields after group convolution. Each convolution kernel will obtain different weights according to the spatial features of its corresponding receptive field, enabling the network model to better identify and highlight the key regions of objects.

[0091] The calculation formula for the RFAConv receptive field attention module can be expressed as:

[0092]

[0093] Where g represents a grouped convolution of size i×i, where g 1×1 G represents a 1×1 grouped convolution. k×k represents a k×k grouped convolution, where k represents the kernel size; Softmax represents the normalization function used to generate importance weights for each feature within the receptive field; AvgPool(X) represents the weights of the input features. Figure X Average pooling is performed to aggregate global information; Norm represents normalization; ReLU represents the activation function, introducing non-linear features; X represents the input feature map; F is obtained by multiplying the attention map A with the spatial features F of the transformed receptive field; A rf Attention weights representing the spatial features of the receptive field are generated by the Softmax function; F rf This represents the spatial features of the receptive field extracted through grouped convolution.

[0094] In a further embodiment, the Upsample module of the original feature pyramid layers P2, P3, and P4 in the neck of the original YOLOv11n model is updated to a DySample module. DySample mainly uses point resampling instead of the traditional kernel-based method. It mainly generates a set of candidate coordinates through a lightweight network (such as a fully connected layer or a small convolutional network) to represent the regions that need to be enhanced in the high-resolution space. Its weight generation is achieved by assigning a weight to each coordinate to reflect the importance of the position. Therefore, during upsampling, contextual features can be extracted from the low-resolution input. Based on the coordinates provided by the sampler, the feature values ​​of the corresponding positions are extracted from the feature map and combined with the weights for weighted fusion to achieve the aggregation of sampling points. The aggregated features are further refined through a neural network to generate the final high-resolution result image.

[0095] The DySample module is primarily based on dynamic upsampling of the sampled data, such as... Figure 5 As shown; the input feature (x) is used to create a sampling set through a sampling point generator. Then, the input features are resampled using the grid_sample function to obtain the upsampled features (x'), which can be represented as:

[0096] X′=grid_sample(X,S),

[0097] Where X represents the input enhanced pest and disease feature map; S represents the sampling set; grid_sample represents the grid-based sampling function, which uses bilinear interpolation to calculate the pixel value of the sampling point. grid_sample calculates the accurate sampling result based on the surrounding pixel values. If the coordinates of the sampling point fall between pixels, grid_sample will perform a weighted average based on the values ​​of four adjacent pixels to generate the interpolation result, represented as:

[0098] value=(1-Δx)(1-Δy)·V i,j +Δx(1-Δy)·V i+1,j +(1-Δx)Δy·V i,j+1 +ΔxΔy·V i+1,j+1 ,

[0099] Where Δx and Δy are the decimal parts of the coordinates, V i,j V i+1,j V i,j+1 V i+1,j+1 It is the pixel value of the four adjacent pixels around the sampling point.

[0100] In the sampling point generator, a static range factor is used to generate the first offset by combining a linear layer and pixel shuffle technology with a fixed first range factor. The static range factor can be 0.25. The offset multiplied by a coefficient of 0.25 satisfies the theoretical boundary condition between overlapping and non-overlapping areas, i.e.:

[0101]

[0102] Where O represents the first offset, used to adjust the sampling point position; linear(X) represents the input feature... Figure X Perform linear layer (fully connected layer) processing; 0.25 is the static range factor.

[0103] In the sampling point generator, in addition to the linear layer and pixel shuffling, a dynamic range factor is introduced. First, a second range factor is generated using the sigmoid function, and then it is used to adjust the first offset. To determine the second offset, it is then added to the original grid position (g) to obtain the sample set. To improve the flexibility of the offset, a point-like "dynamic range factor" is generated by linearly projecting the input features. Using the sigmoid function and a static coefficient of 0.5, the dynamic range factor ranges from [0, 0.5], centered at 0.25.

[0104]

[0105] Where linear1(X) represents the first linear layer, used to generate the dynamic range factor; linear2(X) represents the second linear layer, used to generate the base offset; 0.5 represents the static coefficient, which controls the value range of the dynamic range factor [0,0.5].

[0106] In some embodiments, the step of training a second pest and disease detection model includes:

[0107] Step S201: Obtain a sample dataset, wherein the sample dataset includes multiple training samples and their corresponding supervision labels, the training samples represent a single lemon orchard image, and the supervision labels represent the annotation information of lemon diseases and pests in the lemon orchard image;

[0108] Step S202: Randomly generate a data augmentation strategy combination, wherein the data augmentation strategy combination includes multiple sub-strategy combinations, including shear transformation, translation, rotation, contrast adjustment, histogram equalization, reducing the number of bits in each color channel of the image, brightness adjustment, sharpness adjustment, and simulating occlusion.

[0109] Step S203: Use the sub-strategy combination to perform data augmentation on the sample dataset respectively, so as to determine the data-augmented sample dataset;

[0110] Step S204: Train the second pest and disease detection model using the data-enhanced sample dataset until the second pest and disease detection model is trained to convergence. Evaluate the second pest and disease detection model on the validation set to determine the reward value corresponding to the performance index, wherein the higher the performance index, the greater the reward value.

[0111] Step S205: Adjust the parameters of the system based on the reward value corresponding to the performance index using a reinforcement learning algorithm to find the optimal combination of sub-policies.

[0112] Specifically, the improved model is trained, and its detection accuracy is evaluated using precision (P), recall (R), accuracy, and mean precision (mAP@0.5). The lightweighting effect of the model is evaluated using model size and model computational cost (FLOPs). The specific formulas for calculating precision, recall, and accuracy include:

[0113]

[0114]

[0115] In this model, TP represents a correct positive, indicating that the prediction was positive and the actual result is a positive example; TN represents a correct negative, indicating that the prediction was negative and the actual result is a negative example; FP represents a false positive, indicating that the prediction was positive but the actual result is a negative example; FN represents a false negative, indicating that the prediction was negative but the actual result is a positive example. Precision represents the ratio of correctly detected samples to the total number of samples detected; Recall, also known as recall, represents the ratio of correctly detected samples to the total number of samples in the test set; Accuracy represents the percentage of correctly detected samples out of the total number of samples. The mean precision (mAP@0.5) reflects the overall performance of the model in terms of precision and recall across different categories. The model weights refer to the size or number of learnable parameters in the model. Model computational cost refers to the number of arithmetic operations (such as addition and multiplication) required for a model to perform one forward propagation (from input data through the model to obtain the output result). It is mainly used to measure the computational cost of the model during the inference phase.

[0116] An exemplary network structure of the improved YOLOv11n model is as follows: Figure 6As shown, the improved YOLOv11n model is named the YOLOv11-RRM model. The YOLOv11-RRM model is trained on the lemon pest dataset. The specific process is as follows:

[0117] In data augmentation, reinforcement learning uses autoaugment. During training, SGD is selected as the optimizer for the model. The initial learning rate of the optimizer is set to 0.01, the weight decay is 0.0005, the model is trained 200 times, the batch size is 16, the bounding box loss gain is set to 7.5, and the distance-IoU loss (DFL) gain is set to 1.5 to regress the bounding boxes more accurately.

[0118] The comparison results between the improved YOLOv11n model and the original YOLOv11n model are shown in Table 2:

[0119] Table 2 shows the comparison results between the improved YOLOv11n model and the original YOLOv11n model.

[0120]

[0121] As shown in Table 2, compared with the original YOLOv11n model, YOLOv11-RRM improves mAP@0.5 by 3.8% and recall by 5.1% while maintaining a similar number of parameters. It effectively eliminates false positives and false negatives, reduces size by 10%, and slightly increases FLOPs. The improved YOLOv11n model and the original YOLOv11n model show a comparison of mAP@0.5 for various insect pests. Figure 7 As shown in the results, the accuracy of identifying small target pests has been significantly improved. Based on the original model, the improved model simultaneously achieves a reduction in the number of parameters (Size reduced from 5.5 to 5.0M) and an improvement in the detection accuracy of lemon pests and diseases.

[0122] Step S30: Input the image of the lemon orchard to be detected into the second pest and disease detection model that has been trained to convergence, so as to detect lemon pests and diseases in the image of the lemon orchard to be detected, so as to complete the detection of lemon orchard pests and diseases in complex environments.

[0123] Once the constructed second pest and disease detection model has been trained to a convergent state, it can be put into production use to detect lemon pests and diseases in lemon orchard images. The lemon orchard image to be detected is input into the second pest and disease detection model that has been trained to a convergent state to detect the lemon pests and diseases in the lemon orchard image to complete the detection of lemon pests and diseases in lemon orchards under complex environments.

[0124] In some embodiments, the improved YOLOv11n model is deployed to the backend of a lemon pest and disease detection server, and the front-end WeChat mini-program visualization interface is as follows: Figure 8 As shown, to facilitate user operation, the app displays common lemon pests and diseases. Clicking on a model allows users to view detailed information and control methods. Clicking "Recognize" leads to a model selection page where users can select the appropriate model and take a real-time photo or import an image for recognition. The app then displays the recognition results, the types and numbers of pests and diseases detected. Users can then use the information displayed in the app to control specific pests and diseases, including the use of corresponding pesticides and the pesticide dosage determined by the number of pests and diseases detected. This greatly facilitates the management of pests and diseases in lemon orchards for fruit growers.

[0125] In some embodiments, the step of inputting the lemon orchard image to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the lemon orchard image includes:

[0126] Step S301: In the C3k2_RFA module, the spatial features of each receptive field in the pest and disease feature map are extracted by the group convolution method. The pest and disease feature map includes pest and disease features, which characterize the texture of lesions or the morphology of insect bodies.

[0127] Step S302: Average pooling is used to aggregate the global information of the spatial features of each receptive field, and the information is interacted through 1×1 convolution operations to learn the correlation between different channels;

[0128] Step S303: Normalize the output of the 1×1 group of convolutions using the softmax function to generate the importance weights of the spatial features of each receptive field;

[0129] Step S304: Combine the importance weights with the spatial features of the original receptive field extracted by grouped convolution to output an enhanced pest and disease feature map for subsequent pest and disease detection.

[0130] In a further embodiment, the step of inputting the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the image of the lemon orchard to be detected includes:

[0131] Step S3001: Input the enhanced pest and disease feature map obtained after processing by the C3k2_RFA module into the DySample module. In the sampling point generator, the static range factor is used to generate the first offset by combining the linear layer and pixel shuffling technology with the fixed first range factor.

[0132] Step S3002: In the sampling point generator, a second range factor is first generated using a dynamic range factor, and then the first offset is adjusted using the second range factor to determine the second offset. The second offset is added to the original grid position to generate a sampling set. The second range factor is generated by the Sigmoid function.

[0133] Step S3003: Based on the generated sampling set, the enhanced pest and disease feature map is resampled using bilinear interpolation to generate an upsampled feature map for subsequent pest and disease detection.

[0134] As can be seen from the above embodiments, compared with the prior art, this application addresses the problems that background noise and light intensity may interfere with the identification of pests and diseases in complex orchard environments, and that existing models are unable to fully capture the complex manifestations of pests and diseases, and have insufficient multi-scale feature fusion capabilities. This application has, but is not limited to, the following beneficial effects:

[0135] First, the improved YOLOv11n model in this application introduces the RepViT network as the backbone network, significantly enhancing the model's learning ability. By optimizing the training process, it eliminates the computational and memory overhead caused by skip connections, while ensuring efficient inference. During model inference, by reducing unnecessary computation, it achieves optimal accuracy with minimal latency increase, thereby improving detection performance, especially on resource-constrained devices where the performance improvement is more pronounced.

[0136] Secondly, the improved YOLOv11n model in this application effectively aggregates global information by introducing the RFAConv receptive field attention module, enabling the network to focus on the statistical characteristics within the entire receptive field and generate unique attention weights for each convolutional kernel. This approach allows the model to adaptively focus on key feature regions, significantly improving the detection capability of lemon pests and diseases, especially in complex backgrounds, particularly under background noise or occlusion conditions.

[0137] Third, this application updates the Upsample module of the original YOLOv11n model to the DySample model, resulting in a superior YOLOv11n model that performs better when handling targets with indistinct edges, irregular shapes, and multi-scale features. The DySample upsampling method improves the efficiency of feature extraction, especially in lemon pest and disease detection, helping the model to identify features at different scales. Simultaneously, it enhances the fusion capability of multi-scale features, enabling the model to better handle pest and disease targets of different sizes and shapes, thus improving the comprehensiveness and accuracy of detection.

[0138] Fourth, this application successfully achieves lightweight model deployment by adopting the RepViT network architecture, reducing the number of model parameters and weights. This enables the model to be quickly deployed and run efficiently in real-world scenarios, and is particularly valuable for resource-constrained terminal devices (such as mobile devices or monitoring equipment in farmland). The lightweight design not only reduces computational overhead but also reduces energy consumption and improves versatility.

[0139] Fifth, by introducing a data augmentation strategy based on reinforcement learning, the data augmentation process is automatically generated and optimized, reducing the workload of manually designing data augmentation strategies. This strategy significantly improves the model's generalization ability in complex contexts, enabling it to adapt to different environmental conditions and effectively cope with the complex characteristics of lemon pests and diseases and the changing environmental factors.

[0140] Sixth, the method for detecting lemon orchard pests and diseases in complex environments proposed in this application not only represents significant theoretical innovation and improvement, but also demonstrates outstanding effectiveness in practical applications. By combining a WeChat mini-program app with a server-side backend model deployment, rapid and accurate identification of lemon pests and diseases is achieved. Farmers do not need extensive technical backgrounds; they can obtain real-time pest and disease detection results through simple operations, greatly facilitating farmland management and improving agricultural production efficiency. This convenient tool not only helps farmers improve the yield and quality of lemon crops but also promotes the modernization of agriculture. By providing farmers with an efficient and accurate lemon pest and disease detection service, it not only enables timely identification of pests and diseases, improving the efficiency of pest and disease control, and reducing crop losses, but also reduces pesticide use and environmental pollution. This solution contributes to improving agricultural sustainability, promoting the development of green agriculture, and has significant social and economic value.

[0141] In summary, the lemon pest and disease detection method provided in this application incorporates several innovations in model design, particularly in optimizing the network structure, introducing a receptive field attention mechanism, and replacing the upsampling module. These innovations significantly improve the model's detection performance in complex environments. Furthermore, combined with intelligent data augmentation strategies and a lightweight deployment scheme, the model can be efficiently applied in real-world scenarios, bringing tangible convenience and economic benefits to farmers and promoting the modernization of agriculture.

[0142] Please see Figure 9This application provides a device for detecting pests and diseases in lemon orchards under complex environments, comprising an image acquisition module 1100, a detection model construction module 1200, and a pest and disease detection module 1300, in response to an instruction to detect pests and diseases in a lemon orchard under complex environments. The image acquisition module 1100 is configured to acquire an image of the lemon orchard containing pests and diseases to be detected, wherein the pests and diseases include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly, and citrus psyllid. The detection model construction module 1200 is configured to update the backbone network of a preset first pest and disease detection model to a RepViT network, and introduce RF into the C3k2 module on the neck network of the first pest and disease detection model. The AConv receptive field attention module is used to construct the C3k2_RFA module. The Upsampling module of the original feature pyramid P2, P3, and P4 layers in the neck network is updated to the DySample module to construct the second pest and disease detection model. The pest and disease detection module 1300 is configured to input the lemon orchard image to be detected into the second pest and disease detection model that has been trained to convergence state in order to detect lemon pests and diseases in the lemon orchard image to complete the detection of lemon orchard pests and diseases in complex environments.

[0143] Based on any embodiment of this application, please refer to Figure 10 Another embodiment of this application also provides an electronic device, which can be implemented by a computer device, such as... Figure 10 The diagram shows the internal structure of a computer device. This computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store control information sequences. When the processor executes the computer-readable instructions, it enables the processor to implement a method for detecting pests and diseases in lemon orchards under complex environments. The processor provides computational and control capabilities, supporting the operation of the entire computer device. The memory stores computer-readable instructions, which, when executed by the processor, enable the processor to execute the method for detecting pests and diseases in lemon orchards under complex environments as described in this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0144] In this embodiment, the processor is used to execute... Figure 9 The memory stores the specific functions of each module, and stores the program code and various data required to execute the above modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules in the lemon orchard pest and disease detection device under complex environment of this application, and the server can call the server's program code and data to execute the functions of all modules.

[0145] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the lemon orchard pest and disease detection method under complex conditions described in any embodiment of this application.

[0146] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the method for detecting pests and diseases in a lemon orchard under complex conditions as described in any embodiment of this application.

[0147] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0148] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for detecting pests and diseases in lemon orchards under complex environments, characterized in that, include: In response to an instruction to detect diseases and pests in a lemon orchard under complex conditions, the system acquires an image of the lemon orchard to be detected containing lemon diseases and pests, wherein the lemon diseases and pests include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly, and citrus psyllid. The backbone network of the preset first pest and disease detection model is updated to a RepVi T network. An RFAConv receptive field attention module is introduced into the C3k2 module of the neck network of the first pest and disease detection model to construct a C3k2_RFA module. The Upsampling module of the original feature pyramid P2, P3, and P4 layers in the neck network is updated to a DySample module to construct a second pest and disease detection model. The image of the lemon orchard to be detected is input into the second pest and disease detection model that has been trained to convergence, so as to detect lemon pests and diseases in the image of the lemon orchard to complete the detection of lemon pests and diseases in complex environments.

2. The method for detecting pests and diseases in lemon orchards under complex environments according to claim 1, characterized in that, In the ViTBlock module of the RepViT network, a 1×1 extended convolutional layer is followed by a 3×3 depth-separable convolutional layer for spatial information fusion. The depth-separable convolutional layer is optimized using structural reparameterization technology. A multi-branch topology is introduced through structural reorganization, and the expansion ratio in the convolutional layer is adjusted. The expansion ratio is set to 2 in the channel mixer of all stages.

3. The method for detecting pests and diseases in lemon orchards under complex environments according to claim 2, characterized in that, In the macroscopic design optimization, the RepVi T network uses two 3×3 convolutional layers with a stride of 2 as the starting module. The number of filters in the first convolutional layer is set to 24, and the number of filters in the second convolutional layer is set to 48. A depthwise separable convolutional layer with a stride of 2 and a 1×1 point convolutional layer are used to perform spatial downsampling and channel dimension adjustment, respectively. A RepVi T module is added before the downsampling layer to further deepen the network depth, and a feedforward network module is placed after the 1×1 convolutional layer to remember more potential information. In the micro-design adjustments, all modules preferentially use 3×3 size convolution kernels, and the compressed excitation layers are alternately deployed in the odd number of blocks in each stage.

4. The method for detecting pests and diseases in lemon orchards under complex environments according to claim 1, characterized in that, The step of inputting the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the image of the lemon orchard to be detected includes: In the C3k2_RFA module, the spatial features of each receptive field in the pest and disease feature map are extracted by the grouped convolution method. The pest and disease feature map includes pest and disease features, which characterize the texture of lesions or the morphology of insect bodies. Average pooling is used to aggregate global information of spatial features of each receptive field, and the information is interacted through 1×1 convolution operations to learn the correlation between different channels; The output of the 1×1 convolution is normalized using the softmax function to generate importance weights for the spatial features of each receptive field. The importance weights are combined with the spatial features of the original receptive field extracted by grouped convolution to output an enhanced pest and disease feature map for subsequent pest and disease detection.

5. The method for detecting pests and diseases in lemon orchards under complex environments according to claim 4, characterized in that, The step of inputting the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to a convergent state to detect lemon pests and diseases in the image of the lemon orchard to be detected includes: The enhanced pest and disease feature map obtained after processing by the C3k2_RFA module is input into the DySample module. In the sampling point generator, the static range factor is used to generate the first offset by combining the linear layer and pixel shuffling technology with the fixed first range factor. In the sampling point generator, a second range factor is first generated using a dynamic range factor, and then the first offset is adjusted using the second range factor to determine the second offset. The second offset is added to the original grid position to generate a sampling set. The second range factor is generated by the Sigmoid function. Based on the generated sample set, bilinear interpolation is used to resample the enhanced pest and disease feature map to generate an upsampled feature map for subsequent pest and disease detection.

6. The method for detecting pests and diseases in lemon orchards under complex environments according to claim 1, characterized in that, The steps for training a second pest and disease detection model include: Obtain a sample dataset, wherein the sample dataset includes multiple training samples and their corresponding supervision labels, the training samples represent a single lemon orchard image, and the supervision labels represent the annotation information of lemon diseases and pests in the lemon orchard image; A data augmentation strategy combination is randomly generated, wherein the data augmentation strategy combination includes multiple sub-strategy combinations, including shear transformation, translation, rotation, contrast adjustment, histogram equalization, reducing the number of bits in each color channel of the image, brightness adjustment, sharpness adjustment, and simulating occlusion. The sub-strategy combinations are used to augment the sample datasets to determine the augmented sample datasets. The second pest and disease detection model is trained using the data-enhanced sample dataset until it reaches a convergence state. The second pest and disease detection model is then evaluated on the validation set to determine the reward value corresponding to the performance index, wherein the higher the performance index, the greater the reward value. The reinforcement learning algorithm adjusts its own parameters based on the reward value corresponding to the performance metric to find the optimal combination of sub-policies.

7. The method for detecting pests and diseases in lemon orchards under complex environments according to any one of claims 1 to 6, characterized in that, The basic network architecture of the first pest and disease detection model is the original YOLOv11 n model, while the basic network architecture of the second pest and disease detection model is the improved YOLOv11 n model.

8. A pest and disease detection device for lemon orchards under complex environments, characterized in that, include: The image acquisition module is configured to respond to instructions for detecting diseases and pests in a lemon orchard under complex conditions, and acquire images of the lemon orchard to be detected containing lemon diseases and pests, wherein the lemon diseases and pests include one or more of sooty mold, anthracnose, grease spot, leaf miner, whitefly and citrus psyllid. The detection model construction module is configured to update the backbone network of the preset first pest and disease detection model to a RepVi T network, introduce the RFAConv receptive field attention module into the C3k2 module of the neck network of the first pest and disease detection model to construct the C3k2_RFA module, and update the Upsampling module of the original feature pyramid P2, P3, P4 layers in the neck network to a DySample module to construct the second pest and disease detection model. The pest and disease detection module is configured to input the image of the lemon orchard to be detected into a second pest and disease detection model that has been trained to convergence, so as to detect lemon pests and diseases in the image of the lemon orchard to be detected, thereby completing the detection of lemon orchard pests and diseases in complex environments.

9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.