Tea tree aphid small target detection method

Through the improved Vision Transformer model, combined with global feature modeling and multi-scale feature fusion, the problem of low detection accuracy of small aphids in complex tea garden environments is solved, and the effects of high precision, high recall and real-time monitoring are achieved.

CN120125970APending Publication Date: 2025-06-10SHANDONG ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510223591.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing tea tree aphid detection method has low detection accuracy in complex tea garden backgrounds and is poorly robust to noise and interference, making it difficult to cope with the changing environmental conditions of tea gardens.

Method used

The improved Vision Transformer model combines global feature modeling and multi-scale feature fusion to enhance the detection ability of small targets of aphids through multi-head self-attention mechanism and dynamic attention mechanism, and uses transfer learning and data augmentation techniques to improve the generalization ability of the model.

Benefits of technology

It significantly improves the detection accuracy and recall rate of small target aphids, improves the robustness and computing efficiency of the model in complex tea garden environments, and realizes real-time monitoring and efficient detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125970A_ABST
    Figure CN120125970A_ABST
Patent Text Reader

Abstract

The invention discloses a tea tree aphid small target detection method. The method comprises the following steps: S1, collecting tea garden diversified image data and carrying out preprocessing; s2, constructing a detection model based on Vision Transform, wherein the detection model comprises construction of a feature coding module and design of a small target enhancement module; s3, through multi-scale fusion and a dynamic attention mechanism, enhancing the capability of detecting small targets of aphids; s4, carrying out optimization training on the model by utilizing transfer learning; and S5, deploying the model aphid detection system trained in the step S4 on an unmanned aerial vehicle or edge equipment, and monitoring aphid diseases and insect pests of the tea garden in real time. According to the method, by combining multi-scale feature extraction, an attention mechanism and the Vision Transform model, the problem of low aphid small target detection precision in the tea garden environment is solved, the method has the advantages of high precision, high recall rate, high calculation efficiency, real-time monitoring capability and the like, and powerful technical support is provided for intelligent prevention and control of tea garden diseases and insect pests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and agricultural informatization, and particularly relates to a method for detecting small targets of tea tree aphids. Background Art

[0002] Existing tea tree aphid detection methods mainly rely on traditional image processing techniques or convolutional neural networks (CNNs). However, aphid targets are small in size, highly camouflaged, and sparsely distributed. Traditional methods have low detection accuracy in complex tea garden backgrounds. Traditional image processing techniques usually rely on manually designed feature extraction methods, such as color, texture, and shape features. These methods may perform well in simple backgrounds, but in complex tea garden environments, due to factors such as lighting changes, cluttered backgrounds, and similar colors between aphids and leaves, the detection effect is often poor. In addition, traditional methods have poor robustness to noise and interference and are difficult to cope with the changing environmental conditions in tea gardens.

[0003] In recent years, the Vision Transformer model has performed excellently in the field of computer vision due to its powerful global feature modeling ability. It can capture long-range dependencies in images through the self-attention mechanism, thereby better modeling global context information. However, directly applying it to the detection of small aphid targets still faces challenges, such as limited resolution, loss of small target features, and the need for a large amount of labeled data.

[0004] Therefore, in view of the special requirements for detecting small targets of tea tree aphids, a targeted improved Vision Transformer model is proposed for the efficient and accurate detection of small targets of tea tree aphids. This model combines the global feature modeling ability of the model and optimization strategies for small targets, such as multi-scale feature fusion, high-resolution feature retention, and data augmentation techniques, to improve the detection accuracy and robustness in complex tea garden backgrounds. By introducing a multi-scale feature fusion module, the model can capture the features of aphid targets at different scales, thereby enhancing the detection ability for small targets. At the same time, through the high-resolution feature retention strategy, the model can retain more detailed information when processing high-resolution images and reduce the loss of small target features. In addition, the application of data augmentation techniques can effectively expand the training data set and improve the generalization ability of the model.

[0005] In summary, in response to the challenges of detecting small targets of tea tree aphids, the improved Vision Transformer model is expected to achieve efficient and accurate aphid detection in complex tea garden backgrounds by combining global feature modeling and multi-scale feature fusion strategies, providing strong technical support for tea garden pest control. Summary of the Invention

[0006] The present invention aims to solve the above problems and provides a method for detecting small targets of tea tree aphids. The method includes the following steps: S1: Collect diverse image data of the tea garden and perform preprocessing; S2: Construct a detection model based on Vision Transformer, including constructing a feature encoding module and designing a small target enhancement module; S3: Enhance the detection ability for small aphid targets through multi-scale fusion and dynamic attention mechanism; S4: Optimize and train the model using transfer learning; S5: Deploy the aphid detection system of the model trained in S4 on a drone or edge device for real-time monitoring of tea garden aphid pests and diseases.

[0007] Based on the above technical solution, in step S1, a high-resolution camera and / or drone is used to collect image data in the tea garden; the preprocessing is to perform enhancement processing on the data, including random cropping, rotation, and brightness adjustment to generate diverse training samples; and the image is normalized and standardized.

[0008] Based on the above technical solution, in step S2, the feature encoding module uses a hierarchical structure and combines a multi-scale fusion network to enhance the perception ability of small target features.

[0009] Based on the above technical solution, in step S2, the small target enhancement module adopts a specific multi-head self-attention mechanism to enhance the detailed features of aphid targets.

[0010] Based on the above technical solution, in step S3, the calculation formula of the standard attention mechanism is as follows:

[0011] where: Q, K, and V are the query vector (Query), key vector (Key), and value vector (Value) respectively; d k is the dimension of the key vector, used for scaling to prevent the inner product value from being too large; The dynamic attention mechanism introduces additional learnable parameters and / or dynamic adjustment modules on the basis of standard attention calculation to enhance the attention of regions.

[0012] Based on the above technical solution, the implementation method of the dynamic attention in Vision Transformer is as follows: Multi-head dynamic attention calculation: For the h-th head, the dynamic weight generation formula is:

[0013] Among them, σ is the activation function, and ReLU is the ReLU activation function. xi represents the i th input vector. b 1 (h) is the bias term of the first linear transformation. b 2 (h) is the bias term of the second linear transformation. The updated multi-head dynamic attention is:

[0014] Among them, the Softmax function is used to convert the input vector into a probability distribution. Q h is the query matrix of the h th attention head, K h is the key matrix of the hth attention head, ⊙ is the element-wise multiplication, and V h is the value matrix of the hth attention head. Multi-head attention output merging: The outputs of all attention heads are merged into the final dynamic attention output through a linear transformation:

[0015] Among them: Wo is the weight matrix of a linear transformation, Concat is the concatenation operation, and Head 1 , Head 2 ,…,Head H are the outputs of the 1st to the Hth attention heads. Based on the above technical solutions, in step S4, the optimization training of the model includes the following algorithms and formulas: (1) Transfer learning method The transfer learning steps used include: initializing the Vision Transformer model with the model parameters pre-trained on TA imageNet (such as Transformer weights); freezing some weights (such as the feature extraction layer) or performing fine-tuning optimization. Formula: Model parameter initialization:

[0016] Among them, θ pretrained are the parameters of the pre-trained model, which are loaded from a model that has been trained on a large-scale dataset. (2) Hybrid loss function The hybrid loss function combines cross - entropy loss (CE) and focal loss, and the weight parameters α and β control their relative importance; Formula:

[0017] where α is the weight of the cross - entropy loss, β is the weight of the focal loss, CE( y , y ^) is the cross - entropy loss function, and Focal( y , y ^) is the focal loss function; Cross - entropy loss (CE):

[0018] where yi is the value of the i - th category in the true label vector, log( i ^ y ) is the natural logarithm of the predicted probability i ^ y ^ i ; Focal loss (Focal Loss):

[0019] where γ is a modulating factor used to reduce the attention to easy - to - classify samples; yi is the value of the i - th category in the true label vector, y^ i : the predicted probability of the model for the i - th category; i (3) Optimizer i AdamW optimizer is used for training. This optimizer adds weight decay on the basis of Adam, which helps to improve the generalization ability; Formula: Gradient update:

[0020]

[0021] gt t where gt is the gradient at the current time step t , m t-1 and v t-1 are the momentum value and the exponential moving average of the squared gradient at the previous time step t −1 respectively, and β 1 , β 2 are hyperparameters; Parameter update:

[0022] Among them, θt : represents the current time step t of the model parameters, η is the learning rate, ϵ is the numerical stability term, and λ is the weight decay coefficient; (4) Model training The model training adopts the batch gradient descent method, and calculates the loss and updates the model parameters using a small batch of data each time; Formula: Batch loss:

[0023] Among them, B is the batch size, that is, the number of samples included in a batch, and L j is the loss of the j-th sample; Total training process:

[0024] Among them, min θ indicates that the optimization goal is to minimize a certain function with respect to the model parameters θ , E represents the expectation over all training samples; L( y , y ^) is the loss function, which is used to measure the difference between the true label y and the model prediction y ^; Based on the above technical solution, in step S5, the trained model is deployed to the edge device or embedded system of the drone, and the steps are as follows: 1) Model conversion and optimization: Use frameworks (such as ONNX, TensorRT) to quantize and optimize the trained model to improve the inference speed and reduce the consumption of computing resources; Formula:

[0025] Among them, θ ^ represents the optimized model parameters, θ represents the initial model parameters, Optimize represents the optimization algorithm or optimization process, precision represents the numerical precision used in the optimization process, and hardware represents the hardware device used in the optimization process; 2) Model deployment: Load the optimized model into the edge device or embedded system, and use the edge computing framework (such as TensorFlow Lite or PyTorch Mobile) to load the model:

[0026] Among them, θ ^ represents the optimized model parameters, and device represents the target hardware device on which the model is loaded; 3) Input preprocessing: Preprocess the images captured by the drone camera (such as resolution adjustment, normalization); Formula:

[0027] Among them, I is the original image, and μ and σ are the mean and standard deviation of the dataset.

[0028] Based on the above technical solutions, in step S5, the steps for real-time monitoring of tea garden aphid pests and diseases are as follows: (1) Tea garden aphid detection Use the trained object detection model to detect aphid targets in real time; 1) Object detection algorithm: Perform inference based on the object detection framework:

[0029] Among them, f is the model inference function, I′ is the preprocessed image, and θ^ represents the trained model parameters; 2) Object post-processing: Apply non-maximum suppression (NMS) to remove redundant boxes:

[0030] Among them, NMS represents the non-maximum suppression function, D represents the original detection result, IoU is the intersection over union, and τ is the confidence threshold; (2) Distribution heat map generation Use the detection results to generate an aphid distribution heat map to visually display the spatial distribution of tea garden aphids; 1) Generate an aphid distribution matrix: Construct a two-dimensional matrix H with the same resolution as the image, and map the center points of the detection boxes to the matrix:

[0031] Among them, δ( x − xi , y − yi ) is the Dirac delta function, which is used to determine whether the point ( xi , yi ) falls at the position ( x , y ), and xi, yi are the center coordinates of the i-th target; 2) Apply Gaussian filtering: Smooth the aphid distribution to generate a continuous heatmap:

[0032] Among them, H represents the original two-dimensional histogram, G represents the convolutional kernel (or filter), and ∗ represents the convolution operation; 3) Heatmap visualization: Map the smoothed matrix H′ to a pseudo-color image (for example, high, medium, and low aphid densities are represented by red, yellow, and green):

[0033] Among them, Colormap represents the color mapping function. Its role is to map the values in the histogram H ′ to color values, H ′ represents the processed two-dimensional histogram; Based on the same invention, the present invention also provides an application of a method for detecting small targets of tea tree aphids as described above in detecting small targets such as aphids in a tea garden.

[0034] The present invention has the following beneficial effects: 1. High-precision small target detection: The present invention combines multi-scale feature extraction and attention mechanism to improve the VisionTransformer model, significantly improving the detection accuracy of small target aphids. Traditional methods have low detection accuracy for small targets (such as aphids) in complex tea garden environments, while the present invention can more accurately capture the features of small targets through multi-scale fusion and dynamic attention mechanism, thereby improving the detection accuracy.

[0035] 2. High recall rate: The present invention designs a small target enhancement module to ensure that the model can effectively detect more aphid targets. Compared with the prior art, the present invention significantly improves the recall rate while maintaining high precision, reducing the situation of missed detection.

[0036] 3. High computational efficiency: The present invention optimizes the model structure, reduces the computational complexity, and at the same time maintains high detection performance. The present invention is superior to the prior art in terms of computational efficiency and can achieve real-time detection on resource-constrained devices (such as drones and edge devices).

[0037] 4. Applicable to complex tea garden environments: The present invention designs a robust detection model for the complexity of the tea garden environment (such as light changes and cluttered backgrounds). It can achieve high-precision detection in complex tea garden scenes, overcoming the problem of performance degradation of traditional methods in complex environments.

[0038] 5. With real-time monitoring capabilities: The trained model of the present invention is deployed on drones or edge devices, enabling real-time monitoring. Compared with traditional manual inspection or offline detection methods, the present invention can monitor tea garden aphid pests in real time, give early warnings and handle them in a timely manner, improving the efficiency of tea garden management.

[0039] 6. Application of transfer learning: The present invention uses transfer learning to optimize the training of the model, reducing the training time and data requirements. Through transfer learning, the present invention can quickly train a high-performance model with a small amount of labeled data, reducing the cost of data collection and annotation.

[0040] 7. With broad application prospects: The present invention combines the needs of tea garden ecological management and intelligent pest control, providing an efficient and intelligent solution. It is not only applicable to the detection of tea tree aphids, but also can be extended to other agricultural pest detection scenarios, with broad application prospects.

[0041] In summary, by combining multi-scale feature extraction, attention mechanism and Vision Transformer model, the present invention solves the problem of low detection accuracy of small aphid targets in the tea garden environment, and has advantages such as high accuracy, high recall rate, high computational efficiency and real-time monitoring capabilities. Compared with the prior art, the present invention has significant advantages in detection performance, environmental adaptability and application prospects, providing strong technical support for the intelligent prevention and control of tea garden pests. Brief Description of the Drawings

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only one embodiment of the present invention. For those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained according to the provided drawings.

[0043] Figure 1 : Flow chart of the present invention; Figure 2 : Example of the original data of the present invention. Detailed Embodiments

[0044] The following further illustrates the present invention with reference to the drawings and examples: The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention. In the present invention, unless otherwise specified, the equipment and raw materials used can be purchased from the market or are commonly used in the art. The methods in the following embodiments are conventional methods in the art unless otherwise specified.

[0045] Embodiment 1: A method for detecting small targets of tea tree aphids, the method comprising the following steps: S1: Collect diverse image data of the tea garden and perform preprocessing; S2: Construct a detection model based on Vision Transformer, including constructing a feature encoding module and designing a small target enhancement module; S3: Enhance the detection ability of small aphid targets through multi-scale fusion and dynamic attention mechanism; S4: Optimize and train the model using transfer learning; S5: Deploy the aphid detection system on a drone or edge device to achieve real-time monitoring of tea garden aphid pests and diseases.

[0046] Step S1: Collection and preprocessing of tea aphids image data (Tea aphids image dataset, named TA imageNet) Use a high-resolution camera or drone to collect image data in the tea garden. Perform data augmentation operations, including random cropping, rotation, brightness adjustment, etc., to generate diverse training samples. Normalize and standardize the images, and unify the input format to a size of 224×224 pixels.

[0047] Step S2: Construction of an improved Vision Transformer model (1) Feature encoding module: The input image is divided into image patches of a fixed size (P×P).

[0048] Each image patch generates an embedding vector through a linear transformation and is input into the Transformer network in combination with position encoding information.

[0049] Image patch division and embedding vector generation: The input image X∈R H×W×C is divided into N fixed-size image patches (Patches), each with a size of P×P.

[0050] The number of image patches is:

[0051] where, H represents the height of the input image, W represents the width of the input image, P represents the size of each image patch, in pixels; Each image patch generates an embedding vector through a linear transformation:

[0052] where: Xi is the pixel value matrix of the i-th image patch; Flatten(Xi) means flattening the image patch X i into a one-dimensional vector, We is the linear transformation matrix, be : represents the bias vector; After adding the positional encoding, the final embedding representation input to the Transformer is:

[0053] where, z 1 ; z 2 ; …; z N means concatenating the embedding vectors zi of all image patches together, P represents the positional encoding matrix; (2) Small object enhancement module: 1) Multi-scale feature extraction By introducing a multi-scale feature extraction module, multi-level feature representations are generated. Assuming the input feature is F, the feature representations at different scales are:

[0054] where, Conv sk (F): This is the convolution operation itself, Convsk represents using the k th convolutional kernel or the convolution operation at the k th scale, and F is the input feature map. K is the number of scales.

[0055] After fusing the multi-scale features:

[0056] where, Concat(F s 1,F s 2,…,F sK ) is the concatenation operation, F sk is the feature at the k th scale; 2) Multi-Head Self-Attention Mechanism Through the multi-head self-attention mechanism, feature enhancement is performed on small target regions. The calculation formula of the multi-head self-attention mechanism is as follows:

[0057] where: headi = Attention(Qi, Ki, Vi); h is the number of heads; Wo is the output weight matrix.

[0058] 3) Feature Enhancement Optimization For small target regions, a specific attention mechanism is introduced:

[0059] where Attention(Q, K, V) is the calculation process of the attention mechanism, Q is the query vector, K is the key vector, V is the value vector, and Q small focuses on the features of the aphid target region, which is generated using the Region Proposal Network (RPN) or other localization methods; Step S3: Introduction of the Dynamic Attention Mechanism In step S3, in order to dynamically adjust the weight distribution for different image patches, strengthen the attention to the regions with significant aphid features, and suppress background interference, we introduce a dynamic attention mechanism in the self-attention layer of the Transformer. The following is the detailed algorithm and formula description of this step: (1) Basic Principle of the Dynamic Attention Mechanism The self-attention mechanism captures the correlation between elements in the input sequence by calculating the attention weight matrix. In the standard Transformer, the attention weights are calculated as follows: Standard Self-Attention Mechanism

[0060] where: Q, K, V are the query vector (Query), key vector (Key), and value vector (Value) respectively. d k is the dimension of the key vector, which is used for scaling to prevent the inner product value from being too large. The dynamic attention mechanism introduces additional learnable parameters or dynamic adjustment modules on the basis of the standard attention calculation to enhance the attention to specific regions.

[0061] (2) Improvement of the Dynamic Attention Mechanism In the present invention, for the task of tea garden aphid detection, a dynamic weight adjustment module is designed to adaptively enhance the attention distribution of the target region. The specific steps are as follows: 1) Attention Weight Dynamic Adjustment Module Generate a dynamic weight wi for each image patch xi, and the calculation formula is:

[0062] Where: W 1 , W 2 is a learnable weight matrix; b 1 , b 2 is a bias term; σ is the Sigmoid activation function, and the output range is [0, 1].

[0063] 2) Attention calculation after weight correction After applying dynamic weight correction to each image patch, the updated attention calculation formula is:

[0064] Where, Softmax: performs the Softmax operation on the adjusted similarity matrix and converts it into a probability distribution, representing the attention weights of each query to the key. QK T : calculates the dot product of the query vector Q and the key vector K to obtain the similarity matrix, d k is the dimension of the key vector, and ⊙ represents the element-wise multiplication operation. w is the dynamic weight vector, which generates weights for each image patch.

[0065] (3)Implementation of dynamic attention in Vision Transformer In the multi-head self-attention (MHSA) of Vision Transformer, the dynamic attention is introduced as follows: 1) Multi-head dynamic attention calculation For the h-th head, the dynamic weight generation formula is:

[0066] Where, σ is the activation function, W 2( h ) is the weight matrix of the second fully connected layer, which is used to map the output of the hidden layer to the final output, W 1( h ) is the weight matrix of the first fully connected layer, which is used to map the input feature xi xi to the hidden layer, x i is the input feature vector, representing the i -th input sample. b 1( h ) and b2(h) are the bias terms of the first and second fully connected layers respectively.

[0067] The updated multi-head dynamic attention is:

[0068] Among them, Q h is the query vector of the h-th attention head, K h is the key vector of the h-th attention head, d k is the dimension of the key vector, ⊙ w (h) Multiply the scaled similarity matrix element-wise with the weight vector or matrix w (h) to dynamically adjust the attention weights, V h is the value vector of the h-th attention head; 2) Merging of multi-head attention outputs The outputs of all attention heads are merged into the final dynamic attention output through a linear transformation:

[0069] where: W o is the output weight matrix; H is the number of attention heads.

[0070] Step S4: Model training and optimization In step S4, the model training and optimization include the following algorithms and formulas: (1) Transfer learning method Transfer learning refers to using the model weights trained on a large-scale dataset (such as TA imageNet) as initialization parameters to improve the model performance on a small-sample dataset. The transfer learning steps used include: Initialize the VisionTransformer model with the model parameters (such as Transformer weights) pre-trained on TA imageNet.

[0071] Freeze some weights (such as the feature extraction layer) or perform fine-tuning optimization.

[0072] Formula: Model parameter initialization:

[0073] Among them, θ pretrained are the model parameters obtained by training on TA imageNet.

[0074] (2) Hybrid loss function The hybrid loss function combines cross - entropy loss (CE) and focal loss, and the weight parameters α and β control their relative importance.

[0075] Formula:

[0076] Where, α and β are two hyperparameters, CE( y , y ^) is the cross - entropy loss function, and Focal( y , y ^) is the focal loss function; Cross - entropy loss (CE):

[0077] Where, yi is the one - hot encoding of the true label, and yi^ is the probability predicted by the model.

[0078] Focal loss (Focal Loss):

[0079] Where, γ is a modulating factor used to reduce the attention to easily classified samples, y i is the value of the i - th category in the true label vector, and y^ i : The predicted probability of the model for the i - th category; (3) Optimizer AdamW optimizer is used for training. This optimizer adds weight decay on the basis of Adam, which helps to improve the generalization ability.

[0080] Formula: Gradient update:

[0081]

[0082] Where, g t is the gradient at the current time step t, m t-1 and v t-1 are the first - order and second - order momentum estimates representing the momentum value and the exponential moving average of the squared gradient at the previous time step t−1, respectively. β 1 ,β 2 are hyperparameters.

[0083] Parameter update:

[0084] Where, θt : represents the model parameters at the current time step \(t\), \(\eta\) is the learning rate, \(\epsilon\) is the numerical stability term, and \(\lambda\) is the weight decay coefficient.

[0085] (4) Model Training The model is trained using the batch gradient descent method, and the loss is calculated and the model parameters are updated using a small batch of data each time.

[0086] Formula: Batch Loss:

[0087] where \(B\) is the batch size, that is, the number of samples in a batch, and \(L\) j is the loss of the \(j\)-th sample.

[0088] Total Training Process:

[0089] where \(\min_{\theta}\) indicates that the optimization objective is to minimize a certain function with respect to the model parameters \(\theta\), and \(E\) represents the expectation over all training samples. \(L(y,\hat{y})\) is the loss function, which is used to measure the difference between the true label \(y\) and the model prediction \(\hat{y}\); Step S5: Aphid Detection and System Deployment (1) Model Deployment Deploy the trained model to the edge device or embedded system of the drone (such as NVIDIA Jetson or Raspberry Pi) to support real-time computing.

[0090] 1) Model Conversion and Optimization: Use frameworks (such as ONNX, TensorRT) to quantize and optimize the trained model to improve the inference speed and reduce the consumption of computing resources.

[0091] Formula:

[0092] where \(\theta\) is the original model weight and \(\hat{\theta}\) is the optimized model weight.

[0093] 2) Model Deployment: Load the optimized model into the edge device or embedded system. Use an edge computing framework (such as TensorFlow Lite or PyTorch Mobile) to load the model:

[0094] 3) Input Preprocessing: Preprocess the images captured by the drone camera (such as resolution adjustment, normalization).

[0095] Formula:

[0096] where I is the original image, and μ and σ are the mean and standard deviation of the dataset.

[0097] (2) Tea garden aphid detection Use the trained object detection model to detect aphid targets in real time.

[0098] 1) Object detection algorithm: Perform inference based on the object detection framework:

[0099] where f is the model inference function, I′ is the preprocessed image, represents the detected target, where x i ,y i is the center coordinate of the target, w i ,h i is the width and height, and s i is the confidence level.

[0100] 2) Object post-processing: Apply non-maximum suppression (NMS) to remove redundant boxes:

[0101] where IoU is the intersection over union, and τ is the confidence threshold.

[0102] (3) Distribution heatmap generation Use the detection results to generate an aphid distribution heatmap, visually showing the spatial distribution of aphids in the tea garden.

[0103] 1) Generate an aphid distribution matrix: Construct a two-dimensional matrix H with the same resolution as the image, and map the center points of the detection boxes to the matrix:

[0104] where δ is the Kronecker Delta function, and x i ,y i is the center coordinate of the i-th target.

[0105] 2) Apply Gaussian filtering: Smooth the aphid distribution to generate a continuous heatmap:

[0106] Among them, G is the Gaussian kernel function, and ∗ represents the convolution operation.

[0107] 3) Heatmap visualization: Map the smoothed matrix H′ to a pseudo-color image (e.g., high, medium, and low aphid densities are represented by red, yellow, and green):

[0108] (4) System deployment 1) Real-time inspection module: Integrate a camera and an aphid detection model on the drone to regularly inspect the tea garden: Capture real-time images; Input the model for inference; Update the heatmap display.

[0109] 2) Data transmission and storage: The detection results and heatmaps are transmitted to the ground station or the cloud through a wireless network (such as 5G or LoRa). Store the detection results locally or in the cloud for subsequent analysis.

[0110] 3) System optimization and fault tolerance: Optimize the scheduling of computing tasks to ensure the resource utilization rate of edge devices. Implement an automatic recovery mechanism for detection errors.

[0111] Example 2: Select areas where aphids are common in the Weishan Shangshantang Tea Garden for image acquisition, ensuring consistent lighting conditions and avoiding the influence of shadows and reflections on image quality. Use a high-resolution camera or smartphone to acquire images, ensuring that the image clarity meets the experimental requirements. Prepare the drone equipment for verifying the real-time detection task.

[0112] 2.1 Image acquisition: Randomly select multiple points in the tea garden, and take multiple high-resolution images at each point to ensure coverage of different angles and distances. A total of 3,200 high-resolution images containing aphid targets are acquired. Label the acquired images, marking the aphid targets in each image. Check the labeling quality, and eliminate blurred or mislabeled images to ensure the accuracy of the dataset.

[0113] 2.2 Dataset division Randomly divide the 3,200 images into a training set and a test set: Training set: 80% (2,560 images), for model training; Test set: 20% (640 images), for model evaluation. Perform data augmentation (such as rotation, scaling, flipping, etc.) on the training set to increase data diversity and the generalization ability of the model.

[0114] 2.3 Model selection and training Select a deep learning model suitable for object detection as the method of the present invention and compare it with the traditional CNN model. Use the training set (2560 images) to train the model, set appropriate hyperparameters (such as learning rate, batch size, etc.), and monitor the loss and accuracy during the training process. The early stopping strategy is adopted during the training process to prevent overfitting.

[0115] 2.4 Model Evaluation Use the test set (640 images) to evaluate the trained model. Calculate the mean average precision (mAP) of the model: the mAP of the method of the present invention reaches 88.3%; the mAP of the traditional CNN model is 75.2%. Compared with the traditional CNN model, the mAP of the method of the present invention has increased by 13.1 percentage points. Perform real-time detection tasks on the drone and record the frame rate (FPS): the frame rate of the method of the present invention reaches 15 FPS, meeting the actual monitoring requirements of the tea garden.

[0116] 2.5 Experimental Result Recording Experimental data table:

[0117] The above has shown and described the basic principles and main features of the present invention. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments. Therefore, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to include all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention.

[0118] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A method for small target detection of tea tree aphids, characterized in that: The method comprises the following steps: S1: Collect and preprocess diversified image data of tea garden; S2: Build a detection model based on Vision Transformer, including building a feature encoding module and designing a small target enhancement module; S3: Enhance the detection ability of small aphid targets through multi-scale fusion and dynamic attention mechanism; S4: Use transfer learning to optimize the model training; S5: Deploy the model aphid detection system trained in S4 on drones or edge devices to monitor aphid pests and diseases in tea gardens in real time.

2. A method for small target detection of tea tree aphids according to claim 1, characterized in that: In the step S1, a high-resolution camera and / or a drone are used to collect image data in the tea garden; the preprocessing is to enhance the data, including random cropping, rotation, brightness adjustment, and generate diversified training samples; and the image is normalized and standardized.

3. A method for small target detection of tea tree aphids according to claim 1, characterized in that: In step S2, the feature encoding module uses a hierarchical structure and combines a multi-scale fusion network to improve the perception of small target features.

4. A method for small target detection of tea tree aphids according to claim 1, characterized in that: In step S2, the small target enhancement module adopts a specific multi-head self-attention mechanism to enhance the detail features of the aphid target.

5. A method for small target detection of tea tree aphids according to claim 1, characterized in that: In step S3, the calculation formula of the standard attention mechanism is as follows: Among them, Q, K, and V are query vectors (Query), key vectors (Key), and value vectors (Value) respectively; d k is the dimension of the key vector, used for scaling to prevent the inner product value from being too large; The dynamic attention mechanism introduces additional learnable parameters and / or dynamic adjustment modules based on the standard attention calculation to enhance the attention of the region.

6. A method for small target detection of tea tree aphids according to claim 1, characterized in that: The dynamic attention is implemented in Vision Transformer as follows: Multi-head dynamic attention calculation: For the h-th attention head, the dynamic weight generation formula is: Among them, σ is the activation function, ReLU is the ReLU activation function, xi Indicates i input vector, b 1 (h) is the bias term of the first linear transformation, b2 (h) is the bias term of the second linear transformation; The updated multi-head dynamic attention is: Among them, the Softmax function is used to convert the input vector into a probability distribution. Q h For the h The query matrix of attention heads, K h is the key matrix of the h-th attention head, ⊙ is the element-by-element multiplication, V h is the value matrix of the h-th attention head; Multi-head attention output merging: The outputs of all attention heads are merged into the final dynamic attention output through linear transformation: Among them, Wo is a linear transformation weight matrix, Concat is a concatenation operation, Head1, Head2,…,Head H is the output of the 1st to Hth attention heads.

7. A method for detecting small targets of tea tree aphids according to claim 1, characterized in that: In step S4, the optimization training of the model includes the following algorithms and formulas: (1) Transfer learning method The transfer learning steps used include: initializing the Vision Transformer model using model parameters (such as Transformer weights) pre-trained on TA imageNet; freezing some weights (such as feature extraction layers) or fine-tuning optimization; formula: Model parameter initialization: in, θ pretrained The parameters of the pre-trained model are loaded from a model that has been trained on a large-scale dataset; (2) Hybrid loss function The hybrid loss function combines the cross entropy loss (CE) and focal loss, and the weight parameters α and β control their relative importance; formula: Among them, α is the weight of cross entropy loss, β is the weight of focal loss, CE( y , y ^) is the cross entropy loss function, Focal( y , y ^) is the focal loss function; Cross Entropy Loss (CE): Among them, yi is the first i The value of the category log( y ^ i ) is the predicted probability y ^ i The natural logarithm of Focal Loss: Among them, γ is a regulation factor used to reduce the attention to easy-to-classify samples, and yi is the first i The value of each category, y^ i :Model for i The predicted probability of each category; (3) Optimizer Use the AdamW optimizer for training. This optimizer adds weight decay (WeightDecay) based on Adam, which helps improve generalization ability; formula: Gradient Update: in, gt is the current time step t The gradient of t-1 and v t-1 They represent the previous time step respectively. t The momentum value of −1 and the exponential moving average of the squared gradient, β1, β2 are hyperparameters; Parameter update: in, θt : represents the current time step t The model parameters are: η is the learning rate, ϵ is the numerical stability term, and λ is the weight decay coefficient; (4) Model training The model training adopts the batch gradient descent method, using small batches of data each time to calculate the loss and update the model parameters; formula: Batch loss: Among them, B is the batch size, that is, the number of samples contained in a batch, and L j is the loss value of the jth sample; The overall training process: Among them, min θ The optimization goal is to minimize the model parameters θ A function of , E represents the expectation of all training samples, L( y , y ^) is the loss function used to measure the true label y And the model predicts y ^The difference between.

8. A method for detecting small targets of tea tree aphids according to claim 1, characterized in that: In step S5, the trained model is deployed to the edge device or embedded system of the drone, and the steps are as follows: 1) Model conversion and optimization: Use frameworks (such as ONNX and TensorRT) to quantize and optimize trained models to increase inference speed and reduce computing resource consumption; formula: in, θ ^ represents the optimized model parameters, θ Represents the initial model parameters, Optimize represents the optimization algorithm or optimization process, precision represents the numerical precision used in the optimization process, and hardware represents the hardware device used in the optimization process; 2) Model deployment: Load the optimized model to an edge device or embedded system and use an edge computing framework such as TensorFlowLite or PyTorch Mobile to load the model: in, θ ^ indicates the optimized model parameters, and device indicates the target hardware device where the model is loaded; 3) Input preprocessing: Preprocess the images captured by the drone camera (e.g., resolution adjustment, normalization); formula: Where I is the original image, μ and σ are the mean and standard deviation of the dataset.

9. A method for small target detection of tea tree aphids according to claim 1, characterized in that: In step S5, the steps of real-time monitoring of aphid pests in tea gardens are as follows: (1) Aphid detection in tea gardens Use the trained target detection model to detect aphid targets in real time; 1) Object detection algorithm: Reasoning based on the target detection framework: Where f is the model inference function, I′ is the preprocessed image, and θ^ represents the trained model parameters; 2) Target post-processing: Apply non-maximum suppression (NMS) to remove redundant boxes: Among them, NMS represents the non-maximum suppression function, D represents the original detection result, IoU is the intersection over union ratio, and τ is the confidence threshold; (2) Distribution heat map generation The detection results are used to generate a heat map of aphid distribution, which can visually display the spatial distribution of aphids in tea gardens; 1) Generate aphid distribution matrix: Construct a two-dimensional matrix H with the same resolution as the image and map the center point of the detection box to the matrix: Among them, δ( x − xi , y − yi ) is the Dirac delta function, which is used to determine the point ( xi , yi ) falls in position ( x , y ), xi,yi are the center coordinates of the i-th target; 2) Apply Gaussian filtering: Smooth the aphid distribution to produce a continuous heat map: in, H represents the original two-dimensional histogram, G represents the convolution kernel (or filter), ∗ represents the convolution operation; 3) Heatmap visualization: Map the smoothed matrix H′ into a pseudo-color image (such as red, yellow, and green represent high, medium, and low aphid densities): Among them, Colormap represents the color mapping function. Its function is to convert the histogram H ′ is mapped to color values. H ′ represents the processed two-dimensional histogram.

10. Use of the method for detecting small targets of tea tree aphids as claimed in any one of claims 1 to 9 in detecting small targets of tea trees such as aphids in a tea garden.