Citrus detection method using frequency domain aggregation attention mechanism and multi-scale encoder

By constructing a self-picked citrus detection data set and designing frequency aggregation attention network and multi-scale Transformer encoder, the problem of small target and occlusion target detection is solved, and high-precision citrus detection is achieved, which is suitable for fruit grading and yield statistics in the field of intelligent agriculture.

CN120375087AActive Publication Date: 2025-07-25TIANJIN UNIV

Patent Information

Application Number
CN202510522688.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-25
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing citrus detection methods have problems such as insufficient feature extraction capabilities, insufficient context information utilization, and insufficient retention of high-frequency detailed information when dealing with small targets and occlusion targets. The lack of high-quality data sets limits the generalization ability of the model.

Method used

A self-picked citrus detection data set was constructed, a frequency aggregation attention network (FAN) was designed to enhance the frequency characteristic sensitivity of small targets and occluded targets, and a multi-scale Transformer encoder was developed, combining Haar wavelet fusion module and IoU perceptual query selection to optimize the loss function to improve detection accuracy.

Benefits of technology

High-precision detection of citrus fruits, especially small targets and occlusion targets, is achieved, and is suitable for fruit grading and yield statistics in the field of intelligent agriculture, which significantly improves the convergence speed and detection accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375087A_ABST
    Figure CN120375087A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of citrus detection in agricultural automation, in particular to a citrus detection method based on a frequency domain aggregation attention mechanism and a multi-scale encoder. According to the method, an innovative detection framework is provided for the problem of small target and occlusion target detection, and the detection framework comprises a frequency aggregation attention network (FAN) and a multi-scale Transform encoder. The frequency aggregation attention network decomposes the feature map through two-dimensional discrete wavelet transform, and enhances the frequency domain features of the small target and the shielding target; the multi-scale Transform encoder is combined with convolution feature pyramid operation, and high-frequency detail information is reserved through a wavelet fusion module. In addition, an IoU is adopted to perceive query selection and optimize a loss function, so that the detection precision is remarkably improved. Through frequency domain feature enhancement and multi-scale feature fusion, high-precision detection of citrus fruits, especially small targets and sheltered targets, is realized, and the method is suitable for fruit grading, yield statistics and other scenes in the field of intelligent agriculture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of citrus detection in agricultural automation, and in particular to a citrus detection method based on a frequency-domain aggregation attention mechanism and a multi-scale encoder of a Transformer architecture. Background Art

[0002] As a key link in the fields of agricultural intelligent management and automated robots, the core task of citrus target detection is to use advanced image processing and machine learning technologies to achieve efficient and accurate identification and positioning of citrus fruits. In the process of intelligent orchard management and agricultural automation, the application scenarios of citrus detection technology are becoming increasingly extensive, covering aspects such as fruit quantity statistics, pest and disease prediction, and robot picking. However, the complexity of the natural environment and the diversity of fruit growth patterns make the detection of small citrus and occluded citrus a technical challenge.

[0003] Traditional citrus detection methods mainly rely on object detection algorithms based on handcrafted features, such as template matching methods based on color, shape, and texture features. These methods can achieve certain detection effects under specific conditions, but when faced with complex backgrounds, fruit occlusion, and illumination changes, the detection accuracy and robustness significantly decrease. In recent years, the rise of deep learning technology has brought new breakthroughs to citrus detection. Object detection algorithms based on convolutional neural networks (CNNs), such as YOLO, Faster R-CNN, etc., can effectively improve the detection accuracy and speed by learning the feature patterns in the data. However, existing methods still have limitations in dealing with small targets and occluded targets, mainly reflected in insufficient feature extraction ability, inadequate utilization of context information, and insufficient retention of high-frequency detail information.

[0004] In addition, the lack of existing citrus detection datasets further limits the training effect of the model. Most publicly available datasets only contain a small number of annotated images, and the annotation accuracy is low, making it difficult to meet the detection requirements in complex scenarios. There is a lack of dedicated datasets for citrus picking tasks, which limits the generalization ability of the model in practical applications. Therefore, developing a citrus detection method that can effectively handle small targets and occluded targets and constructing a high-quality citrus detection dataset are of great significance for promoting the intelligent development of the citrus industry. Summary of the Invention

[0005] The present invention discloses a citrus detection method using a frequency domain aggregation attention mechanism and a multi-scale encoder, aiming to solve the problem of accurate detection of small citrus fruits and occluded citrus fruits. First, a self-collected citrus detection dataset containing occluded and small fruits is constructed, providing a key data basis for model training. Second, a frequency aggregation attention network (FAN) is designed to effectively capture subtle features in complex scenes by enhancing the model's sensitivity to the frequency features of small and occluded targets. At the same time, an efficient multi-scale Transformer encoder is innovatively developed, and a Haar wavelet fusion module is introduced to retain the high-frequency detail information of small and occluded targets by eliminating the aliasing effect in the upsampling process. In addition, a multi-level supervision strategy is adopted, combined with IoU-aware query selection and decoder dynamic optimization of the loss function, significantly improving the model convergence speed and detection accuracy. Through multi-dimensional technological innovation, the present invention achieves high-precision detection of citrus fruits, especially small and occluded targets, and is applicable to scenarios such as fruit grading and yield statistics in the field of intelligent agriculture.

[0006] To achieve the above object, the present invention adopts the following technical solutions: Step 1: Construct a self-collected citrus detection dataset, including citrus images with occlusions and small fruits, with an image resolution of 3456×3456, and 47,819 citrus targets are labeled. Step 2: Preprocess the image data obtained in Step 1 to obtain preprocessed citrus images. Step 3: Design a frequency aggregation attention network (FAN), perform two-dimensional discrete wavelet transform (2DDWT) on the input feature map, decompose it into different frequency subbands, and perform information compression in the channel dimension and frequency dimension respectively to enhance the frequency features of small and occluded targets. Step 4: Develop a multi-scale Transformer encoder, combine the self-attention mechanism and convolutional feature pyramid operations, and perform cross-scale feature fusion through the Haar wavelet fusion module to retain high-frequency details. Step 5: Adopt IoU-aware query selection and decoder to optimize the loss function and improve the model convergence speed and detection accuracy.

[0007] Step 6: Import the trained network model, input the test set into the network for network performance testing. The characteristics of the present invention also lie in: Further, the specific process of Step 3 is as follows: Perform two-dimensional discrete wavelet transform (2DDWT) decomposition on the input feature map to obtain low-frequency and high-frequency subbands.

[0008] Perform channel - dimension compression for each frequency sub - band, use large convolutional kernels to expand the receptive field, and enhance the context semantics; perform global average pooling on the frequency sub - bands, calculate weights through a dimensionality - reducing fully - connected layer and an activation layer, and perform a final scaling operation to obtain the enhanced feature map.

[0009] The specific process of step 4 is as follows: First, pass through the top - down path. Align the multi - scale features through 1×1 convolution to ensure the consistency of different - scale feature maps in the channel dimension, thereby providing a unified feature representation for subsequent feature fusion.

[0010] Then, pass through the bottom - up path: Deepen the feature map through 3×3 convolution, expand the receptive field and enhance the feature expression ability, thereby extracting more discriminative feature information and improving the detection accuracy of the model for the target.

[0011] In the process of dual - path feature fusion, a new fusion module, the Haar wavelet fusion module, is proposed as follows: The Haar wavelet fusion module uses the Haar wavelet transform to decompose the large - scale features, cascade them with the small - scale features, enhance the low - frequency information, and obtain the final fusion output through residual operations and inverse wavelet transform.

[0012] Furthermore, the loss function described in step 5 includes: Bounding box regression loss and classification loss Among them, IoU information is incorporated into the classification loss to minimize the difference between the confidence score and the localization, and to promote the convergence of the model; ; Among them, represents the ground - truth bounding box, b represents the predicted bounding box, represents the ground - truth class, C represents the predicted class, and IoU represents the intersection - over - union of the predicted bounding box.

[0013] The beneficial effects of the present invention are: 1. The present invention constructs a self - harvesting citrus detection data set containing occlusions and small fruits, providing a key data basis for model training. A frequency aggregation attention network (FAN) is designed. By enhancing the model's sensitivity to the frequency features of small and occluded targets, it can effectively capture the subtle features in complex scenes and enhance the ability to extract fuzzy and complex features.

[0014] 2. The present invention has developed an efficient multi-scale Transformer encoder and introduced a Haar wavelet fusion module. By eliminating the aliasing effect in the upsampling process, the high-frequency detail information of small targets and occluded targets is retained. A multi-level supervision strategy is adopted, combined with IoU-aware query selection and decoder dynamic optimization of the loss function, which significantly improves the model convergence speed and detection accuracy. Through multi-dimensional technological innovation, the present invention realizes high-precision detection of citrus fruits, especially small targets and occluded targets, and is applicable to scenarios such as fruit grading and yield statistics in the field of intelligent agriculture. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is the overall network framework of the present invention, which significantly improves the detection performance of small and occluded citrus; Figure 2 is a schematic diagram of the structure of the Frequency Aggregation Attention Network (FAN), which enhances the frequency domain sensitivity of the model; Figure 3 is a schematic diagram of the structure of the Haar wavelet fusion module, which reduces frequency domain aliasing and efficiently fuses multi-scale features; Figure 4 is a schematic diagram of the target detection results of the self-collected data set, and the false detection rate and missed detection rate of the model are significantly reduced; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments. Based on the embodiments in this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this disclosure. Constructing a citrus target detector includes parts such as data set construction, attention mechanism design, encoder network structure reconstruction, and loss function optimization; Data set preparation: A self-collected citrus data set is constructed, which contains 1388 citrus canopy layer images, covering a variety of citrus varieties. The image resolution is 3456×3456, and 47819 citrus targets are labeled. The data set is divided into a training set, a validation set, and a test set, with a ratio of 3:1:1; To solve the deficiencies of the prior art, the present invention discloses a citrus detection method using a frequency domain aggregation attention mechanism and a multi-scale encoder, as Figure 1 shown, the method includes, Step 1: Construct an attention mechanism for extracting the frequency domain features of small and occluded citrus; Step 2: Use a top-down and bottom-up feature fusion path in the encoder; Step 3: Construct a new feature fusion module in the top-down path.

[0017] In a specific example of the present disclosure, for Step 1, as Figure 2As shown, the specific operation is as follows: Perform a two-dimensional discrete wavelet transform on the input feature map, which is decomposed into a low-frequency sub-band (LL) and high-frequency sub-bands (LH, HL, HH). The low-frequency sub-band retains the contour information of the image, and the high-frequency sub-bands retain the detail information of the image: ; Among them, and are the low-pass and high-pass filters of the 2DDWT respectively, represents the input feature map, and represent downsampling operations along the rows and columns respectively, represents the four sub-bands after wavelet transform; Compress each frequency sub-band separately in the channel dimension. Use a large convolutional kernel (such as 17×17) to expand the receptive field and enhance the contextual semantic understanding of small targets and occluded targets; ; Among them, represents the fully connected operation, represents global average pooling, X i is the sub-band after wavelet decomposition.

[0018] Y,S i represents the intermediate feature. Perform global average pooling on the frequency sub-band, calculate the weights through a dimensionality reduction fully connected layer and an activation layer, and perform the final scaling operation to obtain the enhanced feature map; ; Among them, represents the activation layer to calculate the weights, C represents the concatenation operation, represents the scaling operation, is the enhanced feature map; For step 2, the specific operation is as follows: Reconstruct the multi-scale Transformer encoder. Through dual-path multi-level feature fusion, the model can fully extract the target location information and semantic details; As Figure 2 shown, the multi-scale Transformer encoder performs multi-scale feature fusion through two paths: top-down and bottom-up; Top-down path: Align the multi-scale features through 1×1 convolution to ensure the consistency of features at different scales; Bottom-up path: Deepen the feature map through 3×3 convolution to enhance the feature expression ability; In the multi-scale Transformer encoder, the Haar wavelet fusion module further eliminates the frequency domain aliasing problem caused by upsampling and continuously extracts the frequency domain information of the features; For step 3, asFigure 3 As shown in the figure, the specific operation is as follows: Use Haar wavelet transform to decompose large-scale features, cascade them with small-scale features, and enhance low-frequency information. The final fusion output is obtained through residual operations and inverse wavelet transform (2DIWT), retaining high-frequency details; ; Among them, represents the large-scale feature map of the lower level, represents the small-scale feature map of the higher level, E n represents n enhancement modules, represents the intermediate feature, represents the final fusion output; Finally, perform IoU-aware decoder and loss function optimization; ; Among them, is used to measure the difference between the predicted bounding box and the ground truth bounding box. The IoU loss function is adopted to directly optimize the localization accuracy of the bounding box; represents the use of the cross-entropy loss function to optimize the classification result. The IoU information is incorporated into the classification loss to minimize the difference between the confidence score and the localization, promoting model convergence; represents the ground truth bounding box, b represents the predicted bounding box, represents the ground truth class, C represents the predicted class, and IoU represents the intersection over union of the predicted bounding box; At the same time, a multi-level supervision strategy is adopted to optimize the loss function through IoU-aware query selection and decoder. During the training process, features of different scales are supervised to ensure that the model can accurately detect targets at different scales.

[0019] Evaluation metrics: AP 50 : Represents the average precision when the IoU (Intersection over Union) threshold is 0.50. IoU is the ratio of the overlapping area between the predicted bounding box and the ground truth bounding box. AP50 measures the detection accuracy of the model when IoU is 0.50.

[0020] AP 75 : Represents the average precision when the IoU threshold is 0.75. This is a more stringent evaluation criterion because a higher IoU threshold requires a larger overlapping area between the predicted bounding box and the ground truth bounding box.

[0021] AP 50-95: It represents the average precision between IoU thresholds from 0.50 to 0.95. This is a comprehensive metric used to evaluate the overall performance of the model at different IoU thresholds.

[0022] AP S , AP L , AP M : It represents the average precision for small, medium, and large objects. Small objects are usually defined as those with an area less than 32². Medium-sized objects are usually defined as those with an area between 32² and 96². Large objects are usually defined as those with an area greater than 96².

[0023] FPS: The number of frames processed per second by the model during the inference stage, which is an indicator of the model's real-time performance. The higher the FPS, the better the performance of the algorithm.

[0024] GFLOPs: It represents the number of floating-point operations per second in billions, used to measure the computational complexity and hardware requirements of the model. The higher the GFLOPs, the greater the computational load of the model and the higher the requirements for hardware.

[0025] Experimental settings: The network architecture is implemented using PyTorch; Adam is used as the optimizer, and there are a total of 150 epochs. The learning rate of the backbone network is set to 10 -4 , and the learning rates of the encoder and decoder are set to 10 -5 . Training was carried out on an NVIDIA GeForce RTX 4090 with 24GB of RAM, and the training time was approximately 4 hours.

[0026] Comparative experiments: In this invention, on the self-collected dataset, comparisons are made with some baseline methods and state-of-the-art object detection methods. Table 1 shows the quantitative results on the dataset, and the method we proposed achieved the best results in all cases. The best and second-best results are marked in bold in the table.

[0027] Table 1: ; Complexity analysis: This invention designs models with different complexities to provide more diverse choices. The experimental results are shown in Table 2.

[0028] Table 2 Comparison of different model complexities, * indicates that the experiment was measured with a batch size of 1.

[0029] .

[0030] In addition, the detection results of the model in this invention are visualized, such as Figure 4As shown. Ground Truth represents the true bounding box of citrus targets, and Prediction represents the bounding box predicted by the model.

Claims

1. A citrus detection method using a frequency domain aggregation attention mechanism and a multi-scale encoder, characterized in that, The implementation is carried out according to the following steps: Step 1: Construct a self-collected citrus detection dataset, including citrus images with occlusions and small fruits and the positions of citrus bounding boxes; Step 2: Design a Frequency Aggregation Attention Network (FAN), perform two-dimensional discrete wavelet transform (2DDWT) on the input feature map, decompose it into different frequency sub-bands, perform information compression in the channel dimension and frequency dimension respectively, and enhance the frequency features of small targets and occluded targets; Step 3: Develop a multi-scale Transformer encoder, combine the self-attention mechanism and convolutional feature pyramid operations, perform cross-scale feature fusion through the Haar wavelet fusion module, and retain high-frequency details; Step 4: Adopt the Intersection over Union (IoU)-aware query selection and decoder, optimize the loss function, and improve the convergence speed and detection accuracy of the model; Step 5: Import the trained network model, input the test set into the network for network performance testing; Step 6: Evaluate and analyze the test results using comprehensive indicators such as mean average precision, mean small object precision, mean recall rate, frame rate, and number of parameters.

2. The citrus detection method according to claim 1, wherein The Frequency Aggregation Attention Network (FAN) specifically includes: Step 2.1: Decompose the input feature map by two-dimensional discrete wavelet transform (2DDWT) to obtain low-frequency and high-frequency sub-bands, and calculate using the following formula; ; Among them, and are the low-pass and high-pass filters of the 2DDWT respectively, represents the input feature map, and represent the downsampling operations along rows and columns respectively, represents the four sub-bands after wavelet transform; Step 2.2: Perform channel dimension compression on each frequency sub-band respectively, use a large convolutional kernel to expand the receptive field, and enhance the context semantics; Step 2.3: Perform global average pooling on the frequency sub-bands, calculate the weights through a dimensionality reduction fully connected layer and an activation layer, and perform the final scaling operation to obtain the enhanced feature map; ; Among them, represents two-dimensional inverse wavelet transform, represents large convolution kernel operation, represents fully connected operation, represents global average pooling, represents the excitation operation in channel attention, and C represents the concatenation operation, represents element-wise multiplication, Y, S i represents intermediate features, represents the enhanced feature map.

3. The citrus detection method according to claim 1, characterized in that, The multi-scale Transformer encoder includes: Step 3.1: Top-down path and bottom-up path, align and deepen the multi-scale features through 1×1 convolution and 3×3 convolution respectively; Step 3.2, Haar wavelet fusion module: Decompose the large-scale features by using Haar wavelet transform, cascade them with the small-scale features, enhance the low-frequency information, and obtain the final fusion output F through residual operation and inverse wavelet transform S ; ; Among them, represents the large-scale feature map of the lower level, represents the small-scale feature map of the higher level, E n represents n enhancement modules, represents the intermediate feature, represents the final fusion output.

4. The citrus detection method according to claim 1, characterized in that, The loss function includes: Step 4.1, Bounding box regression loss and classification loss wherein the IoU information is incorporated into the classification loss to minimize the difference between the confidence score and the localization, and promote the model convergence; ; Among them, represents the true bounding box, b represents the predicted bounding box, represents the true category, C represents the predicted category, and IoU represents the intersection over union of the predicted bounding box.

5. According to claim 1, it is characterized in that, The network parameters set in Step 4 include the dataset batch size batch_size, learning rate learning_rate, number of training epochs epoch, method for initializing model parameters, and optimization method.

Citation Information

Patent Citations

  • Dynamic illumination face image quality enhancement method based on multi-scale attention mechanism

    CN115880225A

  • Crop target detection method and system based on spectrum expansion method

    CN116091917A

  • Light-weight remote sensing image target detection method based on deep learning

    CN118334313A

  • Infrared weak and small target detection method based on wavelet guidance

    CN118587507A

  • Navel orange detection method in agricultural environment

    CN119516332A

Cited By

  • Dense overlapping target detection method based on wavelet enhancement sparse hybrid expert model

    CN121353953A