Oscillating fruit detection and tracking method, device and equipment

By improving the YOLOv8n and DeepSORT algorithms and combining them with the C2f-Dattention, CAA-HSFPN, and SPPF-LSKA modules, the accuracy and robustness issues of vibrating fruit detection and tracking in harvesting robots were solved, achieving efficient fruit picking.

CN120635542APending Publication Date: 2025-09-12CHANGZHOU UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510707880.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing picking robots prolong the fruit recognition and positioning time due to vibration during the separation process of fruits from branches. In addition, the YOLO series algorithms and DeepSORT algorithms have problems with insufficient accuracy and robustness in detecting and tracking vibrating fruits in complex backgrounds.

Method used

A multi-scale strategy is adopted, combined with improvements to YOLOv8n and DeepSORT, and the C2f-Dattention mechanism, CAA-HSFPN lightweight network, and SPPF-LSKA multi-scale detection module are introduced to optimize the detection and tracking algorithms, improve detection accuracy and processing speed, and enhance tracking stability through adaptive noise scale Kalman filtering and CIoU matching.

Benefits of technology

It significantly improves the accuracy and processing speed of fruit detection, enhances the stability and accuracy of fruit tracking, and is suitable for fruit picking tasks in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635542A_ABST
    Figure CN120635542A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing and target tracking, and provides an oscillation fruit detection and tracking method, device and equipment, and the method comprises the steps: collecting fruit image data of an orchard, and carrying out the preprocessing of an image; improving and constructing a multi-scale oscillation fruit detection network model based on YOLOv8n; constructing a multi-scale oscillation fruit detection network model based on a DeepSORT algorithm; constructing a multi-scale oscillation fruit detection and tracking network model; and training the multi-scale oscillation fruit detection and tracking model by adopting the labeled data, and verifying the multi-scale oscillation fruit detection and tracking model. On the basis of improvement of YOLOv8n and DeepSORT, a multi-scale strategy is adopted for detecting and tracking concussion fruits, by introducing a C2f-Dattention attention mechanism, a CAA-HSFPN lightweight network and an SPPF-LSKA multi-scale detection module, the precision and processing speed of fruit detection are remarkably improved, meanwhile, by combining with an optimized DeepSORT tracking algorithm, the accuracy and processing speed of fruit detection are improved, and the accuracy and processing speed of fruit detection are improved. And the stability and accuracy of fruit tracking are further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention generally relate to the field of image processing and target tracking technology, and more particularly to a method, device and apparatus for detecting and tracking vibrating fruit. Background Art

[0002] When a fruit-picking robot is harvesting fruit, it first initializes its modules and positions its robotic arm at an appropriate distance from the fruit tree. A video camera captures and obtains image information of the target fruit, and image processing software identifies and locates the target fruit. The robot's control computer drives the robotic arm to complete the harvesting actions, such as grasping the fruit and separating the fruit from the branch. However, when the harvesting robot separates the fruit from the branch, whether by cutting or twisting, it will cause other fruits on the tree to vibrate. In addition, the effects of wind can also cause the fruit to vibrate, which will increase the time required to identify and locate subsequent fruits, interrupt continuous harvesting, and thus increase the overall fruit picking time. Therefore, the detection and tracking algorithm for vibrating fruit is crucial to improving the efficiency of the harvesting robot.

[0003] In recent years, deep learning technology has made significant progress in the fields of image recognition and object detection. As a representative real-time object detection framework, the YOLO series of algorithms has attracted widespread attention for its fast and efficient performance. However, the YOLO series of algorithms still faces challenges when dealing with small target detection, complex backgrounds, and detection of vibrating fruit under changing lighting conditions. YOLOv8 has made significant progress in the speed and accuracy of target detection, but in complex backgrounds, especially when there are many fruits and severe occlusion, these require further optimization of the algorithm to reduce model complexity and the number of parameters. Faced with the problem of vibrating fruit detection in natural environments, the algorithm must not only be accurate but also robust, which may involve specific data preprocessing, model adjustment, or the use of reinforcement learning strategies.

[0004] At the same time, deep learning technology has also made significant breakthroughs in the field of target tracking. As an efficient multi-target tracking method, the DeepSORT algorithm combines the advantages of target detection and Kalman filtering, and has demonstrated strong stability and accuracy in practical applications. However, DeepSORT still faces certain challenges when dealing with dense targets, target occlusion, and tracking of similar targets. Although the algorithm has achieved a good balance between real-time performance and accuracy, its robustness and stability still need to be further improved in complex scenes with fast-moving and densely arranged targets. To meet these challenges, the algorithm may need to introduce more sophisticated feature extraction methods, more efficient matching strategies, or adopt more adaptable tracking models to enhance its performance in complex environments. Summary of the Invention

[0005] To solve the above problems, the present invention is based on the improvement of YOLOv8n and DeepSORT, and adopts a multi-scale strategy for detection and tracking of oscillating fruits. By introducing the C2f-Dattention attention mechanism, CAA-HSFPN lightweight network and SPPF-LSKA multi-scale detection module, the accuracy and processing speed of fruit detection are significantly improved. At the same time, combined with the optimized DeepSORT tracking algorithm, the stability and accuracy of fruit tracking are further enhanced.

[0006] According to an embodiment of the present invention, a method, apparatus and device for detecting and tracking vibrating fruit are provided.

[0007] In a first aspect of the present invention, a method for detecting and tracking vibrating fruit is provided. The method comprises:

[0008] Step S01: collecting orchard fruit image data and preprocessing the image;

[0009] Step S02: constructing a multi-scale oscillating fruit detection network model based on YOLOv8n improvement;

[0010] Step S03: constructing a multi-scale vibrating fruit detection network model based on the DeepSORT algorithm;

[0011] Step S04: constructing a multi-scale vibrating fruit detection and tracking network model;

[0012] Step S05: Use the labeled data to train the multi-scale vibrating fruit detection and tracking model and verify it.

[0013] Furthermore, the multi-scale oscillating fruit detection network model described in step S02 is improved based on YOLOv8n, including: an input end, a backbone network, a neck module and a prediction end. The input end adopts mosaic data enhancement, adaptive anchor frame calculation and adaptive grayscale filling; the backbone network adopts the first convolution module Conv, the second convolution module Conv, the first cross-stage partial fusion module C2f, the third convolution module Conv, the second cross-stage partial fusion module C2f, the fourth convolution module Conv, the third cross-stage partial fusion module C2f, the fifth convolution module Conv, the C2f-Dattention attention mechanism module and the multi-scale detection module SPPF-LSKA which are arranged in sequence; the 10th, 13th, 15th, 20th and 22nd layers of the neck module all adopt the CAA-HSFPN lightweight structure; the prediction end uses two feature vectors of different scales at the 18th and 25th layers of the neck module for prediction results.

[0014] Furthermore, the C2f-Dattention method includes: realizing cross-fusion of information from different channels through the C2f layer; calculating the importance of each channel and spatial position in the Dattention module, and assigning different weights to each position; performing dimensionality reduction output through a 1×1 convolution layer and batch normalization, and performing a residual connection with the input features.

[0015] Furthermore, the CAA-HSFPN lightweight structure includes:

[0016] The CAA module generates channel descriptors through global average pooling operations;

[0017] Use lightweight fully connected layers to adjust the weight of each channel and introduce a channel attention mechanism;

[0018] The HSFPN module adopts a multi-scale feature pyramid structure for feature fusion and combines the hybrid attention mechanism of space and channel to extract multi-level features at different scales;

[0019] Depthwise separable convolution is used to reduce the amount of computation, and the feature dimension is reduced through a 1×1 convolution layer.

[0020] Furthermore, the CAA-HSFPN lightweight structure also uses the PKINet structure to improve adaptability to complex scenarios.

[0021] Furthermore, the multi-scale oscillating fruit tracking network model described in step S03: uses a multi-scale oscillating fruit detection network model to replace the traditional detector; uses the ResNest50 feature extraction network to replace the CNN network, and uses the adaptive noise scale Kalman filter to replace the traditional Kalman filter algorithm; uses CIoU matching instead of IoU matching.

[0022] Furthermore, the multi-scale oscillating fruit detection and tracking network model described in step S04 includes:

[0023] Detecting fruits and extracting features in a complex environment using the multi-scale oscillating fruit detection network model constructed in step S02;

[0024] The detection result of the multi-scale oscillation fruit detection network model is passed to the multi-scale oscillation fruit tracking network model constructed in step S03 to track the detected fruit target;

[0025] The multi-scale oscillating fruit tracking network module feeds the output results back to the multi-scale oscillating fruit detection network module for correcting the detection error in subsequent frames.

[0026] In a second aspect of the present invention, a device for detecting and tracking vibrating fruit is provided. The device comprises:

[0027] Fruit image acquisition module: used to collect orchard fruit image data and pre-process the images;

[0028] Detection model construction module: used to improve and build a multi-scale oscillating fruit detection network model based on YOLOv8n;

[0029] Tracking model construction module: used to build a multi-scale oscillating fruit detection network model based on the DeepSORT algorithm;

[0030] Detection and tracking model building module: used to build a multi-scale oscillating fruit detection and tracking network model;

[0031] Model training module: used to train and verify the multi-scale vibrating fruit detection and tracking model using labeled data.

[0032] In a third aspect of the present invention, an electronic device is provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method according to the first aspect of the present invention is implemented.

[0033] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present invention is implemented.

[0034] Based on the improvements of YOLOv8n and DeepSORT, this paper adopts a multi-scale strategy for detection and tracking of vibrating fruits. By introducing the C2f-Dattention mechanism, the CAA-HSFPN lightweight network and the SPPF-LSKA multi-scale detection module, the accuracy and processing speed of fruit detection are significantly improved. At the same time, combined with the optimized DeepSORT tracking algorithm, the stability and accuracy of fruit tracking are further enhanced.

[0035] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings, in which:

[0037] Figure 1 A flow chart of a method for detecting and tracking vibrating fruits according to an embodiment of the present invention is shown;

[0038] Figure 2 A technical roadmap of a method for detecting and tracking vibrating fruit according to an embodiment of the present invention is shown;

[0039] Figure 3 It shows a schematic structural diagram of a vibrating fruit detection model according to an embodiment of the present invention;

[0040] Figure 4 The structure diagram of the C2f-Dattention mechanism module according to an embodiment of the present invention is shown;

[0041] Figure 5 shows a structural diagram of a Dattention module according to an embodiment of the present invention;

[0042] Figure 6 It shows a structural diagram of the SPPF-LSKA multi-scale detection module according to an embodiment of the present invention;

[0043] Figure 7 A CAA-HSFPN lightweight network structure diagram according to an embodiment of the present invention is shown;

[0044] Figure 8 A flow chart of a method for detecting and tracking vibrating fruit according to an embodiment of the present invention is shown;

[0045] Figure 9 A diagram showing the structure of a ResNest50 lightweight feature extraction network according to an embodiment of the present invention is shown.

[0046] Figure 10 A diagram of a CIoU matching strategy according to an embodiment of the present invention is shown;

[0047] Figure 11 A block diagram of an apparatus for detecting and tracking vibrating fruit according to an embodiment of the present invention is shown;

[0048] Figure 12 A schematic diagram of a device for detecting and tracking vibrating fruits according to an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0050] According to the embodiments of the present invention, methods, devices and equipment for detecting and tracking vibrating fruits are proposed. Based on the improvements of YOLOv8n and DeepSORT, a multi-scale strategy is adopted to detect and track vibrating fruits. By introducing the C2f-Dattention attention mechanism, CAA-HSFPN lightweight network and SPPF-LSKA multi-scale detection module, the accuracy and processing speed of fruit detection are significantly improved. At the same time, combined with the optimized DeepSORT tracking algorithm, the stability and accuracy of fruit tracking are further enhanced.

[0051] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.

[0052] Figure 1 1 is a flow chart of a method for detecting and tracking vibrating fruit according to an embodiment of the present invention. The method includes:

[0053] Step S01: collecting orchard fruit image data and preprocessing the image;

[0054] Step S02: constructing a multi-scale oscillating fruit detection network model based on YOLOv8n improvement;

[0055] Step S03: constructing a multi-scale vibrating fruit detection network model based on the DeepSORT algorithm;

[0056] Step S04: constructing a multi-scale vibrating fruit detection and tracking network model;

[0057] Step S05: Use the labeled data to train the multi-scale vibrating fruit detection and tracking model and verify it.

[0058] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0059] In order to explain the above-mentioned method of detecting and tracking vibrating fruits more clearly, a specific embodiment is used for illustration below. However, it should be noted that this embodiment is only for better illustrating the present invention and does not constitute an improper limitation to the present invention.

[0060] The following is a more detailed explanation of the method for detecting and tracking vibrating fruit using a specific example:

[0061] like Figure 2As shown, a method for detecting and tracking vibrating fruits includes the following steps.

[0062] Orchard fruit image data (ripe peach fruits) were collected, the images were preprocessed using data enhancement, and a portion of the fruit images were labeled.

[0063] Orchard peaches were selected as the research object, and images of ripe peaches were used as the dataset for the fruit picking scene. The shooting equipment was a camera with a resolution of 3024×4032 pixels. To ensure the diversity of samples, the images were taken in the orchard under different weather conditions, including sunny days, cloudy days, mornings and afternoons. To enhance the robustness and generalization of the algorithm, the images were enhanced, including contrast enhancement factor set to 1.5, brightness enhancement factor set to 1.5, staggering, flipping, rotating, and adding Gaussian noise to increase the diversity of the dataset. Then, a portion of the data was randomly selected for image annotation. Labeling was used as the annotation tool, and the annotation information included the coordinate information of the key points of the ripe peach outline.

[0064] A multi-scale oscillating fruit detection network model was constructed based on an improved YOLOv8n. This model introduced the C2f-Dattention mechanism and employed a lightweight CAA-HSFPN network to reduce model size and improve the speed and accuracy of fruit detection in complex backgrounds. Furthermore, the SPPF-LSKA and CAA-HSFPN architectures were combined to further optimize multi-scale object detection performance.

[0065] like Figure 2The multi-scale oscillating fruit detection network model (CHD-YOLOv8n) includes four parts: input end, backbone network (Backbone), neck module (neck) and prediction end (segment); among them, the input end adopts mosaic data enhancement, adaptive anchor box calculation and adaptive grayscale filling; the backbone network adopts the first convolution module Conv, the second convolution module Conv, the first cross-stage partial fusion module C2f, the third convolution module Conv, the second cross-stage partial fusion module C2f, the fourth convolution module Conv, the third cross-stage partial fusion module C2f, the fifth convolution module Conv, the C2f-Dattention attention mechanism module and the multi-scale detection module SPPF-LSKA; the 10th, 13th, 15th, 20th and 22nd layers of the neck module all adopt CAA-HSFPN (Channel Attention Aggregation----HierarchicalSpatial Feature Pyramid) The prediction end uses two feature vectors of different scales at the 18th and 25th layers of the neck module for prediction results.

[0066] The C2f-Dattention mechanism is introduced to effectively capture multi-scale information and improve the model's ability to understand complex scenes, thereby reducing computational complexity and optimizing feature expression.

[0067] like Figure 4 Figure 2 shows the structure of the C2f-Dattention attention mechanism module. The C2f-Dattention attention mechanism module aims to enhance the network's ability to adaptively focus on different channel and spatial features, thereby improving the quality of feature representation. Improvements to this module include: First, the C2f (Cross-Channel Fusion) layer enables cross-channel fusion of information from different channels. By establishing information flow between different channels, the C2f layer promotes correlation between features, thereby improving the network's ability to capture diverse features. Next, the Dattention (Dynamic Attention) module incorporates a dynamic adaptive attention mechanism. By calculating the importance of each channel and spatial position and assigning different weights to each position, the network focuses more on key areas and important channels.

[0068] In this module, the features of different channels are first fused through the C2f layer to enhance the correlation between channels, and then each feature position is dynamically weighted through the Dattention mechanism. Figure 5As shown in the figure, the Dattention module uses dynamic learning to automatically adjust the attention weights for each spatial location and channel based on the input feature map, optimizing feature representation. Finally, the output is reduced in dimension through a 1×1 convolutional layer and batch normalization (BN). This output is then connected to the input features using a residual connection, preserving the original information while enhancing the representation of important features.

[0069] This C2f-Dattention attention mechanism effectively enhances the network's focus on key channels and spatial regions while avoiding interference from irrelevant information. By dynamically adjusting attention weights, the module can flexibly adapt to different input features, improving the network's adaptability and detection capabilities. The optimized C2f-Dattention module reduces redundant computation and improves network efficiency while maintaining efficient feature extraction, making it suitable for visual tasks in a variety of complex scenarios.

[0070] This patent introduces LSKA in SPP and proposes a new module SPPF-LSKA to further improve the model's feature expression and perception capabilities. Figure 6 As shown in the figure, this module captures contextual information and implements long-range dependencies by cascading depthwise convolutions and depthwise dilated convolutions with horizontal and vertical one-dimensional kernels. The fused feature map is then passed to a 1×1 convolution to generate an attention map. Finally, this attention map is multiplied with the input features to achieve adaptive feature refinement. Thanks to the design of the horizontal and vertical one-dimensional kernels, this process does not increase the model's computational complexity while significantly improving feature expression and perception capabilities.

[0071] Specifically, the improvements to this module include: First, multi-scale pooling processing is performed on the input image through the SPPF (SpatialPyramid Pooling) layer. The SPPF layer helps the network extract global information of targets of different sizes by pooling feature maps of different scales, giving the network a stronger perception of targets of different scales. Then, in the LSKA (Learnable Spatial Key Attention) module, the spatial key point attention mechanism (KeyAttention) is used to optimize the feature map. Through adaptive learning of spatial features, the LSKA module can intelligently focus on key areas based on the spatial structure of the input feature map, thereby improving the ability to focus on small targets and local features.

[0072] The output of LSKA is as follows:

[0073]

[0074] AC =W 1×1 *Z C

[0075]

[0076] Among them, F C is the input feature map, * represents convolution, d represents the expansion rate, represents the Hadamard operation, is the output of the depthwise convolution with kernel size (2d-1)×(2d-1), Z C The kernel size is The output of the depth-expanded convolution, A C The attention map is generated by 1×1 convolution. It's A C With F C The result of Hadamard operation.

[0077] In this module, multi-level feature information is first extracted through multi-scale pooling. Then, the LSKA mechanism is introduced to enhance the feature expression of important spatial locations through adaptive weight adjustment. Finally, the output dimension is reduced through a 1×1 convolutional layer and BN (Batch Normalization). The extracted multi-scale information is fused and then residual connections are performed to further enhance the network's expression capabilities. This improvement scheme not only improves the network's multi-scale target detection capabilities, but also introduces a learnable spatial attention mechanism, allowing the network to focus more on key areas and reduce interference from irrelevant features. The optimized SPPF-LSKA module effectively reduces the number of parameters while improving computational efficiency, making the detection process more efficient and adapting to target detection needs in different scenarios.

[0078] like Figure 7 The figure shows the CAA-HSFPN lightweight network structure. The CAA-HSFPN lightweight network structure mainly improves feature extraction accuracy and computational efficiency by combining the channel attention mechanism (CAA) and the hybrid spatial feature pyramid network (HSFPN). Improvements to CAA-HSFPN include: First, the CAA module generates channel descriptors through a global average pooling operation. Then, a lightweight fully connected layer is used to adjust the weight of each channel, introducing a channel attention mechanism to enhance the network's focus on key channel information and optimize the selection of important channels during feature extraction. Second, the HSFPN module uses a multi-scale feature pyramid structure for feature fusion. Combined with a hybrid spatial and channel attention mechanism, this enables the network to effectively extract multi-level features at different scales, improving the detection capability of large and small objects.

[0079] Next, in the network design, Depthwise Separable Convolution was used to further reduce the amount of computation, while the dimensionality of the features was reduced through a 1×1 convolution layer to optimize computational efficiency. In order to further improve the transmission and processing capabilities of feature information, CAA-HSFPN introduced a residual connection mechanism to fuse the optimized feature map with the input features, reducing information loss while improving network stability and training efficiency. This improvement effectively simplifies the original network structure and improves the network's computing speed and accuracy through lightweight design. Overall, the CAA-HSFPN lightweight network structure optimizes the utilization of computing resources while retaining high precision, and performs well on devices with limited computing resources. It is particularly suitable for low-power application scenarios such as mobile terminals and embedded systems.

[0080] The CAA-HSFPN network improves the adaptability of complex scenarios through PKINet and CAA mechanism. PKINet consists of four stages arranged in sequence. F is a simple feedforward network (FFN) taken from Then output The other path consists of a sequence of N PKI blocks that handles and produces The PKI block contains a PKI module and a CAA module. The input and output of phase 1 are F l-1 ∈R C1×H1×W1 and F l ∈R C1×H1×W1 The structure of stage 1 is as follows Figure 7 As shown, it represents a cross-stage part (CSP) structure. Specifically, after preliminary processing, the stage input F l-1 It is divided into two halves along the channel dimension and passed to two different paths. The specific expressions are as follows:

[0081]

[0082] Among them, Conv represents the convolution operation, DS represents the downsampling operation, and F l-1 Represents the input of stage 1, and Concat refers to the cascade operation.

[0083] A multi-scale oscillating fruit detection network model was constructed based on the DeepSORT algorithm. First, the model replaced the traditional CNN with a lightweight ResNest50 feature extraction network, improving the model's feature representation capabilities. Second, the traditional Kalman filter algorithm was replaced with an adaptive noise-scale Kalman filter, enhancing the model's adaptability and flexibility. Furthermore, CIoU matching replaced traditional IoU matching, improving the model's performance when handling objects of varying scales. These improvements significantly enhanced the robustness and flexibility of the object tracking system.

[0084] like Figure 8 This is a flowchart of the method for detecting and tracking oscillating fruit. DeepSORT is a multi-target tracking algorithm that combines deep learning and Kalman filtering. It detects targets and extracts appearance features, then uses Kalman filtering to predict target motion. It then uses a data association algorithm to match targets across frames, maintaining their unique IDs. This allows for accurate target tracking in complex scenes and effectively recovers tracking even when occluded or the target disappears. This model improves upon the DeepSORT algorithm. First, the method replaces the traditional detector with CHD-YOLOv8n and introduces a lightweight ResNest50 feature extraction network, replacing the traditional CNN network. This effectively improves the model's feature representation capabilities and computational efficiency. Second, the method replaces the traditional Kalman filter algorithm with an adaptive noise-scaling Kalman filter, significantly enhancing the model's adaptability and flexibility in complex environments, making it particularly stable in dynamically changing scenes. Furthermore, the method uses CIoU matching instead of traditional IoU matching, further optimizing the model's accuracy and performance when handling targets of varying scales. Through the above improvements, the present invention effectively improves the robustness and flexibility of the target tracking system, and significantly improves the stability and tracking accuracy of the system in a changing environment.

[0085] like Figure 9The figure below shows the lightweight feature extraction network structure of ResNest50. The lightweight feature extraction network structure of ResNest50 improves the performance and efficiency of the network by optimizing the traditional ResNet structure. First, ResNest50 improves on the basic residual block (ResidualBlock) and introduces the Split-Attention module. By splitting the input feature map into multiple groups for parallel processing, it improves the network's ability to capture information from different channels. Next, in each residual block, convolution operations are used to reduce the dimension of features, while BN normalization and ReLU activation functions are added to improve computational stability and nonlinear expression capabilities. In addition, depthwise separable convolution is introduced in the network structure to replace traditional standard convolution, significantly reducing the amount of computation and parameters and improving operational efficiency. Finally, the entire network transmits features through residual connections, ensuring information fluidity and avoiding the gradient vanishing problem. Through these improvements, ResNest50 can effectively extract multi-level features while maintaining low computational complexity, improving model performance and making the network structure more lightweight.

[0086] like Figure 10 Figure 2 shows a diagram of the CIoU matching strategy. The CIoU matching strategy is used for box matching in object detection tasks to improve the accuracy of bounding box regression. The improvement of this strategy is mainly reflected in the way it calculates the overlap between boxes. First, CIoU measures the degree of match between the predicted box and the ground-truth box by comprehensively considering the center point distance, aspect ratio difference, and the area of ​​the overlapping region. The overlap is calculated using IoU (Intersection over Union). However, CIoU further introduces penalties for center point distance and aspect ratio difference, enhancing the matching strategy's focus on box shape and position. Specifically, CIoU adds penalties for the Euclidean distance between center points and the aspect ratio difference between boxes on top of the IoU calculation. This design effectively addresses the shortcomings of traditional IoU when dealing with objects with large aspect ratio differences or large center point offsets. Then, during object detection training, by minimizing the CIoU loss, the network is forced to not only focus on region overlap but also optimize the shape and position of the target box, thereby improving detection accuracy and robustness. The improved CIoU matching strategy further optimizes the traditional IoU framework. By comprehensively considering position and shape, it improves the performance of the object detection network in complex scenarios. In particular, when the target box shape changes greatly, CIoU can more accurately guide the model to perform box regression.

[0087] The calculation formula of IoU is shown in formula (8):

[0088]

[0089] Where W i represents the width of the prediction box, and H i Indicates the height of the prediction box, S A Represents the area of ​​the prediction box, S B represents the area of ​​the ground-truth box, and W i ×H i It represents the intersection area of ​​the predicted box and the real box, and IoU is used to measure the ratio of their intersection area to their union area.

[0090] The present invention also employs an adaptive noise-scaling Kalman filter (ANSKF), which improves the state estimation performance in dynamic noise environments by improving the traditional Kalman filter. First, the ANSKF introduces a noise-scaling adaptive adjustment mechanism to adjust the weights of process noise and observation noise in real time, thereby adapting to the ever-changing noise environment. In this improved method, the residual of the input data is used as an indicator of noise changes. By monitoring the magnitude of the filter residual, the filter can dynamically adjust the noise scale and optimize the accuracy of the state estimation. Next, an adaptive algorithm is used to calculate the Kalman gain, allowing the filter to more accurately process the relationship between signal and noise as the noise intensity changes. The adaptive mechanism not only allows the Kalman gain to vary with noise but also improves the robustness of the estimation by balancing the effects of process noise and observation noise. Furthermore, this method simplifies the noise processing process of the traditional Kalman filter and, by reducing the computational complexity of the noise estimation, enables the filter to achieve faster calculation speed while maintaining high accuracy. Furthermore, the ANSKF optimizes the noise update strategy and can adaptively adjust the update rate to ensure the stability and performance of the filter in different noise environments. This improvement enables the Kalman filter to extract signals more accurately and reduce estimation errors when facing complex noise environments, thereby improving the adaptability and robustness of the filter in practical applications.

[0091] When using the Kalman filter algorithm for target tracking, there are two key steps: prediction and update. The prediction process relies on the target's motion state at the previous moment, using this information to infer the target's motion state at the current moment t. The update process uses the current observation value to correct the previous prediction state, thereby providing a more accurate target state estimate. By repeatedly executing the prediction and update cycle, the tracking system can continuously optimize the target state estimate. Each cycle is based on the latest observation data, allowing the system to adapt to changes in the target's motion in real time and maintain a high-precision state estimate. The specific mathematical expressions of the prediction and update process are shown below.

[0092] Prediction process:

[0093]

[0094] Update process:

[0095]

[0096] Where, is the state prediction value at time k, x k-1 is the estimated value of the system state at time k-1, u k-1 is the input variable at time k-1, F is the state transfer matrix, and B is the input gain matrix; is the Kalman prediction error covariance matrix when k, P k-1 is the Kalman estimation error covariance matrix at time k-1, F T is the transpose of the state transfer matrix F, Q is the process noise covariance matrix; y k is the observation margin, that is, the error between the observed value and the predicted value, z k is the observation value, H is the observation matrix, K k is the Kalman gain, which is used to estimate the importance of the error, H T is the transpose of the observation matrix H, R is the measurement noise covariance matrix, is the optimal state estimate at time k after the update, is the Kalman estimation error covariance matrix at time k after the update, and I is the identity matrix.

[0097] Construct a multi-scale oscillating fruit detection and tracking network model. Based on steps two and three, the network model of fruit detection and tracking is further integrated to achieve efficient detection and precise tracking of fruits in complex backgrounds. Specifically, this step first utilizes the multi-scale oscillating fruit detection network model constructed in step two. This model adopts the C2f-Dattention attention mechanism and the CAA-HSFPN lightweight network architecture, which can accurately detect fruits and extract features in complex environments. Then, the detection results are passed to the multi-scale oscillating fruit tracking network model designed in step three. This model uses the ResNest50 feature extraction network and the adaptive noise scale Kalman filter algorithm, combined with the CIoU matching method, to accurately track the detected fruit targets.

[0098] By fusing these two models, this step achieves real-time collaboration during fruit detection and tracking by sharing features and information flows. The position information provided by the fruit detection module directly serves as the initial input to the tracking module, ensuring that the fruit's initial position is accurately tracked. The output of the tracking module is further fed back to the detection module, helping the model correct detection errors in subsequent frames, thereby improving detection accuracy and enhancing the system's stability in dynamic environments. Through this seamless collaboration, this step significantly improves fruit detection and tracking performance in multi-scale and complex scenarios, ensuring efficient tracking of fruit targets at different scales and in diverse environments.

[0099] Use labeled data to train and validate a multi-scale vibrating fruit detection and tracking model.

[0100] Lightweight multi-scale oscillating fruit detection and tracking training, multi-scale oscillating fruit detection and tracking model is trained and verified using labeled data.

[0101] To jointly train the network end-to-end, the stochastic gradient descent (SGD) algorithm was employed. During training, all input images were uniformly resized to 640×640 pixels to improve the model's detection accuracy. The network parameters were optimized using the SGD optimizer, with an initial learning rate of 0.001, a weight decay rate of 0.005, and a momentum factor of 0.9. Furthermore, a validation epoch of 20 was set, meaning that the model's accuracy was evaluated on the validation set after every 20 iterations. The dataset was split into a training set and validation set ratio of 8:2. Training continued until the model's accuracy converged. After training, the final model was saved and validated using a test set of 800 images.

[0102] Through the trained multi-scale oscillation fruit detection and tracking model, the image of ripe peaches in the orchard is detected, and the coordinate information of the area where the fruit target is located in the orchard is saved, and the next movement trajectory of the fruit is predicted.

[0103] To verify the advantages of the CHD-YOLOv8n method proposed in this paper in fruit detection, it was compared with representative networks in related fields (such as YOLOv5n, YOLOv7-tiny, YOLOv8n, and YOLOv10n). The experimental results are shown in Table 1. The experimental results show that the CHD-YOLOv8n model proposed in this paper performs well in the fruit target detection task. Compared with YOLOv5n, YOLOv7-tiny, YOLOv8n, and YOLOv10n, CHD-YOLOv8n shows significant improvements in indicators such as precision, recall, and mAP@0.5. In addition, the model also has significant advantages in model parameters and storage size. Its model file is only 4.87MB, which is much smaller than other comparison models and is extremely suitable for edge device deployment. Although YOLOv10n is slightly faster in detection speed, CHD-YOLOv8n still performs well, with an average detection time of 3.55ms per image, showing high processing efficiency. In summary, CHD-YOLOv8n performs well in accuracy, speed, and miniaturization, and is particularly suitable for practical applications in resource-constrained environments.

[0104] Table 1 Comparison results of network model experiments

[0105]

[0106] To verify the effectiveness of the lightweight network used in this paper, we compared it with popular lightweight networks from the past two years, such as ConvNeXtV2, EfficientViT, SwinTransformer, RepViT, and MobileNetV4. The results are shown in Table 2. The CAA-HSFPN lightweight network exhibits excellent performance, particularly in precision, recall, and mAP@0.5, which are significantly improved compared to other mainstream networks. It also has a fast computation speed (3.07 milliseconds), making it suitable for efficient embedded systems and mobile device applications.

[0107] Table 2 Comparison results of lightweight network parameters

[0108]

[0109] To verify the effectiveness of the C2f-Dattention attention mechanism module used in this paper, comparative experiments were conducted with mainstream attention mechanism modules such as AKConv, ContextGuided, DCNV2, and DBB. The results are shown in Table 3. The experimental results show that C2f-Dattention significantly improves object detection performance, with precision improvements ranging from 0.53% to 3.95%. Recall also showed some improvement, with a 4.11% improvement in mAP@0.5. Furthermore, the C2f-Dattention model had an average computation time of 3.01 milliseconds, demonstrating high efficiency and significantly outperforming other methods overall. In summary, C2f-Dattention has significant advantages in performance, making it a reasonable choice to replace the C2f module in YOLOv8n with C2f-Dattention.

[0110] Table 3 Comparison results of different attention mechanisms

[0111]

[0112] To verify the effectiveness of the multi-scale object detection network used in this paper, we compared it with several mainstream multi-scale object detection networks (such as AIFI, BiFPN, RepHGNetV2, and RevCol). The experimental results are shown in Table 4. The experimental results show that SPPF-LSKA exhibits high efficiency and significantly outperforms other methods overall.

[0113] Table 4 Comparative experimental results of multi-scale object detection network

[0114]

[0115] To verify the performance improvement of the improved C2f-Dattention, CAA-HSFPN, and SPPF-LSKA modules for multi-scale fruit object detection in a harvested context, a series of ablation experiments were conducted. The results are shown in Table 5. The experimental results show that integrating the C2f-Dattention, CAA-HSFPN, and SPPF-LSKA modules significantly improves the model's precision and recall. Ultimately, the proposed CHD-YOLOv8n model achieves significant improvements in precision, recall, and mAP, with improvements ranging from 2.28% to 6.8% compared to the original YOLOv8n model. The model size is also reasonable, at approximately 4.87 MB. These experimental results validate the effectiveness of the proposed improved method for object detection tasks.

[0116] Table 5. Comparison results of ablation experiments

[0117]

[0118] To verify the improved tracking performance of the improved tracking model presented in this paper, the detectors were completely replaced with CHD-YOLOv8n and compared with SORT, DeepSORT, and the improved DeepSORT. The results are shown in Table 6. Experimental results show that compared to the traditional SORT algorithm, the DeepSORT algorithm improves the Mean Over Time (MOTA) and Mean Time Per Second (MOTP) metrics by 10.3% and 5.6%, respectively, while reducing the number of ID switches by 37.84%. This result demonstrates that the introduction of convolutional neural networks significantly improves the accuracy of DeepSORT in target tracking tasks. Further experiments show that the improved method based on DeepSORT improves the Mean Over Time (MOTA) and Mean Time Per Second (MOTP) by 6.2% and 11.3%, respectively, and significantly reduces the number of ID switches by 34.78%. These improvements not only further enhance tracking performance over DeepSORT but also effectively reduce the number of ID switches, significantly improving the stability and reliability of the tracking system. In summary, the improved algorithm proposed in this paper demonstrates greater efficiency and robustness when handling multi-target tracking tasks in complex scenarios.

[0119] Table 6 Comparison results of fruit tracking model experiments

[0120]

[0121] This study, focusing on ripe peaches in an orchard, aimed to test the accuracy of peach fruit recognition in real-world environments. To this end, a multi-scale oscillating fruit detection and tracking model was deployed on a JETSON AGX ORINCLB development kit running Ubuntu 20.04.6LTS. Experiments were conducted using the Python 3.8 programming language and PyTorch 1.14 environment. The model's performance in various scenarios was evaluated by analyzing a test set of 400 ripe peaches. This set of 400 images covered a variety of environmental conditions, including sunny and cloudy days, and at different time periods. To further enhance the model's performance in oscillating fruit tracking, a video dataset of 30 videos was collected under windy conditions, all of which were saved in .MP4 format.

[0122] Experimental results demonstrate that the improved CHD-YOLOv8n model significantly improves detection accuracy, particularly in complex backgrounds. The model's size is compressed to 4.88MB, making it suitable for efficient deployment on edge devices. Furthermore, the optimized DeepSORT algorithm effectively reduces ID switching by over 50%, significantly improving tracking stability in occlusion and oscillation scenarios. Data augmentation significantly enhances the model's generalization capabilities, further improving its adaptability in diverse environments.

[0123] Verification results show that the multi-scale oscillation fruit detection and tracking model performs well in peach fruit recognition and tracking, and can meet the requirements of real-time detection and tracking.

[0124] Deployment results on the JETSON AGX ORIN CLB development kit show that for a 3024×4032 pixel image, detection and processing take an average of 31ms, equivalent to processing 32.15 frames per second. This demonstrates that this model can effectively detect and track ripe peaches in real time on the JETSON AGX ORIN CLB, capable of completing target detection and tracking tasks in natural scenarios, and provides strong technical support for intelligent harvesting robots.

[0125] Based on the same inventive concept, the present invention also proposes a device for detecting and tracking vibrating fruits. The implementation of the device can refer to the implementation of the above method, and the repeated parts will not be repeated. Figure 11 As shown, the device 100 includes:

[0126] Fruit image acquisition module 101: used to collect orchard fruit image data and pre-process the images;

[0127] Detection model construction module 102: used to improve and construct a multi-scale oscillating fruit detection network model based on YOLOv8n;

[0128] Tracking model construction module 103: used to construct a multi-scale vibrating fruit detection network model based on the DeepSORT algorithm;

[0129] Detection and tracking model building module 104: used to build a multi-scale oscillating fruit detection and tracking network model;

[0130] Model training module 105: used to train the multi-scale vibrating fruit detection and tracking model using labeled data and verify it.

[0131] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0132] like Figure 12As shown, the device includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. An input / output (I / O) interface is also connected to the bus.

[0133] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0134] The processing unit performs the various methods and processes described above, such as method steps S01 to S05. For example, in some embodiments, method steps S01 to S05 can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more of the method steps S01 to S05 described above can be executed. Alternatively, in other embodiments, the CPU can be configured to execute method steps S01 to S05 by any other appropriate means (for example, by means of firmware).

[0135] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), and the like.

[0136] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0137] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0138] In addition, although adopting specific order to describe each operation, this should be understood as requiring such operation to be carried out in the specific order shown or in sequential order, or requiring all illustrated operations to be carried out to obtain desired result.Under certain environment, multitasking and parallel processing may be advantageous.Similarly, although comprising some specific implementation details in the above discussion, these should not be construed as limiting the scope of the present invention.Some features described in the context of independent embodiment can also be realized in single realization in combination.On the contrary, the various features described in the context of independent realization also can be realized in multiple realizations individually or in the mode of any suitable subcombination.

[0139] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A method for detecting and tracking vibrating fruits, characterized in that: The method includes: Step S01: collecting orchard fruit image data and preprocessing the image; Step S02: constructing a multi-scale oscillating fruit detection network model based on YOLOv8n improvement; Step S03: constructing a multi-scale vibrating fruit detection network model based on the DeepSORT algorithm; Step S04: constructing a multi-scale vibrating fruit detection and tracking network model; Step S05: Use the labeled data to train the multi-scale vibrating fruit detection and tracking model and verify it.

2. The method for detecting and tracking vibrating fruits according to claim 1, wherein: The multi-scale oscillating fruit detection network model described in step S02 is improved based on YOLOv8n, and includes: an input end, a backbone network, a neck module and a prediction end. The input end adopts mosaic data enhancement, adaptive anchor frame calculation and adaptive grayscale filling; the backbone network adopts the first convolution module Conv, the second convolution module Conv, the first cross-stage partial fusion module C2f, the third convolution module Conv, the second cross-stage partial fusion module C2f, the fourth convolution module Conv, the third cross-stage partial fusion module C2f, the fifth convolution module Conv, the C2f-Dattention attention mechanism module and the multi-scale detection module SPPF-LSKA, which are arranged in sequence; the 10th, 13th, 15th, 20th and 22nd layers of the neck module all adopt the CAA-HSFPN lightweight structure; the prediction end uses two feature vectors of different scales in the 18th and 25th layers of the neck module for prediction results.

3. The method for detecting and tracking vibrating fruits according to claim 2, wherein: The C2f-Dattention method includes: cross-fusion of information from different channels through the C2f layer; calculating the importance of each channel and spatial position in the Dattention module, assigning different weights to each position; and performing dimensionality reduction output through a 1×1 convolutional layer and batch normalization, and performing a residual connection with the input features.

4. The method for detecting and tracking vibrating fruits according to claim 2, wherein: The CAA-HSFPN lightweight structure includes: The CAA module generates channel descriptors through global average pooling operations; Use lightweight fully connected layers to adjust the weight of each channel and introduce a channel attention mechanism; The HSFPN module adopts a multi-scale feature pyramid structure for feature fusion and combines the hybrid attention mechanism of space and channel to extract multi-level features at different scales; Depthwise separable convolution is used to reduce the amount of computation, and the feature dimension is reduced through a 1×1 convolution layer.

5. The method for detecting and tracking vibrating fruits according to claim 2, wherein: The CAA-HSFPN lightweight structure also uses the PKINet structure to improve adaptability to complex scenarios.

6. The method for detecting and tracking vibrating fruits according to claim 1, wherein: The multi-scale oscillating fruit tracking network model described in step S03: a multi-scale oscillating fruit detection network model is used to replace the traditional detector; the ResNest50 feature extraction network is used to replace the CNN network, and the adaptive noise scale Kalman filter is used to replace the traditional Kalman filter algorithm; CIoU matching is used instead of IoU matching.

7. The method for detecting and tracking vibrating fruits according to claim 1, wherein: The multi-scale oscillating fruit detection and tracking network model described in step S04 includes: Detecting fruits and extracting features in a complex environment using the multi-scale oscillating fruit detection network model constructed in step S02; The detection result of the multi-scale oscillation fruit detection network model is passed to the multi-scale oscillation fruit tracking network model constructed in step S03 to track the detected fruit target; The multi-scale oscillating fruit tracking network module feeds the output results back to the multi-scale oscillating fruit detection network module for correcting the detection error in subsequent frames.

8. A device for detecting and tracking vibrating fruits, characterized in that: The device implements the method according to any one of claims 1 to 7, comprising: Fruit image acquisition module: used to collect orchard fruit image data and pre-process the images; Detection model construction module: used to improve and build a multi-scale oscillating fruit detection network model based on YOLOv8n; Tracking model construction module: used to build a multi-scale oscillating fruit detection network model based on the DeepSORT algorithm; Detection and tracking model building module: used to build a multi-scale oscillating fruit detection and tracking network model; Model training module: used to train and verify the multi-scale vibrating fruit detection and tracking model using labeled data.

9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Offshore multi-target tracking detection method for unmanned ship

    CN121305346A