Multi-target fruit tracking detection and counting method based on complex orchard

Through the improved YOLOv8n network model and dynamic Kalman filtering algorithm of variable forgetting factor, combined with the IoU-Re-ID data correlation method, the problems of low efficiency and poor accuracy of fruit detection, tracking and counting in the orchard are solved, and continuous tracking and accurate counting of fruits in the video sequence are achieved.

CN120182329APending Publication Date: 2025-06-20南宁桂电电子科技研究院有限公司 +1
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510260663.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The prior art has problems of low efficiency and poor accuracy in fruit detection, tracking and counting in orchards, especially when fruits are small, densely distributed and severely occluded.

Method used

The improved fruit detection model based on the YOLOv8n network model is adopted, and the dynamic Kalman filtering algorithm of variable forgetting factors and the IoU-Re-ID data association method are combined to achieve continuous tracking and accurate counting of fruits in the video sequence.

Benefits of technology

It significantly improves the accuracy of fruit detection and the robustness of the model, reduces noise accumulation and prediction errors, and realizes continuous tracking and accurate counting of fruits in video sequences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182329A_ABST
    Figure CN120182329A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-target fruit tracking detection and counting method based on a complex orchard. The method comprises the following steps: acquiring a to-be-detected video sequence; the to-be-detected video sequence is input to a fruit detection model, a detection result is obtained, the fruit detection model is constructed through an improved YOLOv8n network model and is obtained through training of a training set, and the training set is fruit image data; inputting the fruit image data and the detection result into a trajectory prediction model to obtain a prediction result of the fruit, the trajectory prediction model being obtained by introducing a dynamic Kalman filtering algorithm of a variable forgetting factor; and performing data association on the detection result and the prediction result to obtain fruit position information and a counting result. According to the invention, the prediction precision can be further improved, noise accumulation and prediction errors are reduced, and continuous tracking of fruits in a video sequence is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of computer vision recognition technology and multi-target tracking technology, and particularly relates to a multi-target fruit tracking detection and counting method based on a complex orchard. Background Art

[0002] With the continuous growth of the global population, fruits, as an indispensable part of human daily diet, have seen a sharp increase in demand. Against this background, the concept of sustainable agricultural development is particularly important, and effective monitoring of orchard growth conditions has become the key to making scientific management decisions. For this reason, fruit farmers and agronomists focus on the research and application of automatic fruit counting systems, not only requiring technological innovation to improve production efficiency, but also emphasizing the intelligent management of orchards, aiming to monitor and predict fruit yields in advance through real-time detection and improve the overall production efficiency.

[0003] Early fruit detection, tracking and synchronous counting methods mainly relied on manual methods, which consumed a large amount of manpower, had high costs and low efficiency. With the development of machine vision technology, fruit detection methods based on traditional image processing technology have gradually been applied to actual orchard management. These methods usually use techniques such as color segmentation and edge detection to identify fruit targets. For example, color segmentation is based on the typical color of fruits to distinguish the background from the target, which has good effects when the fruit color is relatively single, but performs poorly when the background is complex or the fruit is occluded.

[0004] With the rise of further machine learning technology, researchers began to try to use classifier models for fruit detection. Typical methods include support vector machine (SVM), random forest and K-nearest neighbor, etc. By extracting the features of fruit images, these classifiers can accurately identify fruit targets. However, machine learning methods rely on manual feature extraction and have poor adaptability to complex scenarios, making it difficult to cope with the complex and changeable environment in the orchard.

[0005] In the complex natural environment of the orchard, due to the diverse colors and shapes of fruits themselves, and being affected by factors such as weather, light and occlusion, the realization of the two fields of target detection and target tracking in intelligent fruit operations has been affected to a certain extent, which poses challenges to fruit recognition and yield estimation. For existing traditional fruit detection, tracking and counting methods, there are certain limitations when encountering the above problems, resulting in a lack of fast and accurate automated fruit counting methods. Therefore, to solve the problem of time-consuming and laborious manual counting in the orchard, it is of great significance to develop an accurate and reliable automatic fruit counting method based on machine vision. Summary of the Invention

[0006] To solve the above technical problems, the present invention proposes a multi-object fruit tracking, detection and counting method based on a complex orchard, which demonstrates the superiority of the improved version of the fruit detection network architecture model for the existing traditional fruit detection methods, where the detection model cannot correctly identify the target fruit when the fruits are small and densely distributed, and occlusion phenomena often occur, and even misidentification may occur; in terms of tracking performance, it can significantly reduce the situation where the ID of the target may be discontinuous and jump when the fruit moves rapidly or interacts with other objects in the video frame, further improving the prediction accuracy, reducing the accumulation of noise and prediction errors, and achieving continuous tracking of the fruit in the video sequence.

[0007] The present invention provides a multi-object fruit tracking, detection and counting method based on a complex orchard, including:

[0008] Obtain the video sequence to be detected;

[0009] Input the video sequence to be detected into the fruit detection model to obtain the detection result, where the fruit detection model is constructed by an improved YOLOv8n network model and obtained through training with a training set, and the training set is fruit image data;

[0010] Input the fruit image data and the detection result into the trajectory prediction model to obtain the prediction result of the fruit, where the trajectory prediction model is obtained by introducing a dynamic Kalman filtering algorithm with a variable forgetting factor;

[0011] Perform data association on the detection result and the prediction result to obtain the fruit position information and the counting result.

[0012] Optionally, obtaining the training set includes:

[0013] Collect fruit images and fruit videos;

[0014] Perform frame splitting on the fruit video to obtain the split images;

[0015] Perform annotation processing on the split images and the fruit images to obtain the annotated images;

[0016] Amplify the annotated images to obtain the training set.

[0017] Optionally, the improved YOLOv8n network model includes:

[0018] Replace the original backbone network with the EfficientNetB0 network, optimize the detection head structure of the original YOLOv8n network model, and introduce a multi-scale dilated attention mechanism to obtain the improved YOLOv8n network model.

[0019] Optionally, inputting the fruit image data and the detection result into a trajectory prediction model to obtain the prediction result of the fruit includes:

[0020] Based on the fruit image data and the detection result, obtaining the position information of the fruit;

[0021] According to the position information of the fruit, estimating the position and velocity of the target fruit in the current frame;

[0022] Correcting the position and velocity in the current frame to obtain the prediction result.

[0023] Optionally, estimating the position and velocity of the target fruit in the current frame includes:

[0024]

[0025] where X t is the state vector at the previous moment, F is a state transition matrix of size 8×8, P t is the covariance matrix at the previous moment, ω is a random vector of process noise, Q is the process noise covariance matrix, is the predicted state vector at the current moment, is the predicted covariance matrix at the current moment, f(x t , w) is the state transition function, f(x t ) is the deterministic part of the state transition function, N(0,Q) is the Gaussian distribution, F T is the transpose of the state transition matrix F.

[0026] Optionally, updating the position and velocity in the current frame includes:

[0027] z = h(x) + r, r ~ N(0,R)

[0028]

[0029] where z is the observation value, r is the observation noise, h(x) is the observation function, K is the Kalman gain, H is the observation matrix, λ t is the variable forgetting factor, X t+1 is the updated state vector, P t+1 is the updated covariance matrix.

[0030] Optionally, performing data association on the detection result and the prediction result to obtain the fruit position information and the counting result includes:

[0031] Extracting the shallow appearance feature and the depth appearance feature of the detection result;

[0032] Matching the depth appearance feature and the prediction result to obtain the first association matching result;

[0033] Match the prediction results that do not match successfully with the shallow appearance features to obtain the second associated matching results;

[0034] Obtain the final association matrix according to the first associated matching result and the second associated matching result;

[0035] Obtain the fruit position information and the counting result according to the association matrix.

[0036] Optionally, obtaining the final association matrix according to the first associated matching result and the second associated matching result includes:

[0037]

[0038] where d IoU (A, B) is the IoU distance, representing the gap between the detection result and the bounding box of the tracking target, d ReID (a, b) is the embedding feature distance, representing the difference between the detection result and the appearance feature of the tracking target, d fused is the fused IoU distance and embedding feature distance, used for the final matching decision, A is the bounding box of the detection result, B is the predicted bounding box of the tracking target, a is the appearance feature vector of the detection result, and b is the appearance feature vector of the tracking target.

[0039] Compared with the prior art, the present invention has the following advantages and technical effects:

[0040] (1) The present invention embeds the idea module of small-sample object detection into the YOLO network to better adapt to problems such as small targets and severe occlusions in the fruit detection scenario, and improve the detection effect and the robustness of the model.

[0041] (2) The present invention proposes a counting method based on video sequences to solve the problems of single perspective and insufficient real-time performance in single-image detection; based on the BotSORT object tracking algorithm, the traditional Kalman filtering algorithm is improved to a Kalman filtering algorithm with a variable forgetting factor to achieve continuous tracking of fruits in the video sequence, thereby improving the accuracy and real-time performance of fruit tracking.

[0042] (3) The present invention uses the IoU-Re-ID intersection over union and re-identification data association method for matching, and ensures the coherence of tracking by establishing the association between frames in the video sequence, thereby improving the accuracy of counting. Description of the Drawings

[0043] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0044] Figure 1 It is a flowchart of a multi - target fruit tracking, detection and counting method based on a complex orchard in an embodiment of the present invention;

[0045] Figure 2 It is the overall experimental flowchart of an embodiment of the present invention;

[0046] Figure 3 It is a schematic diagram of data set collection and construction in an embodiment of the present invention. Among them, (a) is a picture of the picture data set, (b) is a picture of the Synthetic - apples video data set, (c) is a picture of the annotation process of the video data set, and (d) is a picture of data augmentation;

[0047] Figure 4 It is the network structure diagram of the improved YOLOv8n model in an embodiment of the present invention;

[0048] Figure 5 It is the flowchart of the Dynamic KalmanFilte system in an embodiment of the present invention;

[0049] Figure 6 It is the flowchart of the Dynamic KalmanFilte Tracker algorithm in an embodiment of the present invention;

[0050] Figure 7 It is the trend chart of tracking, detection and real - time counting of different methods in video frames in an embodiment of the present invention;

[0051] Figure 8 It is the different tracking error situations in an embodiment of the present invention. Among them, (a) is a picture with a huge jump in ID, (b) is a picture with video frame loss - low - score detection and the detection ID not matching the real ID, and (c) is a picture with the ID changed due to occlusion. Detailed implementation manners

[0052] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.

[0053] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer - executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0054] This embodiment proposes a multi - target fruit tracking, detection and counting method based on a complex orchard, as Figure 1-2 shown, and specifically includes the following steps:

[0055] Obtain the video sequence to be detected;

[0056] Input the video sequence to be detected into the fruit detection model to obtain the detection result. Among them, the fruit detection model is constructed through an improved YOLOv8n network model and obtained through training with a training set, and the training set is fruit image data;

[0057] Input the fruit image data and the detection result into the trajectory prediction model to obtain the prediction result of the fruit. Among them, the trajectory prediction model is obtained through the dynamic Kalman filtering algorithm introducing a variable forgetting factor;

[0058] Perform data association on the detection result and the prediction result to obtain the fruit position information and the counting result.

[0059] Specifically, step 1: Collect, annotate, divide, and perform data augmentation on the dataset of fruit images and videos in a complex orchard environment to complete the collection and construction of the dataset;

[0060] Step 2: Based on the YOLOv8n network model structure, replace the backbone network with the EfficientNetB0 module, innovatively design the ASFFHead four-head structure for the second time, and introduce the MSDA multi-scale dilated attention mechanism method to construct an improved fruit detection model, and perform frame-by-frame detection of the fruits in the video sequence through the object detection model;

[0061] Step 3: Extract the appearance position and motion features of the fruit, use the improved Dynamic Kalman Filter to dynamically predict the future state of the target, and construct a feature extraction and state prediction model;

[0062] Step 4: Based on the object detection result and state prediction, perform object matching to construct the IoU-Re-ID data association method;

[0063] Step 5: As Figure 8 shown in (a)-(c), finally combine the continuous tracking results to count the number of fruits and perform fruit counting.

[0064] Furthermore, obtaining the training set includes:

[0065] Collect fruit images and fruit videos;

[0066] Perform frame splitting on the fruit video to obtain the split images;

[0067] Perform annotation processing on the split images and fruit images to obtain annotated images;

[0068] Perform amplification on the annotated images to obtain the training set.

[0069] Specifically, step 1 specifically includes:

[0070] Use the fruit pictures in the orchard as the dataset of the detection model, which meets challenging factors such as small targets, occlusion, dense targets, and complex natural environments. Label them with the LabelImg annotation tool, and select a video dataset with a duration of 24s and a frame rate of 10f / s, a total of 250 frames containing 5 trees. Manually annotate the bounding boxes and IDs with the DarkLabel tool and save them in the MOT data format. To enhance data diversity, perform data preprocessing such as flipping, rotating, scaling, adding noise, and enhancing color contrast to improve the robustness and adaptability of the fruit target detection model to different actual application scenarios.

[0071] Furthermore, the improved YOLOv8n network model includes:

[0072] Replace the original backbone network with the EfficientNetB0 network, optimize the detection head structure of the original YOLOv8n network model, and introduce a multi-scale dilated attention mechanism to obtain the improved YOLOv8n network model.

[0073] Specifically, in step 2, replace the feature extraction module, optimize the detection head, and fuse the attention mechanism. The specific construction of the improved fruit target detection model includes:

[0074] (1) The EfficientNetB0 network is an efficient convolutional neural network architecture that takes into account the complexity of the network during design. Through a compound scaling method, its compound scaling method is:

[0075]

[0076] In the formula, d, ω, and r respectively represent the depth, width, and resolution of the network, Φ represents the compound scaling coefficient, and α, β, and γ represent the corresponding scaling bases;

[0077] By uniformly scaling and adjusting the depth, width, and resolution of the network, the performance of the convolutional neural network can be improved, enabling better handling of data changes in various scenarios, having better feature extraction capabilities, and being able to better detect and identify fruits in different scales, densities, and occlusion situations;

[0078] EfficientB0 is the smallest model in the EfficientNet series. First, an image with an input size of 224×224×3 is passed through a Conv3×3 dimensionality increase operation to obtain a feature map of 112×112×32. Then, a series of MBConv (Mobile inverted Bottleneck Convolution) modules are used to process the feature map to obtain a feature map of 7×7×320. Finally, the results are output using Conv1×1, Pooling, and a Fully connected (FC) layer. Compared with traditional convolutional neural networks, the MBConv module mainly consists of a 1×1 ordinary convolution for dimensionality increase, a 3×3 Depthwise Conv, an SE module attention mechanism, a 1×1 ordinary convolution for dimensionality reduction, and a Dropout layer, reducing the number of model parameters and computational complexity and having a fast running speed. This design of EfficientNetB0 not only considers the accuracy of the model but also has higher computational efficiency and is a lightweight model, which is particularly important for real-time fruit detection;

[0079] (2) In a complex orchard counting environment, problems such as obvious changes in fruit tree density, severe fruit occlusion, and uneven distribution in images are often faced. To overcome these problems, the MSDA module is introduced. The MSDA module (multi-scale dilated attention mechanism) is an improvement derived from the analysis of the Vision Transformer (ViT) model, and it is found that there is a contradiction between computational complexity and receptive field size when dealing with the global attention mechanism; therefore, the MSDA multi-scale dilated attention module is designed to simulate local and sparse interactions at different scales. By parallelly stacking dilated convolutional layers with different dilation rates (r = 1, r = 2, r = 3), the ability to capture features of targets at different scales is enhanced, enriching the information of the feature map; in this way, the model can better identify occluded apples and effectively distinguish different fruits in dense areas, optimizing the apple counting and tracking effects in the entire orchard environment;

[0080] (3) The core idea of the FASFF module is to adaptively fuse feature information at different levels, adaptively learn the low-level features with less semantic information and the high-level features with less detailed information, improving the network's detection ability for apples of different sizes, especially small target apples in the distance; and by integrating multi-level features, better handle partially occluded apples, distinguish target apples from complex backgrounds such as branches or other fruits; specifically, taking FASFF1 as an example, first adjust feature maps of different scales (X1, X2, X3, X4) to the same size (X 1→1 、X 2→1 、X 3→1 、X 4→1), and set trainable weight parameters (α1, β1, γ1, λ1) for each layer of feature maps. Finally, feature map fusion is performed according to these weight parameters, and its calculation formula is:

[0081] F1 = α1X 1→1 + β1X 2→1 + Υ1X 3→1 + λ1X 4→1

[0082] In the formula, F1 is the newly fused feature output, a1, b1, g1, and l1 are weight parameters respectively, and X 1_>1 、X 2_>1 、X 3_>1 、X 4_>1 respectively represent the feature maps obtained by adjusting the feature maps of each layer to the same size as the X1 feature map through sampling;

[0083] In YOLOv8, the detection head is a key part of the model. Its function is to directly process the feature maps to generate object detection results. By default, it contains three detection layers (P3, P4, P5), and each layer is specifically responsible for object recognition at different scales. Although this design performs well in most cases, it still faces certain challenges when dealing with small objects. Especially in a complex orchard environment, apples may be partially blocked by leaves or other fruits, or it is difficult to identify due to changes in lighting conditions, which affects its detection performance. Secondly, apples are in different growth stages and have different sizes. During the detection process, apples of different scales will appear. Therefore, a detection network with better feature fusion ability is needed;

[0084] To solve the inconsistency problem between the feature maps of target fruits at different scales, the ASFF (Adaptive Spatial Feature Fusion) module is introduced into the detection head of YOLOv8, and on this basis, a secondary innovation is carried out. A fourth output layer is added to change the traditional three detection heads into four, forming the FASFF module, reducing feature conflicts, solving the problem of feature loss due to cross-scale fusion, and increasing the secondary extraction of the small object detection layer;

[0085] Through these improvements, the robustness of the detection model has been significantly enhanced, enabling it to better adapt to changing environmental conditions, thereby improving the accuracy and reliability of fruit detection and providing more accurate state position information for subsequent tracking.

[0086] Furthermore, the fruit image data and detection results are input into the trajectory prediction model, and the predicted results of the fruits obtained include:

[0087] Based on the fruit image data and detection results, obtain the position information of the fruit;

[0088] According to the position information of the fruit, estimate the position and speed of the target fruit in the current frame;

[0089] Correct the position and velocity of the current frame to obtain the prediction result.

[0090] Specifically, in step 3, extract the appearance position and motion characteristics of the fruit, and use the improved Dynamic Kalman Filter to dynamically predict the future state of the target, and construct a feature extraction and state prediction model. The process is as follows:

[0091] (1) In the initialization stage, use the object detection algorithm YOLO to detect the fruit in the image of the t-th frame, obtain the position information of the fruit, including coordinates, class labels, and confidence scores, and define the initial state matrix X t and the initial covariance matrix P t , as shown in the following formula:

[0092] X t =[p t , v t T

[0093]

[0094] In the formula, the state matrix X t is composed of the position p t of the detection frame and the velocity state variable v t . They respectively represent the uncertainty of measurement, and their calculations depend on the initial measurement noise weight w. Usually, w p ≤ w v . The covariance matrix P t is initialized as a diagonal matrix, where the elements on the diagonal are the variances (the squares of the standard deviations) of each variable, and the non-diagonal terms are 0, indicating that the variables are initially uncorrelated. I 4×4 represents the 4×4 identity matrix;

[0095] (2) In the prediction stage, use the state vector X t and covariance matrix P t of the target at the previous moment to estimate the position and velocity of the target in the current frame through the state transition model. The calculation formula is:

[0096]

[0097] X t+1 = f(x t , w) = f(x t ) + w, w ~ N(0, Q)

[0098] ​Among them, F is a state transition matrix of size 8×8, used for state update in the prediction step, and Q is the process noise covariance matrix, which is used to represent the error introduced by the uncertainty of motion during prediction, and its definition is based on the standard deviations of position and velocity;

[0099] (3) In the update stage of the traditional Kalman filter algorithm, it is necessary to combine the detected observation information to correct the predicted state matrix and covariance matrix .

[0100] By calculating the Kalman gain and using the innovation (i.e., the difference between the predicted state and the actual observation), the updated state vector and covariance matrix can more accurately reflect the actual position and motion state of the target. This correction process combines the motion information of the prediction model and the measurement information of the observation model, ensuring the tracking accuracy and adaptability to the dynamic changes of the target. The calculation formulas are as follows:

[0101] Kalman gain:

[0102] State update:

[0103] Covariance update:

[0104] (4) To improve the self - adaptability of the traditional Kalman filter, a variable forgetting factor λ t is introduced to dynamically adjust the covariance update process. The modified covariance update equation is:

[0105]

[0106] Among them, λ t ∈[0,1] is the forgetting factor, P t is the covariance matrix after filtering in the previous frame. By dynamically adjusting λ t , it is possible to more flexibly handle noise changes;

[0107] Variable forgetting factor: λ t =λ min +(λ max -λ min )·exp(-γ·||v t || 2 );

[0108] Square norm of the measurement residual: ||v t || 2 =(z t -HX t - ) T (z t -HXt - );

[0109] Furthermore, data association is performed on the detection results and the prediction results to obtain the fruit position information and the counting result, including:

[0110] Extract the shallow appearance features and the depth appearance features of the detection results;

[0111] Match the depth appearance features with the prediction results to obtain the first association matching result;

[0112] Match the prediction results with unsuccessful matches with the shallow appearance features to obtain the second association matching result;

[0113] Obtain the final association matrix according to the first association matching result and the second association matching result;

[0114] Obtain the fruit position information and the counting result according to the association matrix.

[0115] Specifically, in step 4, target matching is performed based on the object detection results and the state prediction, and an IoU-Re-ID data association method is constructed, specifically including:

[0116] Use the FastReID library to extract the Re-ID features of the target, and combine the exponential moving average (EMA) method to smooth the features, and update the average appearance feature state of the tracklet in real time. The calculation formula is:

[0117]

[0118] where α is the smoothing coefficient (0 < α < 1), is the smoothed feature, f t+1 is the feature at the current moment;

[0119] On this basis, combine IoU and Re-ID features, calculate their feature similarities respectively to capture the motion and appearance information of the target; secondly, filter out the IoU matching candidates with low correlation through Re-ID; finally, fuse the IoU and Re-ID features based on the minimum value rule to generate the final association matrix. The calculation formula is:

[0120]

[0121] This strategy can improve the accuracy of the target ID by optimizing simultaneously in both the spatial and appearance dimensions, effectively improving the MOTA multi-object tracking accuracy index and the IDF1 index, and thus enhancing the overall performance of target tracking.

[0122] In Step 5, the ImprovedYOLO improved object detection algorithm is combined with the Dynamic KF Tracker dynamic Kalman filter, significantly enhancing the overall performance of fruit detection and tracking; by optimizing the network structure and introducing an improved Kalman filter tracking algorithm with a variable forgetting factor, the accurate detection, continuous tracking, and precise counting of fruits are effectively achieved.

[0123] The following elaborates on this embodiment in conjunction with the attached drawings:

[0124] As Figure 1 shown, this embodiment provides a multi-object fruit tracking, detection, and counting method based on a complex orchard, including the following steps:

[0125] Step 1: Collect, annotate, partition, and perform data augmentation on the dataset of fruit images and videos in a complex orchard environment to complete the collection and construction of the dataset;

[0126] Step 2: Based on the YOLOv8n network model structure, replace the backbone network with the EfficientNetB0 module, innovatively design the ASFFHead four-head structure for the second time, and introduce the MSDA multi-scale dilated attention mechanism method to construct an improved fruit detection model, and perform frame-by-frame detection of fruits in the video sequence through the object detection model;

[0127] Step 3: Extract the appearance position and motion features of the fruits, and use the improved Dynamic Kalman Filter to dynamically predict the future state of the object to construct a feature extraction and state prediction model;

[0128] Step 4: Based on the object detection results and state prediction, perform object matching to construct an IoU-Re-ID data association method;

[0129] Step 5: Finally, combine the continuous tracking results to count the number of fruits for fruit counting.

[0130] As Figure 3 (a)-(d) shown, collect, annotate, partition, and perform data augmentation on the dataset of fruit images and videos in a complex orchard environment to complete the collection and construction of the dataset.

[0131] In this embodiment, the apple dataset in the detection part contains 7,311 images, which are annotated through LabelImg and randomly divided into a training set, a validation set, and a test set according to a ratio of 7:2:1; the video dataset has a duration of 24 s, a frame rate of 10 f / s, and a total of 250 frames including 5 trees, which are annotated using Darklabel;

[0132] To simulate different lighting conditions and angular changes that may occur in an orchard environment, the original data is preprocessed, including flipping, rotating, scaling, adding noise, enhancing color contrast, etc., which helps the model generalization and improves the model robustness.

[0133] As Figure 4 shown, as one of the latest versions of the YOLO series, YOLOv8 performs excellently in object detection and complex scenarios. Especially in the real-time fruit monitoring task, it can better balance the detection speed and accuracy. Therefore, YOLOv8n is selected as the baseline network model for apple detection tasks in complex orchard environments.

[0134] (1) Replace the backbone network with the EfficientNetB0 module. By using a new model scaling method, network expansion, enhanced feature extraction, and lightweight are achieved.

[0135] (2) Optimize on the ASFFHead detection to form a cross-scale fusion module for secondary extraction, increasing the detection accuracy of small target fruits.

[0136] (3) Introduce the MSDA multi-scale dilated attention module to simulate local and sparse interactions at different scales. By parallelly stacking dilated convolutional layers with different dilation rates, the feature capture ability for apples of different scales is enhanced.

[0137] The above improvements enable the model to have better feature extraction and feature fusion capabilities, optimizing the apple detection effect in the entire orchard environment.

[0138] To evaluate the performance of the detection model, metrics such as mean average precision (mAP), computational volume (GFLOPs), FPS (Frames Per Second), etc. are used as the measurement indicators for the model. These indicators comprehensively evaluate the accuracy, spatial complexity, computational performance, and real-time performance of the model. The calculation formulas are as follows:

[0139]

[0140] To verify the performance of the improved fruit detection model, the experimental results are shown in Table 1:

[0141] Table 1

[0142]

[0143]

[0144] It can be seen from the table that the key performance indicators such as the accuracy rate (P), mean average precision (mAP), and GFLOPs of the method with improved mean average precision are all better than other methods.

[0145] AsFigure 5-6 As shown in Figure 5-6 , the Dynamic KF Tracker has made multiple optimizations in motion prediction and appearance association. It uses the Kalman filtering algorithm with a variable forgetting factor to predict the motion trajectory of the fruit, updates it in combination with the new detection box, and proposes Camera Motion Compensation (CMC) to handle the motion interference of the camera, ensuring the continuity of tracking and making the prediction more accurate. Finally, it associates with the target fruit through IoU and ReID to complete the update of the trajectory.

[0146] To analyze the performance of the fruit tracking algorithm, the tracking performance of the model is comprehensively evaluated in combination with multiple metrics. Among them, the three relatively important metrics for this embodiment include MOTA (Multiple Object Tracking Accuracy), IDF1 (Identification F1 Score), and HOTA (Higher Order Tracking Accuracy).

[0147]

[0148] To improve the tracking accuracy, this embodiment adopts a counting strategy that combines YOLOv8n and the improved YOLO as object detection models with multiple tracking algorithms to evaluate the performance of different methods in fruit tracking. The experimental design involves four common tracking algorithms: SORT, DeepSORT, ByteTrack, and BoTSORT, and compares them with the Dynamic Kalman Filter Tracker (Dynamic KF Tracker) proposed in this embodiment. The tracking detection and real-time counting trend graphs of different methods in video frames are as Figure 7 shown, and the experimental data are shown in Table 2:

[0149] Table 2

[0150]

[0151] In this embodiment, through the comparison of experimental results, it can be seen that in the face of the obvious limitations of traditional deep learning-based object detection models in complex orchard environments, they can only effectively analyze static single-frame images, are difficult to cope with the actual dynamic changes in orchards, and cannot meet the needs of large-scale orchard detection and counting. The proposed dynamic Kalman filter tracker is optimized on the basis of the ImprovedYOLO detector, introducing a Kalman filter with a variable forgetting factor, and combining IoU and Re-ID features to optimize the target matching accuracy, effectively improving the tracking accuracy and stability of the target. Experimental results show that this scheme exceeds other algorithms in performance indicators such as MOTA (95.0%), IDF1 (65.5%), and HOTA (82.4%), showing the best performance. By paying attention to the overall trend and morphological changes of the curves of each method and the ground truth, Method9 fits well with the ground truth, showing superior detection ability in most frames and maintaining high tracking stability.

[0152] To comprehensively evaluate the performance of these methods, on the basis of the object tracking algorithm, two metrics, R 2 (coefficient of determination) and RMSE (root mean square error), are used to further compare the performance of the 3 methods with better tracking effects in the counting task. The calculation formulas are as follows:

[0153]

[0154] In this embodiment, their respective linear regression equations are calculated, as shown in Table 3:

[0155] Table 3

[0156]

[0157] Experimental results show that the two methods based on BotSORT have poor tracking effects. Some fruits are occluded in the 40th frame, and when they are detected again in the following frames, the ID changes due to tracking errors. Therefore, their correlation coefficients R 2 are both small, 0.72 and 0.78 respectively; the method proposed in this embodiment performs best in this object counting task, with the highest R 2 : 0.85, and the RMSE is 1.57, indicating that the method has a high degree of fitting between the predicted fruit ID and the actual number of manually counted fruits, with a small error when estimating the number of target fruits. The overall effect is significantly higher than the first two methods, and it can better maintain the continuity of the target in a complex orchard environment, thereby improving the counting accuracy.

[0158] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-target fruit tracking, detection and counting method based on a complex orchard, characterized in that: include: Obtain a video sequence to be detected; Inputting the video sequence to be detected into a fruit detection model to obtain a detection result, wherein the fruit detection model is constructed by an improved YOLOv8n network model and obtained by training with a training set, and the training set is fruit image data; Inputting the fruit image data and the detection result into a trajectory prediction model to obtain a prediction result of the fruit, wherein the trajectory prediction model is obtained by introducing a dynamic Kalman filter algorithm with a variable forgetting factor; The detection result and the prediction result are data-associated to obtain the fruit position information and the counting result.

2. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 1, characterized in that: Acquiring the training set includes: Collect fruit images and fruit videos; Performing frame processing on the fruit video to obtain frame-divided images; Annotating the frame-processed image and the fruit image to obtain annotated images; The annotated image is amplified to obtain the training set.

3. The multi-target fruit tracking, detection and counting method based on a complex orchard according to claim 1 is characterized in that: The improved YOLOv8n network model includes: The original backbone network is replaced with the EfficientNetB0 network, the detection head structure of the original YOLOv8n network model is optimized, and a multi-scale hole attention mechanism is introduced to obtain an improved YOLOv8n network model.

4. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 1, characterized in that: Inputting the fruit image data and the detection result into the trajectory prediction model to obtain the prediction result of the fruit includes: Based on the fruit image data and the detection result, obtaining the position information of the fruit; According to the position information of the fruit, estimating the position and speed of the target fruit in the current frame; The position and speed of the current frame are corrected to obtain the prediction result.

5. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 4, characterized in that: Estimating the position and speed of the target fruit in the current frame includes: Among them, X t is the state vector at the previous moment, F is the state transfer matrix of size 8×8, P t is the covariance matrix of the previous moment, ω is a random vector of process noise, Q is the process noise covariance matrix, is the state vector predicted at the current moment, is the covariance matrix predicted at the current moment, f(x t , w) is the state transfer function, f(x t ) is the deterministic part of the state transfer function, N(0,Q) is a Gaussian distribution, F T is the transpose of the state transfer matrix F.

6. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 5, characterized in that: Updating the position and speed of the current frame includes: z=h(x)+r,r~N(0,R) Where z is the observation value, r is the observation noise, h(x) is the observation function, K is the Kalman gain, H is the observation matrix, and λ t is the variable forgetting factor, X t+1 is the updated state vector, P t+1 is the updated covariance matrix.

7. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 1, characterized in that: Performing data association between the detection result and the prediction result to obtain the fruit position information and the counting result includes: Extracting shallow appearance features and deep appearance features of the detection result; Matching the deep appearance feature with the prediction result to obtain a first associated matching result; Match the unsuccessful prediction results with the shallow appearance features to obtain the second associated matching results; Obtaining a final correlation matrix according to the first correlation matching result and the second correlation matching result; According to the association matrix, the fruit position information and counting results are obtained.

8. The method for tracking, detecting and counting multiple fruits in a complex orchard according to claim 7, characterized in that: According to the first association matching result and the second association matching result, obtaining a final association matrix includes: Among them, d IoU (A, B) is the IoU distance, which indicates the gap between the detection result and the bounding box of the tracked target, d ReID (a, b) is the embedding feature distance, which indicates the difference between the detection result and the appearance feature of the tracked target. fused In order to fuse the IoU distance and the embedded feature distance for the final matching decision, A is the bounding box of the detection result, B is the predicted bounding box of the tracked target, a is the appearance feature vector of the detection result, and b is the appearance feature vector of the tracked target.

Citation Information

Cited By

  • Diagnosis method and system for glenoid cavity and humeral head defect area

    CN120809145A

  • Multi-target tracking-based sturgeon residual feed dynamic counting method

    CN120823483A

  • Multi-target tracking method and system based on non-appearance chain trajectory correlation model

    CN120912642A

  • Buoy safety visual detection method and system based on deep learning

    CN121214363A

  • Deep Learning-Based Visual Detection Method and System for Buoy Safety

    CN121214363B