Method for counting xanthoceras sorbifolia bunge fruits based on unmanned aerial vehicle video
By improving the YOLOv11 backbone network and PAN-FPN feature fusion network, and combining the SEAM attention mechanism and adaptive Kalman filter algorithm, the problems of occlusion and environmental changes in UAV video fruit counting are solved, improving detection accuracy and stability, and achieving efficient fruit quantity estimation.
Patent Information
- Application Number
- CN202511011618.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-11-07
AI Technical Summary
Existing drone video fruit counting methods lack datasets in real-world scenarios, resulting in insufficient detection accuracy, especially for fruits with occlusion or complex backgrounds. Furthermore, camera movement and lighting changes lead to tracking loss and frequent ID switching.
An improved YOLOv11 backbone network and PAN-FPN feature fusion network are adopted, combined with the SEAM attention mechanism and adaptive Kalman filter algorithm. Through feature extraction, feature fusion and detection head network, high-precision localization and tracking of Xanthoceras sorbifolium fruits are achieved. The IoU-ReID fusion mechanism is used for matching to improve occlusion detection and environmental adaptability.
This improves the detection accuracy and stability of fruit counting in drone video, reduces errors caused by occlusion and environmental changes, and achieves more efficient fruit quantity estimation.
Smart Images

Figure CN120913067A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a shiitake fruit counting method based on unmanned aerial vehicle video. BACKGROUND
[0002] Yield estimation is a key link in shiitake production management. Traditional shiitake fruit counting mainly relies on manual sampling estimation, which estimates the total yield by multiplying the number of fruits on a number of randomly selected shiitake trees by the total number of fruit trees in the orchard. This method is not only time-consuming and inefficient, but also because of the significant difference in fruit number between individual trees, the selected samples often fail to adequately represent the overall situation, resulting in insufficient estimation accuracy, which in turn affects decision-making quality and causes economic losses.
[0003] In recent years, the development of artificial intelligence has driven the rapid rise of precision agriculture. Intelligent yield estimation technology based on computer vision has developed into one of the key areas of modern agricultural research. Researchers have developed fruit counting methods based on deep learning, which achieve counting by detecting and tracking the same fruit between video frame sequences. However, these studies are mostly based on video counting using ground cameras, and the application scope is mainly limited to single or small number of fruit tree samples, which cannot meet the needs of high-throughput and rapid detection in actual large-scale agricultural production environment. With the rapid development of unmanned aerial vehicle technology, its high efficiency, wide coverage and flexibility make it an ideal tool for agricultural monitoring, and fruit counting methods based on unmanned aerial vehicle video have gradually become a research hotspot. However, the existing methods have the following main problems: first, the lack of public data sets of shiitake in real scenes affects the training effect of the detection model; second, under the aerial view of the unmanned aerial vehicle, the existing detection methods have insufficient detection accuracy for small proportion, severely occluded and complex background fruits; finally, in the fruit tracking stage, the problems of tracking loss and frequent ID switching caused by camera movement and changes in lighting environment have not been effectively solved. SUMMARY
[0004] The present application provides a shiitake fruit counting method based on unmanned aerial vehicle video to overcome the above technical problems.
[0005] In order to achieve the above purpose, the technical scheme of the present application is:
[0006] A shiitake fruit counting method based on unmanned aerial vehicle video, specifically comprising:
[0007] S1: obtaining a sequence of unmanned aerial vehicle video frame images containing shiitake fruits;
[0008] and labeling the boundary box of the shiitake fruits in the sequence of unmanned aerial vehicle video frame images to obtain a video frame sample image, and randomly dividing the video frame sample image into a training set and a test set;
[0009] S2: Constructing a Xanthoceras sorbifolia Bunge fruit detection model;
[0010] And the Xanthoceras sorbifolia Bunge fruit detection model comprises a feature extraction network obtained by improving a YOLOv11 backbone network, a feature fusion network improved from a PAN-FPN, and a detection head network;
[0011] The feature extraction network extracts the fruit contour features in the video frame sample image to obtain a multi-scale feature map; the feature fusion network fuses the multi-scale feature map to obtain a fused feature map; and the detection head network detects the positioning frame of the Xanthoceras sorbifolia Bunge fruit according to the fused feature map;
[0012] S3: Training and verifying the constructed Xanthoceras sorbifolia Bunge fruit detection model according to the training set and the test set, obtaining an optimized Xanthoceras sorbifolia Bunge fruit detection model, and realizing positioning frame detection of the Xanthoceras sorbifolia Bunge fruit according to the optimized Xanthoceras sorbifolia Bunge fruit detection model to obtain a positioning frame detection score set;
[0013] And based on a preset score threshold, the positioning frame detection score set is divided into a high-score detection frame set and a low-score detection frame set;
[0014] S4: Based on a camera global motion compensation mechanism, a global affine transformation matrix for describing camera motion parameters is obtained according to the video frame sample image, and based on an improved adaptive Kalman filtering algorithm, the Xanthoceras sorbifolia Bunge fruit is tracked according to the global affine transformation matrix to obtain a tracking detection frame;
[0015] S5: Based on an IoU-ReID fusion mechanism, the high-score detection frame set and the tracking detection frame are associated and matched in similarity to obtain a primary matching detection frame and a primary non-matching detection frame;
[0016] Based on an IoU similarity mechanism, the primary non-matching detection frame and the low-score detection frame set are associated and matched in similarity to obtain a secondary matching detection frame and a secondary non-matching detection frame, and the secondary non-matching detection frame is discarded; and based on the primary matching detection frame and the secondary matching detection frame, the number of Xanthoceras sorbifolia Bunge fruits is confirmed.
[0017] Further, the feature extraction network obtained by improving the YOLOv11 backbone network in S2 comprises a first convolutional layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an SPPF module and a C2PSS module connected in sequence;
[0018] And each feature extraction module comprises a second convolutional layer and a first C3k2 module with different convolutional kernels;
[0019] The second convolutional layer is used for convolution operation on its input;
[0020] The first C3k2 module is configured to perform a feature extraction operation on the output of the second convolutional layer;
[0021] The first convolutional layer is configured to perform a convolution operation on the video frame sample image to obtain a convolution feature map;
[0022] The first feature extraction module is configured to perform a feature extraction operation on the convolution feature map to obtain a first feature map;
[0023] The second feature extraction module is configured to perform a feature extraction operation on the first feature map to obtain a second feature map;
[0024] The third feature extraction module is configured to perform a feature extraction operation on the second feature map to obtain a third feature map;
[0025] The fourth feature extraction module is configured to perform a feature extraction operation on the third feature map to obtain a fourth feature map;
[0026] The SPPF module is configured to perform a maximum pooling operation on the fourth feature map to obtain a fifth feature map;
[0027] The C2PSS module includes a third convolutional layer, a Split module, a plurality of stacked PSSBlock modules, a first concatenation layer, and a fourth convolutional layer;
[0028] The third convolutional layer is configured to perform a convolution operation on the fifth feature map to obtain a sixth feature map;
[0029] The Split module is configured to perform a feature segmentation operation on the sixth feature map to obtain a seventh feature map;
[0030] The PSSBlock module includes an SEAM module, a fifth convolutional layer, and a sixth convolutional layer connected in sequence;
[0031] The SEAM module is configured to extract local features of the occluded fruits in the seventh feature map to obtain an eighth feature map; the fifth convolutional layer is configured to perform a convolution operation on the ninth feature map to obtain a tenth feature map; the ninth feature map is a feature map obtained by element-wise addition of the seventh feature map and the eighth feature map; and the sixth convolutional layer is configured to perform a convolution operation on the tenth feature map to obtain an eleventh feature map;
[0032] The first concatenation layer is configured to perform a concatenation operation on the seventh feature map and a twelfth feature map to obtain a thirteenth feature map, and the twelfth feature map is a feature map obtained by element-wise addition of the ninth feature map and the eleventh feature map; the fourth convolutional layer is configured to perform a convolution operation on the thirteenth feature map to obtain a fourteenth feature map; and the first feature map, the second feature map, the third feature map, and the fourteenth feature map are the obtained multi-scale feature maps.
[0033] Further, the feature fusion network for improving the PAN-FPN in S2 comprises a first upsampling layer, a first fusion module, a second upsampling layer, a second fusion module, a third upsampling layer, a third fusion module, a seventh convolutional layer, a fourth fusion module, an eighth convolutional layer, a fifth fusion module, a ninth convolutional layer, and a sixth fusion module connected in sequence.
[0034] Each fusion module comprises a second splicing layer and a second C3k2 module with different convolution kernels.
[0035] The second splicing layer is configured to perform a splicing operation on the input thereof, and the second C3k2 module is configured to perform a feature extraction operation on the output of the second splicing layer.
[0036] The first upsampling layer is configured to perform an upsampling operation on the fourteenth feature map to obtain a fifteenth feature map.
[0037] The first fusion module is configured to obtain a sixteenth feature map according to the fifteenth feature map and the third feature map.
[0038] The second upsampling layer is configured to perform an upsampling operation on the fourteenth feature map to obtain a seventeenth feature map.
[0039] The second fusion module is configured to obtain an eighteenth feature map according to the seventeenth feature map and the second feature map.
[0040] The third upsampling layer is configured to perform an upsampling operation on the eighteenth feature map to obtain a nineteenth feature map.
[0041] The third fusion module is configured to obtain a twentieth feature map according to the nineteenth feature map and the first feature map.
[0042] The seventh convolutional layer is configured to perform a convolution operation on the twentieth feature map to obtain a twenty-first feature map.
[0043] The fourth fusion module is configured to obtain a twenty-second feature map according to the twenty-first feature map and the eighteenth feature map, the eighth convolutional layer is configured to perform a convolution operation on the twenty-second feature map to obtain a twenty-third feature map, the fifth fusion module is configured to obtain a twenty-fourth feature map according to the twenty-third feature map and the sixteenth feature map, the ninth convolutional layer is configured to perform a convolution operation on the twenty-fourth feature map to obtain a twenty-fifth feature map, and the sixth fusion module is configured to obtain a twenty-sixth feature map according to the twenty-fifth feature map and the fourteenth feature map. The twentieth feature map, the twenty-second feature map, the twenty-fourth feature map, and the twenty-sixth feature map are the obtained fusion feature maps.
[0044] Further, the detection head network in S2 comprises detection heads with different detection resolutions connected to the output ends of the third fusion module, the fourth fusion module, the fifth fusion module, and the sixth fusion module, respectively.
[0045] Further, the method for obtaining the optimized Xanthoceras sorbifolia Bunge fruit detection model in S3 specifically comprises the following steps:
[0046] S31: training the constructed Xanthoceras sorbifolia Bunge fruit detection model according to the training set to obtain a trained Xanthoceras sorbifolia Bunge fruit detection model:
[0047] S32: constructing a fusion loss function for fruit detection, and evaluating the trained Xanthoceras sorbifolia Bunge fruit detection model according to the test set to determine whether the output of the trained Xanthoceras sorbifolia Bunge fruit detection model converges;
[0048] If yes, the trained Xanthoceras sorbifolia Bunge fruit detection model is the optimized Xanthoceras sorbifolia Bunge fruit detection model;
[0049] Otherwise, the weight parameters of the trained Xanthoceras sorbifolia Bunge fruit detection model are adaptively adjusted based on the back propagation algorithm, and step S31 is repeatedly executed.
[0050] Further, the fusion loss function constructed in S32 is
[0051] L Inner-MPD =λ(1-IoU inner )+(1-λ)(1-MPDIoU)
[0052]
[0053] B inner =B gt ⊙resizefactor
[0054]
[0055] In the formula, L Inner-MPD represents the fusion loss function; λ represents a balance parameter for adjusting the weight of the loss term; IoU inner represents the intersection over union between the predicted bounding box B pred and the auxiliary bounding box B inner ; B inner represents the auxiliary bounding box inside the real bounding box B gt ; resizefactor represents a scaling factor; MPDIoU represents the minimum point distance between the bounding boxes; d i represents the minimum point distance between two bounding boxes; d c represents the distance between the centers of two bounding boxes; w and h respectively represent the width and height of the normalized reference scale; and IOU represents the intersection over union between the predicted bounding box and the real bounding box.
[0056] Further, S4 specifically comprises the following steps:
[0057] S41: Based on the camera global motion compensation mechanism: extract two preset key feature points in the continuous two frame video frame sample graphs, and set the corresponding relationship of the two preset key feature points;
[0058] Based on the RANSAC algorithm, the global affine transformation matrix for describing the camera motion parameters is estimated and obtained according to the corresponding relationship;
[0059] S42: Define the position state vector and the position observation vector of the detection tracking frame of the improved Kalman filtering algorithm, and configure a unique identification ID for each Xanthoceras sorbifolia fruit;
[0060] And the expression of the position state vector and the position observation vector is
[0061]
[0062] Z (k) =[z xc (t),z yc (t),z ω (t),z h (t)] T
[0063] In the formula: X (k) represents the position state vector; x c (t), y c (t) represents the center coordinates of the detection tracking frame; ω(t), h(t) represents the width and height of the detection tracking frame; represents the corresponding velocity change value of the state vector; Z (k) represents the position observation vector; z xc (t), z yc (t) represents the center coordinates of the tracking frame when measuring; z ω (t), z h (t) represents the width and height of the tracking frame when measuring; T represents transposition;
[0064] S43: Obtain the position state vector estimate value and the prediction error covariance according to the position state vector, and the expression is
[0065]
[0066] In the formula: represents the estimate value of the position state vector at k-1 time; represents obtain the estimate value of the position state vector at k time; Φ (k,k-1) represents the state transition matrix; P (k|k-1) , P (k-1) respectively represent the prediction error covariance and the error covariance; represents Φ (k,k-1)an estimated value of the position state vector; Q (k) denotes process noise; and (K,K-1) denotes a system noise matrix;
[0067] S44: based on the estimated value of the position state vector, obtaining the measurement noise by introducing an adjustment update parameter;
[0068] and the obtaining formula of the measurement noise is
[0069]
[0070] wherein d k denotes an adjustment update parameter; b denotes a forgetting factor; and v (k) denotes a system observation variable; H (k) denotes a system observation matrix; K (k-1) denotes a gain parameter; Y k denotes an intermediate variable; R (k-1) denotes the measurement noise at time k-1; and R (k) denotes the measurement noise at time k;
[0071] S45: updating the position state vector according to the measurement noise and the prediction error covariance to obtain an updated state vector;
[0072] and the obtaining formula of the updated state vector is
[0073]
[0074] P (k) = [I-K (k) H (k) ]P (k|k-1)
[0075]
[0076] wherein S (k) denotes an intermediate variable; P (k) denotes an updated prediction error covariance; denotes the updated state vector;
[0077] S46: performing motion compensation on the updated state vector by a global affine transformation matrix to realize trajectory tracking on the fruit of Xanthoceras sorbifolia Bunge to obtain a tracking detection box.
[0078] Further, the S5 specifically comprises the following steps:
[0079] S51: performing similarity correlation matching on the high-resolution detection box set and the tracking detection box based on an IoU-ReID fusion mechanism to obtain a primary matched detection box and a primary unmatched detection box;
[0080] And the expression based on the IoU-ReID fusion mechanism is
[0081] S(i,j) = lambda * S iou (i,j) + (1-lambda) * S app (i,j)
[0082]
[0083] In the formula, S(i,j) represents the matching similarity of the high-score detection box and the tracking detection box; lambda represents a weight parameter for dynamically adjusting the proportion of the IoU similarity and the ReID similarity according to the track state; S app (i,j) represents the ReID similarity; S iou (i,j) represents the IoU similarity. represents the predicted bounding box of the i-th track, that is, the tracking detection box; represents the j-th detection bounding box, that is, the high-score detection box; f i ,f j respectively represent the feature vector of the predicted bounding box of the i-th track and the feature vector of the j-th detection bounding box; tau represents the confidence score of the tracking detection box; tau high ,tau low represent a preset threshold; lambda high ,lambda low respectively represent the weight parameters in the high-confidence and low-confidence cases.
[0084] S52: Based on the IoU similarity mechanism, the similarity correlation matching of the initial unmatched detection box and the low-score detection box set is performed, the secondary matching detection box and the secondary unmatched detection box are obtained, and the secondary unmatched detection box is discarded.
[0085] And the expression of the IoU similarity mechanism is
[0086]
[0087] And based on the unique ID, the de-duplication operation of the initial matching detection box and the secondary matching detection box is performed, and then the actual number of the fruit of the Chinese toona fruit is confirmed.
[0088] Beneficial effects: the present application provides a Chinese toona fruit counting method based on unmanned aerial vehicle video,
[0089] The fruit contour features in the video frame sample picture are extracted through the feature extraction network, a multi-scale feature map is obtained, that is, the C2PSS module is constructed by fusing the SEAM attention mechanism into the C2PSA module of the YOLOv11n backbone network, so that the local features of the occluded fruits can be better captured, and the spatial correlation between the occluded area and the non-occluded area is established, and the detection ability of the model for the occluded fruits is improved; the fusion feature map is obtained through the feature fusion network; the bounding box detection of the Xanthoceras sorbifolia Bunge fruit is realized according to the fusion feature map through the detection head network, that is, a small target detection layer is added in the Head part of the YOLOv11n network, so that finer features can be extracted in the high-resolution stage, and the detection accuracy of the small Xanthoceras sorbifolia Bunge fruit is improved; in the model training stage, the positioning and recognition ability of the model for the small Xanthoceras sorbifolia Bunge fruit under the occlusion condition is improved through the constructed fusion loss function. In view of the problems of tracking loss and frequent ID switching caused by camera motion and light environment change, based on the camera global motion compensation mechanism combined with the improved adaptive Kalman filtering algorithm, the measurement noise parameter and the gain matrix can be dynamically adjusted, the adaptability of the system to the environment change is improved, the width and height of the tracking box are directly estimated by redefining the state vector, the positioning accuracy of the bounding box is improved, and the ID switching problem is effectively reduced; through the IoU-ReID fusion mechanism, the matching strategy can be flexibly adjusted in different scenes, the spatial information is more relied on when the target motion is relatively stable, and the appearance features are more relied on in the occlusion or fast motion scene, so that more stable tracking performance is realized. BRIEF DESCRIPTION OF DRAWINGS
[0090] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, below the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without paying creative labor.
[0091] Figure 1 The flowchart of the Xanthoceras sorbifolia Bunge fruit counting method based on the unmanned aerial vehicle video of the present application;
[0092] Figure 2 The schematic diagram of the Xanthoceras sorbifolia Bunge fruit detection model constructed in the present embodiment;
[0093] Figure 3 The network architecture schematic diagram of the SEAM module in the present embodiment;
[0094] Figure 4 The simplified schematic diagram of the network structure after adding the P2 detection layer in the present embodiment;
[0095] Figure 5 The overall framework diagram of the Xanthoceras sorbifolia Bunge fruit counting method in the present embodiment;
[0096] Figure 6 Fig. 3 is a fruit detection effect comparison chart in the embodiment;
[0097] Figure 7 Fig. 4 is a fruit tracking and counting effect comparison chart in the embodiment. DETAILED DESCRIPTION
[0098] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0099] The embodiment provides a shiitake mushroom fruit counting method based on a UAV video, as shown in the figure, specifically comprising: Figure 1
[0100] S1: acquiring a UAV video frame sequence image containing shiitake mushroom fruits;
[0101] and performing boundary box labeling on the shiitake mushroom fruits in the UAV video frame sequence image, acquiring a video frame sample image, and randomly dividing the video frame sample image into a training set and a test set;
[0102] S2: constructing a shiitake mushroom fruit detection model; as shown in the figure, and the shiitake mushroom fruit detection model comprises a feature extraction network obtained by improving a YOLOv11 backbone network, a feature fusion network improved by a PAN-FPN, and a detection head network; fruit contour features in the video frame sample image are extracted through the feature extraction network to obtain a multi-scale feature map; Figure 2
[0103] Specifically, the feature extraction network obtained by improving the YOLOv11 backbone network comprises a first convolutional layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an SPPF module and a C2PSS module connected in sequence;
[0104] And each feature extraction module comprises a second convolutional layer and a first C3k2 module with different convolution kernels;
[0105] The second convolutional layer is used for convolution operation on the input thereof;
[0106] The first C3k2 module is used for feature extraction operation on the output of the second convolutional layer;
[0107] The first convolutional layer is used for convolution operation on the video frame sample image to obtain a convolution feature map;
[0108] The first feature extraction module is used to extract features from the convolutional feature map to obtain the first feature map;
[0109] The second feature extraction module is used to extract features from the first feature map to obtain the second feature map;
[0110] The third feature extraction module is used to extract features from the second feature map to obtain the third feature map;
[0111] The fourth feature extraction module is used to extract features from the third feature map to obtain the fourth feature map;
[0112] The SPPF module is used to perform max pooling on the fourth feature map to obtain the fifth feature map;
[0113] The C2PSS module includes a third convolutional layer, a Split module, several stacked PSSBlock modules, a first splicing layer, and a fourth convolutional layer;
[0114] The third convolutional layer is used to perform a convolution operation on the fifth feature map to obtain the sixth feature map;
[0115] The Split module is used to perform feature segmentation on the sixth feature map to obtain the seventh feature map;
[0116] The PSSBlock module consists of a SEAM module, a fifth convolutional layer, and a sixth convolutional layer connected in sequence;
[0117] The SEAM module is used to extract local features of the occluded fruit in the seventh feature map to obtain the eighth feature map; the fifth convolutional layer is used to perform a convolution operation on the ninth feature map to obtain the tenth feature map; and the ninth feature map is the feature map obtained by adding the seventh feature map and the eighth feature map element by element; the sixth convolutional layer is used to perform a convolution operation on the tenth feature map to obtain the eleventh feature map.
[0118] The first concatenation layer is used to concatenate the seventh feature map and the twelfth feature map to obtain the thirteenth feature map. The twelfth feature map is obtained by adding the ninth feature map and the eleventh feature map element by element. The fourth convolutional layer is used to perform a convolution operation on the thirteenth feature map to obtain the fourteenth feature map. The first feature map, the second feature map, the third feature map, and the fourteenth feature map are the obtained multi-scale feature maps.
[0119] The feature extraction network (Backbone) in this embodiment is the core of its target detection performance, which efficiently extracts multi-scale semantic features of images through a multi-stage hierarchical structure. The network is constructed based on the improved C3K2 module, which optimizes the computational efficiency by adjusting the convolution kernel configuration while retaining the rich expression ability of deep features. In the shallow stage, the network preferentially extracts local detail features, such as edges and textures, and gradually captures global semantic information as the level deepens. To further enhance the feature capturing ability of the model for occluded targets, this embodiment introduces the SEAM attention mechanism into the C2PSA module of the original YOLOv11 to construct the C2PSS module, which improves the model's detection ability for occluded fruits by establishing the spatial correlation between occluded and non-occluded regions. In the detection of Xingguo fruits, occlusion is an important factor affecting the accuracy of the model. When Xingguo fruits are partially occluded by leaves, branches, or other Xingguo fruits, the traditional YOLOv11 cannot accurately identify the position or type of Xingguo fruits, resulting in detection errors. To address the occlusion problem in the Xingguo fruit detection task under the unmanned aerial vehicle aerial scene, this embodiment proposes the C2PSS module based on the YOLOv11 model. This module replaces the original attention mechanism in the PSSBlock module with the spatial enhancement attention mechanism SEAM, which effectively improves the detection performance of fruits in the occluded state and significantly enhances the adaptability of the YOLOv11 model to complex occlusion scenarios. The network structure of the C2PSS module is shown in Figure 2 , and the network structure of the SEAM module is shown in Figure 3 . The traditional C2PSA module is an important component of the YOLO11 model backbone network. By introducing the pyramid slice (PSA) attention mechanism, it significantly enhances the feature extraction and processing capabilities of the model and improves the feature processing and representation based on the squeeze and excitation (SE) attention mechanism, thereby improving the target detection performance. However, in specific applications such as Xingguo fruit detection, this module has limited effectiveness in detecting fruits in occluded scenarios. To overcome this deficiency, this embodiment constructs the C2PSS module, which replaces the original attention mechanism with SEAM, as shown in Figure 3 . SEAM is first implemented through the DcovN module, which combines deep separable convolution and residual connection to reduce the number of parameters while learning the importance of channels. In this embodiment, the DcovN module uses 1x1 convolution to fuse the outputs of different depth convolutions to compensate for the loss of channel relationships. In addition, a two-layer fully connected network is used to strengthen the relationships between channels, thereby reducing the information loss caused by occlusion.
[0120] The multi-scale feature maps are fused through the feature fusion network to obtain a fused feature map.
[0121] Specifically, the improved PAN-FPN feature fusion network comprises a first up-sampling layer, a first fusion module, a second up-sampling layer, a second fusion module, a third up-sampling layer, a third fusion module, a seventh convolutional layer, a fourth fusion module, an eighth convolutional layer, a fifth fusion module, a ninth convolutional layer and a sixth fusion module connected in sequence and different in network layer parameters.
[0122] Each fusion module comprises a second splicing layer and a second C3k2 module different in convolution kernel.
[0123] The second splicing layer is configured to perform a splicing operation on an input thereof, and the second C3k2 module is configured to perform a feature extraction operation on an output of the second splicing layer.
[0124] The first up-sampling layer is configured to perform an up-sampling operation on the fourteenth feature map to obtain a fifteenth feature map.
[0125] The first fusion module is configured to obtain a sixteenth feature map according to the fifteenth feature map and the third feature map.
[0126] The second up-sampling layer is configured to perform an up-sampling operation on the fourteenth feature map to obtain a seventeenth feature map.
[0127] The second fusion module is configured to obtain an eighteenth feature map according to the seventeenth feature map and the second feature map.
[0128] The third up-sampling layer is configured to perform an up-sampling operation on the eighteenth feature map to obtain a nineteenth feature map.
[0129] The third fusion module is configured to obtain a twentieth feature map according to the nineteenth feature map and the first feature map.
[0130] The seventh convolutional layer is configured to perform a convolution operation on the twentieth feature map to obtain a twenty-first feature map.
[0131] The fourth fusion module is configured to obtain a twenty-second feature map according to the twenty-first feature map and the eighteenth feature map, the eighth convolutional layer is configured to perform a convolution operation on the twenty-second feature map to obtain a twenty-third feature map, the fifth fusion module is configured to obtain a twenty-fourth feature map according to the twenty-third feature map and the sixteenth feature map, the ninth convolutional layer is configured to perform a convolution operation on the twenty-fourth feature map to obtain a twenty-fifth feature map, and the sixth fusion module is configured to obtain a twenty-sixth feature map according to the twenty-fifth feature map and the fourteenth feature map.
[0132] The feature fusion network in this embodiment realizes multi-scale feature fusion through an improved PAN-FPN architecture. The core of the network is composed of a C3K2 module for dynamic cross-scale connection. The network first performs bidirectional transmission of different levels of features output by the Backbone, such as shallow high-resolution features and deep semantic features: the top-down path transmits deep semantic information to the shallow layer to enhance the small target detection capability; the bottom-up path fuses detailed features to optimize the large target positioning accuracy. The C3K2 module further simplifies the calculation in this stage, reduces the parameter quantity through grouped convolution and adaptive channel compression, and at the same time retains the cross-layer feature interaction capability. In view of the problem of uneven distribution of multi-scale targets in complex scenes, the feature fusion network further introduces an adaptive feature selection mechanism to dynamically adjust the feature fusion weight according to the target size, and in the detection of small-sized velvet fruit, the small target is given a higher weight of shallow features with high resolution. The adaptive feature selection mechanism introduced is a known prior art, and will not be described in detail here.
[0133] The detection head network realizes the positioning frame detection of velvet fruit according to the fused feature map.
[0134] Specifically, the detection head network includes detection heads with different detection resolutions connected to the output ends of the third fusion module, the fourth fusion module, the fifth fusion module, and the sixth fusion module, respectively.
[0135] The embodiment considers that small target detection has been a key bottleneck restricting the performance of fruit detection algorithm based on unmanned aerial vehicle. Due to the particularity of aerial view and the inherent size characteristics of Xanthoceras sorbifolia Bunge, the fruit target in the obtained image usually presents obvious small target characteristics. Specifically, the proportion of Xanthoceras sorbifolia Bunge in the aerial image is usually less than 5% of the whole image, and the pixel size is usually between 10x10 and 30x30. This small size feature brings a series of unique challenges to the target detection algorithm. First, the feature loss problem is particularly prominent. In the multiple downsampling process of the traditional target detection network, the feature information of the small target is easily diluted or completely lost, so that the network cannot capture enough discriminative features. Second, the representation of small targets on high-level feature maps is extremely limited, occupying only a small number of pixels, which cannot provide sufficient context information. In addition, the light changes, shadows and occlusion effects in the natural environment further increase the difficulty of small target Xanthoceras sorbifolia Bunge recognition. As the latest single-stage detection algorithm in the target detection field, the traditional YOLOv11n adopts a feature pyramid network structure to realize multi-scale feature fusion. Its original architecture includes P3(80x80), P4(40x40) and P5(20x20) three detection layers, which are responsible for the detection of targets of different scales. Although this structure performs well in general target detection tasks, it has obvious limitations in processing small target detection such as unmanned aerial vehicle aerial Xanthoceras sorbifolia Bunge. As the feature map with the highest resolution in the traditional YOLOv11n, P3 detection layer should theoretically bear the main responsibility for small target detection. However, for tiny fruit targets with pixel size between 10x10 and 30x30, the 80x80 feature map resolution is still insufficient. On the P3 feature map, such small targets may only correspond to 1-3 pixel points, which seriously limits the detection performance. P4 and P5 detection layers are mainly aimed at medium and large targets, and their low-resolution feature maps have more limited representation ability for small targets, resulting in most small target fruits being missed in the inference process.
[0136] To solve the above problems, as shown in Figure 4 The embodiment introduces a P2 small target detection layer in the YOLOv11n architecture to enhance the model's detection ability for tiny fruit targets. The P2 layer generates a 160x160 high-resolution feature map, which is 4 times the resolution of the original P3 layer, providing more detailed feature representation for small targets. Figure 4The feature extraction network structure after introducing the P2 detection layer is shown. In the specific implementation of the P2 detection layer, first, the second stage feature map (160x160) output by the C3k2 module in the improved YOLOv11n backbone network, that is, the output of the first feature extraction module, is used as the basis for constructing the P2 layer. This feature map retains rich spatial detail information at a lower degree of down-sampling, which is beneficial to accurate positioning and identification of small targets. Second, to enhance the semantic representation ability of the P2 layer, a top-down feature fusion path is designed. First, the P3 feature map (80x80) is reduced in channel dimension by a 1x1 convolutional layer and then expanded to 160x160 by an up-sampling operation, followed by weighted fusion with the base feature map. Finally, the same detection head structure is applied to the P2 detection layer as other detection layers, and the anchor box size is optimized and adjusted for small target characteristics. The P2 layer gives priority to the coverage of small size targets in the design of anchor boxes to ensure that fruit targets in the range of 10x10 to 30x30 pixels are fully covered by prior boxes. In summary, the introduction of the P2 small target detection layer effectively compensates for the shortcomings of the traditional YOLOv11n in small target Xanthoceras sorbifolia fruit detection. By providing more detailed feature representation and optimized detection strategies, the model's fruit detection performance in the unmanned aerial vehicle aerial scene is improved. This technical improvement provides more reliable algorithm support for fruit monitoring and yield estimation in precision agriculture.
[0137] S3: training and verifying the Xanthoceras sorbifolia fruit detection model constructed according to the training set and the test set to obtain an optimized Xanthoceras sorbifolia fruit detection model;
[0138] In specific embodiments, the method for obtaining the optimized Xanthoceras sorbifolia fruit detection model specifically comprises:
[0139] S31: model training of the Xanthoceras sorbifolia fruit detection model constructed according to the training set to obtain a trained Xanthoceras sorbifolia fruit detection model:
[0140] S32: constructing a fusion loss function for fruit detection and model evaluation of the trained Xanthoceras sorbifolia fruit detection model according to the test set to determine whether the output of the trained Xanthoceras sorbifolia fruit detection model converges;
[0141] If yes, the trained Xanthoceras sorbifolia fruit detection model is the optimized Xanthoceras sorbifolia fruit detection model;
[0142] Otherwise, the weight parameters of the trained Xanthoceras sorbifolia fruit detection model are adaptively adjusted based on the back propagation algorithm, and step S31 is repeatedly executed;
[0143] And according to the optimized Xanthoceras sorbifolia fruit detection model, a bounding box detection of the Xanthoceras sorbifolia fruit is realized to obtain a set of bounding box detection scores;
[0144] And based on the preset score threshold, the positioning frame detection score set is divided into a high-score detection frame set and a low-score detection frame set;
[0145] Specifically, the constructed fusion loss function includes:
[0146] In this embodiment, the intersection over union (IOU) is used as a standard measure for evaluating the quality of boundary box prediction in the target detection field, that is, by defining the ratio of the intersection area of the predicted boundary box and the real boundary box to the union area, the mathematical expression is:
[0147]
[0148] In the formula: B p ,B gt respectively represent the predicted boundary box and the real boundary box. The traditional IOU is widely used in target detection and boundary box regression due to its simplicity and effectiveness, but it has obvious limitations. First, when the predicted box and the real box are completely non-overlapping, the IOU value is zero, which cannot distinguish different degrees of non-overlapping situation, leading to gradient disappearance problem, hindering the effective training of the model. Secondly, IOU has insufficient ability to distinguish boundary boxes of different overlapping degrees, especially in the high IOU value area, it is difficult to accurately distinguish the quality difference of boundary boxes. Thirdly, the sensitivity of IOU to small target detection is not enough, the same pixel deviation causes much greater IOU value change for small targets than for large targets. Finally, the traditional IOU only considers the overlapping area and ignores the geometric characteristics of the boundary box, such as the center point distance, the aspect ratio and other key information, which limits its application effect in complex scenes. In order to overcome the limitations of traditional IOU, the minimum point distance IOU (MPDIoU) introduces the concept of minimum point distance between boundary boxes, which significantly improves the ability to distinguish completely non-overlapping boundary boxes, and the calculation formula of MPDIoU is:
[0149]
[0150] In the formula: d i represents the minimum point distance between two boundary boxes; d cdistance between two bounding box center points; w, h represent the width and height of the normalized reference scale respectively; IOU represents the intersection over union of the predicted bounding box and the real bounding box; the minimum point distance is defined as the minimum Euclidean distance between all point pairs of two bounding boxes, which is usually simplified to calculate the minimum distance between eight feature points including four top points and four edge midpoints of the bounding box; MPDIoU solves the gradient vanishing problem of traditional IOU in the non-overlapping case by combining geometric distance information with traditional IOU; by introducing a distance penalty term, MPDIoU can still provide effective gradient information when the predicted box and the real box are not overlapped, guiding the model to adjust the predicted box position to be closer to the real box; in addition, MPDIoU considers the minimum point distance and the center point distance between the bounding boxes, providing more detailed geometric position guidance for bounding box regression, which is particularly suitable for target detection tasks in complex scenes.
[0151] In the embodiment, Inner-IOU is an innovative method for optimizing target detection performance by introducing the concept of auxiliary bounding box, the core idea of which is to dynamically generate an auxiliary bounding box inside the real bounding box, and use the IOU relationship between the auxiliary box and the predicted box to guide the model training, the expression of which is
[0152] B inner = B gt ⊙resizefactor
[0153]
[0154] In the formula, IoU inner represents the intersection over union between the predicted bounding box B pred and the auxiliary bounding box B inner ; B inner represents the real bounding box B gtan inner auxiliary bounding box; resizefactor represents a scaling factor; MPDIoU represents the minimum point distance between bounding boxes; and represents a scaling operator; wherein the inner auxiliary bounding box is generated by dynamically adjusting the scaling factor, and the size of the inner auxiliary bounding box is usually set to a certain proportion of the real bounding box, and the scaling factor usually ranges from 0 to 1; the two main advantages of this dynamic adjustment method are: first, by adjusting the size of the inner auxiliary bounding box, the model can learn more accurate bounding box positioning ability, especially for small target detection; second, the introduction of the inner auxiliary bounding box increases the diversity of training samples and improves the adaptability of the model to different scale targets. Inner-IOU reduces the uncertainty of bounding box prediction by forcing the predicted box to be closer to the target internal region, and is especially suitable for target detection tasks in complex backgrounds. The Inner-MPDIOU fusion loss function is an innovative solution that combines the advantages of MPDIoU and Inner-IOU loss functions, that is, the constructed fusion loss function is:
[0155] L Inner-MPD =λ(1-IoU inner )+(1-λ)(1-MPDIoU)
[0156] In the formula: λ represents a balance parameter for adjusting the relative importance of the two loss terms; (1-IoU inner ) represents the Inner-IOU loss, which is calculated by calculating the IOU of the predicted box and the inner auxiliary box and taking the inverse; (1-MPDIoU) represents the MPDIoU loss, which is calculated by calculating the MPDIoU of the predicted box and the real box and taking the inverse. Through this fusion method, the Inner-MPDIOU loss function combines the advantages of MPDIoU in geometric discrimination and the advantages of Inner-IOU in scale adaptability. Specifically, the MPDIoU component provides more detailed geometric position guidance by considering the minimum point distance and center point distance between bounding boxes, effectively solving the gradient disappearance problem in non-overlapping or partially overlapping cases; and the Inner-IOU component enhances the model's detection ability for different scale targets through the dynamic adjustment mechanism of the auxiliary bounding box, especially improving the positioning accuracy of small targets. By adjusting the λ parameter, the contributions of the two loss terms can be flexibly balanced during training, so as to obtain the best performance in different application scenarios. Therefore, in fruit detection and other application scenarios, the Inner-MPDIOU fusion loss function can obtain higher detection accuracy and better bounding box positioning effect, while showing faster convergence speed and more stable optimization process.
[0157] S4: based on a camera global motion compensation mechanism, a global affine transformation matrix for describing camera motion parameters is obtained according to a video frame sample image, and based on an improved adaptive Kalman filtering algorithm, a track detection frame of a Xanthoceras sorbifolia Bunge fruit is tracked according to the global affine transformation matrix;
[0158] As shown in the specific steps include: Figure 5
[0159] S41: based on a camera global motion compensation mechanism: two preset key feature points in two continuous video frame sample images are extracted, and a corresponding relationship of the two preset key feature points is set;
[0160] Based on a RANSAC algorithm, a global affine transformation matrix for describing camera motion parameters is estimated and obtained according to the corresponding relationship;
[0161] At present, the traditional SORT series algorithm performs well in a static camera scene, but in the case of camera motion, since the Kalman filter assumes that the target motion model is relatively stable, the prediction accuracy will be greatly reduced. To solve this problem, the embodiment introduces a camera motion compensation mechanism (CMC), the core idea of which is to estimate the camera motion and separate its influence from the target motion. Its implementation process mainly includes three steps of key point extraction and matching, motion model estimation and state vector compensation: first, key points are extracted in two continuous images and a corresponding relationship is established; second, a global affine transformation matrix A is estimated using a RANSAC algorithm, which describes the camera motion parameters; finally, the state vector predicted by the Kalman filter is compensated using the transformation matrix, which includes applying affine transformation to the position component, adjusting the size component according to the scaling factor, and making corresponding corrections to the velocity component, etc. Through this camera motion compensation mechanism, the embodiment significantly improves the target tracking accuracy in dynamic scenes, even if the camera is displaced or vibrates, it can still maintain stable tracking performance, effectively solving the problem of target loss and false tracking commonly faced by traditional multi-target tracking algorithms in camera motion scenes, thereby enhancing the tracking accuracy and environmental adaptability of the algorithm;
[0162] S42: In order to predict and estimate the motion trajectory of the Xanthoceras sorbifolia target in the two-dimensional plane in this embodiment, Kalman filtering is used as the core prediction module. The traditional Kalman filtering algorithm is designed based on a constant speed model. Although it can accurately predict the next trajectory of the target under ideal conditions, it has obvious limitations under the motion conditions of the unmanned aerial vehicle platform. When the unmanned aerial vehicle platform moves, its constant speed model is difficult to accurately describe the relative motion of the target, resulting in a deviation between the tracking frame and the actual target position, which affects the matching accuracy. The initial state matrix of the target is determined by estimating the aspect ratio of the Xanthoceras sorbifolia target to determine the tracking frame parameters. The accuracy of the bounding box directly affects the tracking and counting accuracy. Even a small deviation can cause two key problems: (1) The boundary boxes of targets with close positions may overlap and interfere with each other, leading to tracking failure and serious ID switching problems; (2) Insufficient accuracy of the boundary box increases the probability of misidentification, leading to non-Xanthoceras sorbifolia targets being incorrectly counted, affecting the counting accuracy.
[0163] In order to solve the above problems, the position state vector and the position observation vector of the detection tracking frame of the Kalman filtering algorithm are redefined and improved in this embodiment. By directly estimating the width and height of the tracking frame, the accuracy of the boundary box is improved, and a unique ID is configured for each Xanthoceras sorbifolia fruit.
[0164] The expressions of the position state vector and the position observation vector are
[0165]
[0166] Z (k) =[z xc (t),z yc (t),z ω (t),z h (t)] T
[0167] In the formula, X (k) represents the position state vector; x c (t), y c (t) represent the center coordinates of the detection tracking frame; ω(t), h(t) represent the width and height of the detection tracking frame; represents the speed change value corresponding to the state vector; Z (k) represents the position observation vector; z xc (t), z yc (t) represent the center coordinates of the tracking frame at the time of measurement; z ω (t), z h (t) represent the width and height of the tracking frame at the time of measurement; T represents the transpose.
[0168] S43: In practical applications, the system model needs to be updated and corrected continuously according to actual and predicted data. Therefore, the reasonable setting of the measurement noise R(k) and process noise Q(k) parameters is crucial. However, when the UAV camera experiences movement or changes in light and environment, the fixed measurement noise R(k) and process noise Q(k) parameters cannot adapt to real-time scenarios, leading to tracking loss and accuracy decline. To solve this problem, the embodiment introduces an adaptive Kalman filtering mechanism, dynamically adjusts the gain of related modules by fine-tuning the weight, and improves the system's environmental adaptability. The specific implementation process is as follows:
[0169] According to the position state vector, the position state vector estimate value and the prediction error covariance are obtained, and the expression is
[0170]
[0171] In the formula: represents the estimate value of the position state vector at time k-1; represents The estimate value of the position state vector at time k is obtained; Φ (k,k-1) represents the state transition matrix; P (k|k-1) , P (k-1) represent the prediction error covariance and error covariance, respectively; represents the estimate value of Φ (k,k-1) ; Q (k) represents the process noise; Γ (K,K-1) represents the system noise matrix;
[0172] S44: Based on the position state vector estimate value, the measurement noise is obtained by introducing an adjustment update parameter; and the measurement noise acquisition formula is
[0173]
[0174] In the formula: d k represents the adjustment update parameter; b represents the forgetting factor, with a value range from 0.95 to 0.99; v (k) represents the system observation variable; H (k) represents the system observation matrix; K (k-1) represents the gain parameter; Y k represents the intermediate variable; R (k-1) represents the measurement noise at time k-1; R (k) represents the measurement noise at time k; by introducing the adjustment update parameter d k into the Kalman filtering algorithm in this embodiment, the measurement noise R (k) can be dynamically adjusted based on d k , and R (k) tends to a stable value R (k-1) as the frame number k changes.
[0175] S45: updating the position state vector according to the measurement noise and the prediction error covariance to obtain an updated state vector;
[0176] and the formula for obtaining the updated state vector is
[0177]
[0178] P (k) = [I - K (k) H (k) ]P (k|k-1)
[0179]
[0180] In the formula, S (k) represents an intermediate variable; P (k) represents the updated prediction error covariance; represents the updated state vector;
[0181] In this embodiment, the improved Kalman filtering algorithm can dynamically adjust the noise parameters and the gain matrix according to the environmental changes, significantly improving the tracking accuracy of the Xanthoceras sorbifolia target. At the same time, the redefined state vector directly estimates the width and height of the tracking frame, further enhancing the bounding box accuracy and effectively reducing the ID switching problem, laying a technical foundation for Xanthoceras sorbifolia fruit counting;
[0182] S46: performing motion compensation on the updated state vector through a global affine transformation matrix to realize trajectory tracking of the Xanthoceras sorbifolia fruit and obtain a tracking detection frame.
[0183] S5: performing similarity correlation matching on the high-score detection frame set and the tracking detection frame based on the IoU-ReID fusion mechanism to obtain primary matched detection frames and primary unmatched detection frames;
[0184] performing similarity correlation matching on the primary unmatched detection frames and the low-score detection frame set based on the IoU similarity mechanism to obtain secondary matched detection frames and secondary unmatched detection frames, discarding the secondary unmatched detection frames, and confirming the number of Xanthoceras sorbifolia fruits based on the primary matched detection frames and the secondary matched detection frames.
[0185] Specifically, the following steps are included:
[0186] S51: performing similarity correlation matching on the high-score detection frame set and the tracking detection frame based on the IoU-ReID fusion mechanism to obtain primary matched detection frames and primary unmatched detection frames;
[0187] and the expression based on the IoU-ReID fusion mechanism is
[0188] S(i,j) = λ · S iou (i,j) + (1 - λ) · S app (i,j)
[0189]
[0190] wherein S(i,j) represents the matching similarity of the high-score bounding box and the tracking bounding box; λ represents a weight parameter for dynamically adjusting the proportion of the IoU similarity and the ReID similarity according to the track state; S app (i,j) represents the ReID similarity; S iou (i,j) represents the IoU similarity. represents the predicted bounding box of the i-th track, i.e., the tracking bounding box; represents the j-th detection bounding box, i.e., the high-score bounding box; f i f j respectively represent the feature vector of the predicted bounding box of the i-th track and the feature vector of the j-th detection bounding box; τ represents the confidence score of the tracking bounding box; τ high , τ low represents a preset threshold; λ high , λ low respectively represent the weight parameters in the high-confidence and low-confidence cases.
[0191] S52: based on the IoU similarity mechanism, the similarity correlation matching is performed on the initial unmatched detection bounding box and the low-score detection bounding box set, the secondary matched detection bounding box and the secondary unmatched detection bounding box are obtained, and the secondary unmatched detection bounding box is discarded;
[0192] and the expression of the IoU similarity mechanism is
[0193]
[0194] and based on the unique identifier ID, the de-duplication operation is performed on the initial matched detection bounding box and the secondary matched detection bounding box, and then the actual number of the fruit of the Xanthoceras sorbifolia Bunge is confirmed.
[0195] In this embodiment, a multi-target tracking (MOT) algorithm is adopted, which is a basic task in computer vision, and the core challenge is how to maintain the consistency of the target identity in a complex scene. The method described in this embodiment innovatively fuses motion information and appearance features, and proposes an efficient IoU-ReID fusion mechanism, which significantly improves the target tracking performance. Among them, the extraction of the ReID feature adopts the network architecture of the Transformer, which can capture the global and local features of the target, and has strong robustness to occlusion and pose changes. The feature vector extracted by it can pass through the cosine similarity S appThe method described in the embodiment calculates the similarity of the target appearance by using (i, j), and proposes an adaptive fusion strategy, i.e., dynamically adjusting the weight of IoU and ReID according to the confidence of tracking, and λ is a weight parameter dynamically adjusted according to the track state. The adaptive fusion strategy can make the BoT-SORT tracking algorithm, i.e., the method described in steps S4 to S5 in the embodiment, flexibly adjust the matching strategy in different scenes, rely more on spatial information when the target motion is relatively stable, and more utilize appearance features in the occlusion or fast motion scene, so as to realize more robust tracking performance.
[0196] The fruit detection model constructed in the embodiment is combined with the BoT-SORT tracking algorithm to realize accurate counting of the Xanthoceras sorbifolia Bunge fruits in the orchard. In the fruit detection stage, the Xanthoceras sorbifolia Bunge fruit detection model outputs the bounding box coordinates, confidence and category information of each detection target, providing high-precision detection results for subsequent tracking. In the tracking stage, the BoT-SORT tracking algorithm stabilizes the prediction of the fruit motion track through the adaptive Kalman filter and the camera motion compensation mechanism, and assigns a unique identifier (ID) to each fruit, ensuring the consistency of the cross-frame target. The counting module processes the detection and tracking results: first, the fruit track is de-duplicated according to the unique ID to avoid repeated statistics; then, the number of valid tracks is accumulated, and finally the total number of Xanthoceras sorbifolia Bunge fruits in the orchard is output. As shown in Figure 6 to Figure 7 Figure 6 (a) represents a frame image to be detected; Figure 6 (b) represents the detection result graph of the traditional YOLOv11n; Figure 6 (c) represents the detection result graph of the method described in the embodiment; Figure 7 (a) represents the tracking effect graph of the traditional tracking algorithm; Figure 7 (b) represents the tracking effect graph of the method described in the embodiment; it can be known from the result graph that the method described in the embodiment effectively solves the error problem caused by target occlusion, motion blur or repeated detection in the traditional counting, and significantly improves the accuracy and reliability of automatic counting.
[0197] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for counting Xanthoceras sorbifolia fruit based on UAV video, characterized in that, Specifically comprising: S1: obtaining a sequence of unmanned aerial vehicle video frame images containing Xanthoceras sorbifolia fruits; And the Xanthoceras sorbifolia fruits in the sequence of unmanned aerial vehicle video frame images are labeled with a bounding box to obtain a video frame sample image, and the video frame sample image is randomly divided into a training set and a test set; S2: constructing a Xanthoceras sorbifolia fruit detection model; And the Xanthoceras sorbifolia fruit detection model comprises a feature extraction network obtained by improving a YOLOv11 backbone network, a feature fusion network obtained by improving a PAN-FPN, and a detection head network; The fruit contour features in the video frame sample image are extracted through the feature extraction network to obtain a multi-scale feature map; the multi-scale feature map is fused through the feature fusion network to obtain a fused feature map; and the bounding box detection of the Xanthoceras sorbifolia fruit is realized through the detection head network according to the fused feature map; S3: training and verifying the constructed Xanthoceras sorbifolia fruit detection model according to the training set and the test set, obtaining an optimized Xanthoceras sorbifolia fruit detection model, and realizing the bounding box detection of the Xanthoceras sorbifolia fruit according to the optimized Xanthoceras sorbifolia fruit detection model to obtain a bounding box detection score set; And based on a preset score threshold, the bounding box detection score set is divided into a high-score detection box set and a low-score detection box set; S4: based on a camera global motion compensation mechanism, a global affine transformation matrix for describing camera motion parameters is obtained according to the video frame sample image, and based on an improved adaptive Kalman filtering algorithm, trajectory tracking of the Xanthoceras sorbifolia fruit is performed according to the global affine transformation matrix to obtain a tracking detection box; S5: based on an IoU-ReID fusion mechanism, the high-score detection box set and the tracking detection box are associated and matched in similarity to obtain a primary matching detection box and a primary non-matching detection; Based on an IoU similarity mechanism, the primary non-matching detection box and the low-score detection box set are associated and matched in similarity to obtain a secondary matching detection box and a secondary non-matching detection box, and the secondary non-matching detection box is discarded; And based on the primary matching detection box and the secondary matching detection box, the number of Xanthoceras sorbifolia fruits is confirmed.
2. The method according to claim 1, wherein, The feature extraction network obtained by improving the YOLOv11 backbone network in S2 comprises a first convolutional layer, a first feature extraction module, a second feature extraction module, a third feature extraction module, a fourth feature extraction module, an SPPF module and a C2PSS module connected in sequence; And each feature extraction module comprises a second convolutional layer and a first C3k2 module with different convolution kernels; The second convolutional layer is used for convolution operation on its input; The first C3k2 module is used for feature extraction operation on the output of the second convolutional layer; The first convolutional layer is used for convolution operation on the video frame sample image to obtain a convolution feature map; The first feature extraction module is used for feature extraction on the convolution feature map to obtain a first feature map; The second feature extraction module is used for feature extraction on the first feature map to obtain a second feature map; The third feature extraction module is used for feature extraction on the second feature map to obtain a third feature map; The fourth feature extraction module is used for feature extraction on the third feature map to obtain a fourth feature map; The SPPF module is used for maximum pooling operation on the fourth feature map to obtain a fifth feature map; The C2PSS module comprises a third convolutional layer, a Split module, a plurality of stacked PSSBlock modules, a first splicing layer, and a fourth convolutional layer; The third convolutional layer is configured to perform convolutional operation on the fifth feature map to obtain a sixth feature map; The Split module is configured to perform feature segmentation operation on the sixth feature map to obtain a seventh feature map; The PSSBlock module comprises an SEAM module, a fifth convolutional layer, and a sixth convolutional layer connected in sequence; The SEAM module is configured to extract local features of the occluded fruits in the seventh feature map to obtain an eighth feature map; the fifth convolutional layer is configured to perform convolutional operation on a ninth feature map to obtain a tenth feature map; the ninth feature map is a feature map obtained by element-wise addition of the seventh feature map and the eighth feature map; and the sixth convolutional layer is configured to perform convolutional operation on the tenth feature map to obtain an eleventh feature map; The first splicing layer is configured to perform splicing operation on the seventh feature map and a twelfth feature map to obtain a thirteenth feature map, the twelfth feature map being a feature map obtained by element-wise addition of the ninth feature map and the eleventh feature map; and the fourth convolutional layer is configured to perform convolutional operation on the thirteenth feature map to obtain a fourteenth feature map; the first feature map, the second feature map, the third feature map, and the fourteenth feature map are the obtained multi-scale feature maps.
3. The method of claim 2, wherein, The improved PAN-FPN feature fusion network in S2 comprises a first upsampling layer, a first fusion module, a second upsampling layer, a second fusion module, a third upsampling layer, a third fusion module, a seventh convolutional layer, a fourth fusion module, an eighth convolutional layer, a fifth fusion module, a ninth convolutional layer, and a sixth fusion module connected in sequence; Each fusion module comprises a second splicing layer and a second C3k2 module with different convolution kernels; The second splicing layer is configured to perform splicing operation on its input, and the second C3k2 module is configured to perform feature extraction operation on the output of the second splicing layer; The first upsampling layer is configured to perform upsampling operation on the fourteenth feature map to obtain a fifteenth feature map; The first fusion module is configured to obtain a sixteenth feature map according to the fifteenth feature map and the third feature map; The second upsampling layer is configured to perform upsampling operation on the fourteenth feature map to obtain a seventeenth feature map; The second fusion module is configured to obtain an eighteenth feature map according to the seventeenth feature map and the second feature map; The third upsampling layer is configured to perform upsampling operation on the eighteenth feature map to obtain a nineteenth feature map; The third fusion module is configured to obtain a twentieth feature map according to the nineteenth feature map and the first feature map; The seventh convolutional layer is configured to perform convolutional operation on the twentieth feature map to obtain a twenty-first feature map; The fourth fusion module is configured to obtain a twenty-second feature map according to the twenty-first feature map and the eighteenth feature map; and the eighth convolutional layer is configured to perform convolutional operation on the twenty-second feature map to obtain a twenty-third feature map; The fifth fusion module is configured to obtain a twenty-fourth feature map according to the twenty-third feature map and the sixteenth feature map; and the ninth convolutional layer is configured to perform convolutional operation on the twenty-fourth feature map to obtain a twenty-fifth feature map; The sixth fusion module is configured to obtain a twenty-sixth feature map according to the twenty-fifth feature map and the fourteenth feature map; and the twentieth feature map, the twenty-second feature map, the twenty-fourth feature map and the twenty-sixth feature map are the obtained fusion feature maps.
4. The method according to claim 3, wherein, The detection head network in S2 includes detection heads respectively connected to output ends of the third fusion module, the fourth fusion module, the fifth fusion module and the sixth fusion module and having different detection resolutions.
5. The method of counting Xanthoceras sorbifolia fruit based on UAV video according to claim 4, characterized in that, In S3, the method for obtaining the optimized Xanthoceras sorbifolia Bunge fruit detection model includes the following steps: In S31, the Xanthoceras sorbifolia Bunge fruit detection model is trained according to the training set to obtain a trained Xanthoceras sorbifolia Bunge fruit detection model. In S32, a fusion loss function for fruit detection is constructed, and the trained Xanthoceras sorbifolia Bunge fruit detection model is evaluated according to the test set to determine whether the output of the trained Xanthoceras sorbifolia Bunge fruit detection model converges. If yes, the trained Xanthoceras sorbifolia Bunge fruit detection model is the optimized Xanthoceras sorbifolia Bunge fruit detection model. Otherwise, the weight parameters of the trained Xanthoceras sorbifolia Bunge fruit detection model are adaptively adjusted based on a back propagation algorithm, and step S31 is repeatedly executed.
6. The method of counting the fruit of Xanthoceras sorbifolia Bunge based on the video of the UAV according to claim 5, characterized in that, The fusion loss function constructed in S32 is L Inner-MPD = λ(1 - IoU inner )+(1 - λ)(1 - MPDIoU) B inner = B gt ⊙ resizefactor In the formula: L Inner-MPD The fusion loss function is represented by λ, which is the balancing parameter used to adjust the weights of the loss terms; IoU inner Represents the predicted bounding box B pred With auxiliary bounding box B inner The crossover ratio between them; B inner Represents the true bounding box B gt The inner auxiliary bounding box; resizefactor represents the scaling factor; MPDIoU represents the minimum point distance between bounding boxes; d i d represents the minimum point distance between two bounding boxes; c This represents the distance between the center points of the two bounding boxes; w and h represent the width and height of the normalized reference scale, respectively. IOU represents the intersection over union of the predicted bounding box and the real bounding box.
7. The method of counting Xanthoceras sorbifolia fruit based on UAV video according to claim 6, characterized in that, S4 specifically includes the following steps: In S41, based on a camera global motion compensation mechanism, two preset key feature points in two continuous video frame samples are extracted, and the corresponding relationship of the two preset key feature points is set. Based on a RANSAC algorithm, the global affine transformation matrix for describing the camera motion parameters is estimated and obtained according to the corresponding relationship. In S42, the position state vector and the position observation vector of the detection tracking box of the improved Kalman filtering algorithm are defined, and a unique identification ID is configured for each Xanthoceras sorbifolia Bunge fruit. The expressions of the position state vector and the position observation vector are Z (k) = [z xc (t),z yc (t),z ω (t),z h (t)] T where X (k) represents a position state vector; x c (t), y c (t) represents the center coordinates of the detected tracking frame; ω(t), h(t) represents the width and height of the detected tracking frame; represents a velocity change value corresponding to the state vector; Z (k) represents a position observation vector; z xc (t), z yc (t) represents the center coordinates of the tracking frame at the time of measurement; z ω (t), z h (t) represents the width and height of the tracking frame at the time of measurement; T represents transposition; In S43, the position state vector estimation value and the prediction error covariance are obtained according to the position state vector, and the expressions are where: denotes an estimate of the position state vector at time k - 1; denotes an estimate of the position state vector at time k; Φ (k,k-1) denotes a state transition matrix; P (k|k-1) denotes a measurement matrix; R (k-1) denotes a prediction error covariance and an error covariance, respectively; denotes an estimate of Φ (k,k-1) ; Q (k) denotes a process noise; Γ (K,K-1) denotes a system noise matrix; In S44, the measurement noise is obtained by introducing an adjustment update parameter based on the position state vector estimation value. The measurement noise is obtained according to the following formula where: d k denotes the adjustment update parameter; b denotes the forgetting factor; v (k) denotes the system observation variable; H (k) denotes the system observation matrix; K (k-1) denotes the gain parameter; Y k denotes the intermediate variable; R (k-1) denotes the measurement noise at time k - 1; R (k) denotes the measurement noise at time k; In S45, the position state vector is updated according to the measurement noise and the prediction error covariance to obtain an updated state vector. The updated state vector is obtained according to the following formula P (k) = [I - K (k) H (k) ]P (k|k-1) where: S (k) denotes an intermediate variable; P (k) denotes the updated prediction error covariance; denotes the updated state vector; In S46, the updated state vector is motion compensated by the global affine transformation matrix to realize trajectory tracking of the Xanthoceras sorbifolia Bunge fruit to obtain a tracking detection box.
8. The method according to claim 7, wherein, S5 specifically includes the following steps: In S51, the IoU-ReID fusion mechanism is used to perform similarity correlation matching on the high-score detection box set and the tracking detection box to obtain primary matched detection boxes and primary unmatched detection boxes. The expression of the IoU-ReID fusion mechanism is S(i,j) = λ · S iou (i,j) + (1 - λ) · S app (i,j) In the formula, S(i, j) represents the matching similarity of the high-score detection frame and the tracking detection frame; λ represents a weight parameter for dynamically adjusting the proportion of the IoU similarity and the ReID similarity according to the track state; S app (i, j) represents the ReID similarity; S iou (i, j) represents the IoU similarity; represents the predicted bounding box of the i th track, that is, the tracking detection frame; represents the j th detection bounding box, that is, the high-score detection frame; f i ,f j respectively represent the feature vector of the predicted bounding box of the i th track and the feature vector of the j th detection bounding box; τ represents the confidence score of the tracking detection frame; τ high ,τ low represents a preset threshold; λ high ,λ low respectively represent the weight parameters in the high-confidence and low-confidence cases; In S52, the IoU similarity mechanism is used to perform similarity correlation matching on the primary unmatched detection boxes and the low-score detection box set to obtain secondary matched detection boxes and secondary unmatched detection boxes, and the secondary unmatched detection boxes are discarded. The expression of the IoU similarity mechanism is The primary matched detection boxes and the secondary matched detection boxes are de-duplicated based on the unique identification ID, and the actual number of the Xanthoceras sorbifolia Bunge fruits is determined.
Citation Information
Cited By
Systems and Methods for Video-Based Orchard Item Counting and Fruit Weight Estimation Using Motion-Guided Tracking and Elliptical Volume Modelling
AU2026203124B1