Method and apparatus for fruit recognition and counting
Patent Information
- Application Number
- CN202610680132.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-18
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]本申请提供一种果实识别与计数方法及装置,用以解决现有技术中过度依赖个体追踪易受遮挡导致轨迹断裂与重复计数的技术缺陷,实现基于果树级结构匹配与运动一致性的高精度去重与精准计数
[0017] The fruit identification and counting method and apparatus provided in this application aggregate individual fruits into fruit tree clusters through horizontal constraint clustering, extract spatial topological structure features for tree-level matching, and combine motion consistency constraints of dominant motion screening to achieve accurate deduplication at the fruit level. This can effectively avoid repeated counting across frames and significantly improve the accuracy, robustness and real-time performance of fruit counting in complex orchard environments.
Smart Images

Figure CN122780940A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method and apparatus for fruit recognition and counting. Background Technology
[0002] Currently, to automate orchard yield estimation, existing yield estimation robots typically employ a combination of deep learning object detection and multi-object tracking (MOT). The specific technical process generally involves: First, acquiring continuous orchard video sequences via cameras, and using deep learning object detection networks (such as the YOLO series) to extract color and semantic features from the images, locating the bounding box coordinates of individual fruits in each frame; then, to prevent the same fruit from being counted repeatedly in different video frames, the system usually introduces algorithms such as Kalman filtering to predict the fruit's trajectory, and combines this with matching mechanisms such as Mahalanobis distance or Intersection over Union (IoU), using algorithms like the Hungarian algorithm to perform data association and trajectory tracking on individual fruits in consecutive frames. Finally, the yield estimation result is obtained by counting the number of successfully tracked trajectories in the video stream.
[0003] However, the existing fruit identification and counting methods have significant drawbacks in complex orchard applications. First, limited by natural environmental factors such as foliage shading, drastic changes in lighting, and dense fruit overlap, existing single-frame detection based on bounding boxes is prone to missed detections or misidentifying multiple adjacent fruits as a single target. Second, when addressing cross-frame repetition counting, existing technologies are highly susceptible to fragmented fruit tracking trajectories due to sudden changes in perspective caused by robot movement and dense occlusion, leading to severe identity switching and deduplication failures. Furthermore, when faced with fruit targets in orchards that exhibit high visual similarity and frequently disappear and reappear, the reliability of cross-frame data association is extremely low, resulting in a persistently high matching error rate. Consequently, the accuracy of the final yield estimation and the robustness of the algorithm fail to meet the demands of actual production. Summary of the Invention
[0004] This application provides a fruit recognition and counting method and apparatus to solve the technical defects of the prior art that rely too much on individual tracking and are prone to occlusion, resulting in trajectory breakage and repeated counting, and to achieve high-precision deduplication and accurate counting based on fruit tree-level structural matching and motion consistency.
[0005] This application provides a method for fruit identification and counting, including the following steps.
[0006] Obtain the fruit position information of each frame in the video sequence; Clustering is performed based on the spatial distribution characteristics of fruit location information to obtain fruit tree clusters for each frame of the image; Extract the spatial structure features of each fruit tree cluster, and match the fruit tree clusters in the adjacent previous frame image and the next frame image based on the spatial structure features; Calculate the spatial translation vector based on the successfully matched fruit tree clusters, and translate the fruit position information of the previous frame image according to the spatial translation vector. The fruit position information after translation is locally matched with the fruit position information of the next frame image. The matched fruit position information is removed and the unmatched fruit position information is counted to obtain the fruit count result.
[0007] According to the fruit identification and counting method provided in this application, the fruit position information of each frame in a video sequence image is obtained, including: The video sequence images are input into a pre-trained fruit detection model, which outputs the fruit location information of each frame. The fruit detection model includes a feature extraction layer, a feature fusion layer, and an output prediction layer; The feature extraction layer includes the first 13 layers of the VGG16 backbone network, and a coordinate attention mechanism module is cascaded after each convolutional block of the VGG16 backbone network to obtain feature maps that retain spatial location information. The feature fusion layer employs a bidirectional feature pyramid network to perform bidirectional information flow weighted fusion of feature maps at different scales. Furthermore, the bidirectional feature pyramid network internally uses depthwise separable convolutions to process the weighted fused feature maps. The output prediction layer includes a classification head and a regression head. The regression head is configured to output a set of fruit location points using the Huber loss function.
[0008] According to the fruit identification and counting method provided in this application, clustering is performed based on the spatial distribution characteristics of fruit location information to obtain fruit tree clusters for each frame of the image, including: Calculate the distance between any two fruit locations along the horizontal axis of the image coordinate system; Fruit locations whose horizontal distance is less than the horizontal distance threshold are grouped into the same candidate cluster; Candidate clusters containing more than the minimum cluster size threshold are identified as fruit tree clusters.
[0009] According to the fruit identification and counting method provided in this application, the spatial structure features include topological features, which are extracted through the following steps: Divide the set of fruit locations within the fruit tree cluster into an upper subset and a lower subset according to the height direction; Perform Delaunay triangulation on the fruit positions in the upper and lower subsets respectively to obtain the corresponding upper and lower edge sets; Calculate the Euclidean lengths of each edge in the upper and lower edge sets, and generate the upper edge length distribution histogram and the lower edge length distribution histogram respectively. The upper and lower side length distribution histograms are used as topological features.
[0010] According to the fruit identification and counting method provided in this application, a spatial translation vector is calculated based on successfully matched fruit tree clusters, including: Calculate the centroid difference between the matching fruit tree clusters in the previous frame and the next frame, and use it as the initial translation vector; A random sampling consensus algorithm is used to perform robust regression on all initial translation vectors to remove outliers, thus obtaining spatial translation vectors.
[0011] According to the fruit recognition and counting method provided in this application, the method involves locally matching the translated fruit position information with the fruit position information of the subsequent frame image, including: For each translated fruit position point, find the nearest neighbor fruit position point in the next frame image and establish candidate matching pairs; Calculate the displacement vectors for all candidate matching pairs; The dominant motion distance is determined based on the distance distribution of the displacement vector, and the dominant motion direction is determined based on the angular distribution of the displacement vector. Select the matching pairs whose displacement vectors satisfy the first and second conditions from the candidate matching pairs, and use them as the target matching pairs; The first condition is that the deviation between the displacement vector and the dominant motion distance is less than a first threshold, and the second condition is that the deviation between the displacement vector and the dominant motion direction is less than a second threshold.
[0012] According to the fruit identification and counting method provided in this application, the fruit counting results are obtained, including: In the next frame image, mark the fruit position information corresponding to the target matching pair as duplicate detection and remove it; The location information of unmatched fruits in the next frame is included in the count as new detection targets; The newly detected targets in each frame are accumulated to obtain the fruit count result.
[0013] This application also provides a fruit identification and counting device, including the following modules: The acquisition module is used to acquire the fruit position information of each frame in the video sequence image; The clustering module is used to cluster fruits based on the spatial distribution characteristics of their location information to obtain fruit tree clusters for each frame of the image. The feature extraction and matching module is used to extract the spatial structure features of each fruit tree cluster and match the fruit tree clusters in the adjacent previous frame image and the next frame image based on the spatial structure features. The translation module is used to calculate the spatial translation vector based on the successfully matched fruit tree clusters, and to translate the fruit position information of the previous frame image according to the spatial translation vector. The counting module is used to locally match the fruit position information after translation with the fruit position information of the next frame image, remove the matching fruit position information and count the unmatched fruit position information to obtain the fruit counting result.
[0014] This application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the fruit recognition and counting method as described above.
[0015] This application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the fruit identification and counting method as described above.
[0016] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the fruit recognition and counting method as described above.
[0017] The fruit identification and counting method and apparatus provided in this application aggregate individual fruits into fruit tree clusters through horizontal constraint clustering, extract spatial topological structure features for tree-level matching, and combine motion consistency constraints of dominant motion screening to achieve accurate deduplication at the fruit level. This can effectively avoid repeated counting across frames and significantly improve the accuracy, robustness and real-time performance of fruit counting in complex orchard environments. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the fruit identification and counting method provided in this application.
[0020] Figure 2 The AD-P2PNet network structure diagram provided for this application.
[0021] Figure 3 The coordinate attention module structure diagram provided for this application.
[0022] Figure 4 The bidirectional feature pyramid network structure diagram provided for this application.
[0023] Figure 5 The structure diagram of the depthwise separable convolutional module provided in this application.
[0024] Figure 6 A schematic diagram of topological feature extraction based on Delaunay triangulation provided in this application.
[0025] Figure 7 The overall flowchart for cross-frame deduplication provided in this application.
[0026] Figure 8 This is a schematic diagram of the fruit identification and counting device provided in this application.
[0027] Figure 9 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0029] The fruit identification and counting method provided in this application can be implemented by an orchard robot, an intelligent mobile acquisition platform, a server, or any electronic device with image processing and computing capabilities. In this embodiment, an orchard robot is used as the executor for illustrative purposes. This robot is equipped with an image acquisition unit and an embedded processing unit for real-time acquisition of orchard video sequences and execution of subsequent fruit identification and counting steps. It is understood that the method of this application is also applicable to other devices with image acquisition and data processing capabilities, such as drones, handheld terminals, or cloud servers; this application does not specifically limit these devices.
[0030] The following is combined with Figures 1 to 9 This application describes the fruit identification and counting method and apparatus.
[0031] Figure 1 A flowchart illustrating the fruit identification and counting method provided in this application is shown below. Figure 1 As shown, the method includes the following: Step 101: Obtain the fruit position information of each frame in the video sequence.
[0032] Among them, the video image sequence refers to the set of image frames containing fruit tree targets that are continuously collected by the orchard robot during its movement.
[0033] In the embodiments of this application, there can be multiple ways to acquire video sequence images.
[0034] For example, high-definition visible light cameras mounted on orchard robots can continuously capture images as the robot moves along rows of fruit trees.
[0035] For example, a drone equipped with a multispectral camera can be used to conduct low-altitude flight data collection to obtain video sequence images containing information about the fruit tree canopy.
[0036] Alternatively, the video sequence images can also come from pre-stored image files, such as orchard videos already captured from an SD card, external hard drive, or cloud server. In scenarios requiring nighttime operation, cameras with active lighting can be configured, or thermal imaging cameras can be used for acquisition.
[0037] In addition, to obtain more comprehensive three-dimensional structural information of fruit trees, depth cameras or binocular stereo vision systems can be used to acquire depth information of the fruit while acquiring color images, providing a richer data foundation for subsequent spatial clustering and structural feature extraction.
[0038] It is understood that the above acquisition methods can be used individually or in combination depending on the actual scenario. This application does not strictly limit the specific source and acquisition method of the video sequence images, as long as continuous image frames containing fruit trees can be obtained.
[0039] After acquiring the video sequence images, obtain the fruit position information of each frame in the video sequence images.
[0040] Specifically, each frame of the video sequence is input into a pre-trained fruit detection model, which outputs the pixel coordinates of all fruits in the current frame to form the fruit position information of the current frame.
[0041] The fruit detection model includes a feature extraction layer, a feature fusion layer, and an output prediction layer.
[0042] For example, a fruit detection model could be an AD-P2PNet (Augmented Detection P2PNet) network. It should be noted that this AD-P2PNet network introduces a coordinate attention mechanism (CA), a bidirectional feature pyramid network (BiFPN), and the Huber loss function on top of P2PNet.
[0043] For example, such as Figure 2As shown, the AD-P2PNet network includes a feature extraction layer, a feature fusion layer, and an output prediction layer. The feature extraction layer uses the first 13 convolutional layers of the Visual Geometry Group 16-layer network (VGG16) as its backbone, with a coordinate attention mechanism module cascaded after each convolutional block; the feature fusion layer uses a bidirectional feature pyramid network to weight and fuse feature maps at different scales; the output prediction layer includes a classification head and a regression head, used to output the fruit location point set.
[0044] For example, a fruit detection model can also be a network that only includes the above-mentioned improvements, such as P2PNet-CA which only introduces a coordinate attention mechanism, or P2PNet-BiFPN which only uses BiFPN for feature fusion.
[0045] The following section provides a detailed explanation of each improvement point in conjunction with the accompanying drawings.
[0046] The feature extraction layer consists of the first 13 layers of the VGG16 backbone network, and a coordinate attention mechanism module is cascaded after each convolutional block of the VGG16 backbone network to obtain feature maps that retain spatial location information.
[0047] VGG16 is a classic convolutional neural network architecture. Its first 13 convolutional layers effectively extract features such as edges, textures, and local shapes from images, providing rich semantic information for subsequent fruit detection. After each convolutional block, a cascaded coordinate attention mechanism module enhances the spatial location information of the feature maps extracted by VGG16, enabling the network to more accurately focus on the region where the fruit target is located.
[0048] For example, such as Figure 3 As shown, the coordinate attention mechanism module decomposes global pooling into one-dimensional pooling in the horizontal and vertical directions, aggregating features in the two spatial directions respectively. This allows the model to capture long-range dependencies while retaining accurate spatial location information, thus enhancing the fruit detection model's ability to focus on important features.
[0049] Specifically, the feature extraction layer uses the first 13 convolutional layers of the VGG16 backbone network. To enhance the fruit detection model's ability to locate fruits in complex backgrounds, a CA (Carrier Array) is cascaded after each convolutional block of the VGG16 network, such as... Figure 3As shown, this module decomposes global pooling into one-dimensional pooling in the horizontal and vertical directions, aggregating features in both spatial directions to effectively capture long-range dependencies and preserve accurate spatial location information. Compared to traditional channel attention, the CA mechanism can enhance important features while avoiding excessive computational overhead. Experiments show that adding the CA module after each convolutional block, compared to adding it only to the first four convolutional blocks, can further improve the localization accuracy of the fruit detection model for occluded targets.
[0050] The technical advantage of this design is that the fruit detection model can not only capture long-range dependencies between channels, but also enhance the focus on target features in areas obscured by leaves while preserving accurate spatial location information.
[0051] The feature fusion layer employs a bidirectional feature pyramid network to perform bidirectional information flow weighted fusion of feature maps at different scales. Furthermore, the bidirectional feature pyramid network internally uses depthwise separable convolution to process the weighted fused feature maps.
[0052] Specifically, in the feature fusion layer, this application replaces the original Feature Pyramid Network (FPN) with BiFPN. While the FPN can fuse multi-scale features through top-down and bottom-up paths, it lacks the ability to dynamically adjust the contribution of features at each scale, leading to information simplification. For example... Figure 4 As shown, BiFPN introduces a bidirectional information flow, meaning that each layer simultaneously includes both top-down and bottom-up fusion paths. It employs a learnable weighted fusion strategy at each fusion stage—assigning a trainable scalar weight to each layer's feature map. This weight is optimized along with other network parameters during backpropagation, adaptively adjusting the contribution ratio of features at different scales. Furthermore, BiFPN uses depthwise separable convolutions in the processing of the weighted fused feature maps. By separating spatial convolutions from channel convolutions, it significantly reduces the number of computational parameters while maintaining feature expressiveness.
[0053] For example, such as Figure 5 As shown, depthwise separable convolution decomposes traditional convolution into two independent steps: depthwise convolution and pointwise convolution. By separating spatial convolution from channel convolution, it significantly reduces the number of computational parameters while maintaining feature representation capabilities, thereby improving the computational efficiency of the model.
[0054] The result is that while achieving flexible weighted fusion of fruit features at different scales (large and small fruits), the number of network parameters and computational overhead are significantly reduced.
[0055] The output prediction layer includes a classification head and a regression head. The regression head is configured to output a set of fruit location points using the Huber loss function.
[0056] In the regression head section, this application replaces the original squared error loss with Huber loss.
[0057] Specifically, the expression for Huber loss is: in, The Huber loss function value, with δ as the threshold, measures the error cost between the predicted and actual values. δ is a preset threshold; in this embodiment, δ=1.0. When the prediction error is less than or equal to δ, the Huber loss is expressed as a squared error loss, ensuring sensitivity and convergence speed under small errors. When the error exceeds δ, it switches to linear penalty to avoid excessive influence of outliers on gradient updates, thereby enhancing the robustness of the training process. This represents the predicted value from the fruit detection model. This represents the actual label / actual value.
[0058] It should be noted that during the fruit detection model training phase, an orchard image dataset containing Gaussian heatmaps annotating the fruit center points was used to train AD-P2PNet end-to-end. During training, the classification head used the cross-entropy loss function, the regression head used the Huber loss function, and the total loss was the weighted sum of the two. The optimizer used was Adam, with an initial learning rate set to 1×10⁻⁶. -4 The training process lasted 200 epochs with a batch size of 8. After training, the weights of the fruit detection model that performed best on the validation set were saved for actual inference.
[0059] This effectively prevents abnormal noise under complex lighting conditions from causing excessive interference to the gradient update of the fruit detection model, thereby improving the robustness of the fruit detection model's training and inference.
[0060] Step 102: Cluster the fruit trees according to the spatial distribution characteristics of the fruit location information to obtain the fruit tree clusters of each frame image.
[0061] After obtaining the fruit location information of each frame image, this step aggregates the fruit points belonging to the same fruit tree in the same frame image into fruit tree clusters.
[0062] The specific clustering process is as follows.
[0063] Considering that fruit trees in orchards are planted in rows, fruits on the same tree are densely clustered horizontally, while there are significant gaps between different trees in the horizontal direction; vertically, fruits may be loosely distributed due to differences in canopy height, and fruits in different layers of the same tree may undergo significant vertical displacement due to changes in viewing angle. Therefore, this application adopts an improved density-based spatial clustering algorithm (DBSCAN) based on horizontal constraints. In the clustering process, only the horizontal distance of the fruit's location point is used for determination, and the vertical distance is not included in the calculation.
[0064] The specific steps are as follows: First, for any two fruit location points P in the current frame image i (X i ,Y i ) and P j (X j ,Y j ), calculate its distance along the horizontal axis of the image coordinate system: d x (P i ,P j )=∣Xi-Xj∣.
[0065] Secondly, fruit locations with a horizontal distance less than a preset horizontal distance threshold α are grouped into the same candidate cluster. The horizontal distance threshold can be preset based on the average spacing between adjacent fruit trees in the orchard, for example, taking the image plane projection pixel value of 1.5 meters on the image plane.
[0066] Finally, the number of fruit location points contained in each candidate cluster is counted, and candidate clusters with a number greater than the minimum cluster number threshold are determined as the final fruit tree clusters. The minimum cluster number threshold can be set to 3 to avoid misclassifying isolated noise points as fruit tree clusters. It should be noted that the above threshold can be adjusted according to the actual orchard planting density and image resolution, and this application does not impose strict limitations on it.
[0067] Through the clustering process described above, the fruit points in each frame of the image are divided into several fruit tree clusters, with each fruit tree cluster corresponding to the set of fruit points of one fruit tree in the image. These fruit tree clusters will serve as the basic units for extracting spatial structure features and performing cross-frame matching in subsequent steps.
[0068] Step 103: Extract the spatial structure features of each fruit tree cluster, and match the fruit tree clusters in the adjacent previous frame image and the next frame image based on the spatial structure features.
[0069] After clustering fruit tree clusters in a single frame image, this step extracts the spatial structure features of each fruit tree cluster and matches fruit tree clusters in adjacent frames based on these features to identify the same fruit tree in consecutive frames. The specific process is as follows.
[0070] For each fruit tree cluster, the following multidimensional spatial structure features are extracted: ① Number of fruits: Count the number of fruit locations within the fruit tree cluster, denoted as N.
[0071] ② Width-to-height ratio: Calculate the lateral span of the fruit tree cluster W = max(X) i )-min(X i ) and longitudinal span H=max(Y i )-min(Y i The aspect ratio is R=W / H.
[0072] ③ Fruit density: Calculate the convex hull area A of the fruit tree cluster, and the density is D=N / A.
[0073] ④ Topological Features: This application proposes a topological feature extraction method based on Delaunay triangulation to characterize the spatial distribution pattern of fruit points within the tree canopy. The specific steps are as follows: Given a set of fruit location points P = {p1, p2, ..., p...} n}, where p i =(X i ,Y i ) , representing the spatial coordinates of the fruit.
[0074] First, divide the point set P into an upper subset P along the height direction. upper and the subset P lowe r. The division method can be based on the median of the vertical axis or on a preset percentage (e.g., the first 50% is the upper layer and the last 50% is the lower layer).
[0075] Next, perform Delaunay triangulation on the point set P to generate the edge set: Where E is the edge set. It is an edge in the Delaunay triangulation. The edge set E satisfies the empty circle property, that is, the circumcircle of any triangle does not contain any other points. a and b are index values.
[0076] Specifically, for P upper and P lower Perform Delaunay triangulation to generate the upper edge set E. upper and the following set E lowerDelaunay triangulation satisfies the empty circle property, meaning that the circumcircle of any triangle does not contain any other points.
[0077] Calculate the Euclidean lengths of all sides to form a set of side lengths: Where L is the set of all side lengths. For the edge ( The Euclidean length of ) This represents the two-dimensional Euclidean distance.
[0078] Finally, the side length values in the side length set L are discretized into B intervals (e.g., B=32), and a normalized histogram is calculated to serve as a topological feature characterizing the spatial distribution pattern of the fruit. The subset histogram H is then obtained. upper and the histogram of the subset H lower .
[0079] For example, Figure 6 The illustration of topological feature extraction based on Delaunay triangulation provided in this application is as follows: Figure 6 As shown, for the set of fruit points within a fruit tree cluster, it is first divided into upper and lower parts along the height direction, that is... Figure 6 In the upper and lower regions of the fruit, Delaunay triangulation is performed to generate edge sets. Then, the Euclidean length of each edge is calculated and discretized into a histogram of edge length distribution, which serves as a topological feature characterizing the spatial distribution pattern of the fruit.
[0080] This feature can effectively reflect the spatial aggregation, dispersion and arrangement patterns of fruits, and together with features such as fruit quantity, width-to-height ratio, and density, it constitutes a complete characteristic description of the fruit tree cluster.
[0081] Through the above steps, each fruit tree cluster is represented as a feature vector, including the number of fruits, aspect ratio, fruit density, upper topological histogram, and lower topological histogram.
[0082] Furthermore, for fruit tree clusters in the previous and next frames, the number of fruits, aspect ratio, fruit density, and topological features (i.e., histograms of the side length distribution of the upper and lower parts) obtained based on Delaunay triangulation are extracted. By calculating the similarity of the topological features and combining it with the differences in auxiliary features such as fruit number, aspect ratio, and density, a weighted fusion is performed to obtain a comprehensive similarity score. Fruit tree clusters with the highest comprehensive similarity score and below a preset threshold are identified as matching fruit tree clusters, i.e., the same fruit tree in the previous and next frames.
[0083] Step 104: Calculate the spatial translation vector based on the successfully matched fruit tree clusters, and translate the fruit position information of the previous frame image according to the spatial translation vector.
[0084] After completing the cross-frame matching of fruit tree clusters, this step calculates the inter-frame spatial translation vector based on the successfully matched fruit tree cluster pairs, and uses this translation vector to perform an overall translation of the fruit position information of the previous frame image to achieve coarse alignment between the two frames.
[0085] Specifically, for each successfully matched cluster of fruit trees, the difference between the centroid of the cluster in the previous frame and the centroid of the cluster in the next frame is calculated as the initial translation vector. When there are multiple successfully matched clusters of fruit trees, the Random Sample Consensus (RANSAC) algorithm is used to perform robust regression on all initial translation vectors, eliminating possible outliers and obtaining the globally optimal spatial translation vector.
[0086] In one embodiment, the centroid of a fruit tree cluster is obtained by calculating the average of the coordinates of all fruit locations within the cluster. If there are K pairs of successfully matched fruit tree clusters, then K initial translation vectors t1, t2, ..., t_{t+1} are obtained. K The RANSAC algorithm selects interior points and fits the final translation vector t through random sampling, fruit detection model fitting, and consistency checks. This translation vector can represent the globally consistent motion displacement between two frames.
[0087] Understandably, the RANSAC algorithm can effectively suppress noise introduced by mismatches in individual fruit tree clusters, thus improving the robustness of translation estimation. After translation, the fruit position information of the previous frame is mapped to the coordinate system of the next frame, and the fruit positions in the two frames are spatially aligned, facilitating subsequent accurate local matching and deduplication.
[0088] Step 105: Perform local matching between the translated fruit position information and the fruit position information of the next frame image, remove the matched fruit position information and count the unmatched fruit position information to obtain the fruit count result.
[0089] After completing the spatial translation, this step performs fine matching of the fruit points in the previous frame and the fruit points in the next frame within a local range, and eliminates erroneous matches through motion consistency constraints, and finally obtains the counting results.
[0090] (1) Nearest neighbor matching to establish candidate matching pairs.
[0091] For each fruit location point after translation, the nearest fruit location point is found in the next frame image using a spatial indexing method (such as a kd-tree), establishing candidate matching pairs. A greedy matching strategy is employed to ensure that the matching results are one-to-one relationships.
[0092] Among them, the greedy matching strategy is a conflict resolution mechanism used in nearest neighbor matching to ensure a one-to-one matching result. Its core idea is: when multiple points to be matched (the fruit point of the previous frame after translation) compete for the same target point (the fruit point of the next frame), the target point is assigned to the first matching point according to a certain priority order, and the remaining points search for the second nearest neighbor or other alternative matches.
[0093] It is understandable that, in order to avoid multiple translated fruit points matching the same fruit point in the next frame, the greedy matching strategy can process each translated point in a preset order (e.g., according to the horizontal coordinates of the translated points from smallest to largest). For the current translated point, it searches for the first unoccupied fruit point in the next frame among its nearest neighbors, second nearest neighbors, etc., establishes a matching pair, and marks the fruit point in the next frame as occupied. If all nearest neighbors have been occupied, the current translated point is considered unmatched.
[0094] (2) Determine the dominant motion parameters.
[0095] The displacement vectors of all candidate matching pairs are statistically analyzed, and their distance and orientation angle are calculated respectively. Histogram analysis is used to determine the most frequently occurring distance value as the dominant motion distance and the most frequently occurring orientation angle as the dominant motion direction.
[0096] (3) Motion consistency screening.
[0097] Based on the assumption of global consistency in camera motion, candidate matching pairs are filtered. Only matching pairs that simultaneously meet the following conditions are retained as target matching pairs: the deviation of the displacement vector from the dominant motion distance is less than a first threshold, and the deviation of the displacement vector from the dominant motion direction is less than a second threshold. This filtering effectively eliminates erroneous matches caused by factors such as occlusion and detection errors.
[0098] (4) Elimination and counting.
[0099] In the next frame, the fruit locations corresponding to the target pairs are marked as duplicate detections and removed. Unmatched fruit locations in the next frame are included as new detection targets in the count. By accumulating the new detection targets in each frame, the fruit count result for the entire video sequence can be obtained.
[0100] For example, such as Figure 7 As shown, the overall process of cross-frame deduplication includes three main stages: (a) fruit tree-level clustering based on horizontal constraints, which aggregates fruit points into fruit tree clusters; (b) topological feature extraction and fruit tree cluster matching based on Delaunay triangulation, which determines the same fruit tree in the previous and next frames; and (c) fruit-level local matching and deduplication based on motion consistency constraints, which achieves accurate counting.
[0101] The fruit identification and counting method provided in this application improves the matching granularity from individual fruits to fruit tree clusters, utilizes the spatial structural features of fruit trees for cross-frame matching, and combines motion consistency constraints to achieve precise deduplication at the fruit level. This effectively avoids the problem of repeated counting caused by occlusion and changes in viewpoint, and significantly improves the accuracy and robustness of fruit counting in complex orchard environments.
[0102] The fruit identification and counting device provided in this application is described below. The fruit identification and counting device described below can be referred to in correspondence with the fruit identification and counting method described above.
[0103] Figure 8 This is a schematic diagram of the fruit identification and counting device provided in this application. Figure 8 As shown, this application provides a fruit identification and counting device, which may include: The acquisition module 801 is used to acquire the fruit position information of each frame in the video sequence image; Clustering module 802 is used to cluster fruits based on the spatial distribution characteristics of fruit location information to obtain fruit tree clusters for each frame of image; The feature extraction and matching module 803 is used to extract the spatial structure features of each fruit tree cluster and match the fruit tree clusters in the adjacent previous frame image and the next frame image based on the spatial structure features. Translation module 804 is used to calculate the spatial translation vector based on the successfully matched fruit tree clusters, and to translate the fruit position information of the previous frame image according to the spatial translation vector. The counting module 805 is used to locally match the translated fruit position information with the fruit position information of the next frame image, remove the matched fruit position information and count the unmatched fruit position information to obtain the fruit counting result.
[0104] In one embodiment, the acquisition module 801 is specifically used for: The video sequence images are input into a pre-trained fruit detection model, which outputs the fruit location information of each frame. The fruit detection model includes a feature extraction layer, a feature fusion layer, and an output prediction layer; The feature extraction layer includes the first 13 layers of the VGG16 backbone network, and a coordinate attention mechanism module is cascaded after each convolutional block of the VGG16 backbone network to obtain feature maps that retain spatial location information. The feature fusion layer employs a bidirectional feature pyramid network to perform bidirectional information flow weighted fusion of feature maps at different scales. Furthermore, the bidirectional feature pyramid network internally uses depthwise separable convolutions to process the weighted fused feature maps. The output prediction layer includes a classification head and a regression head. The regression head is configured to output a set of fruit location points using the Huber loss function.
[0105] In yet another embodiment, clustering module 802 is specifically used for: Calculate the distance between any two fruit locations along the horizontal axis of the image coordinate system; Fruit locations whose horizontal distance is less than the horizontal distance threshold are grouped into the same candidate cluster; Candidate clusters containing more than the minimum cluster size threshold are identified as fruit tree clusters.
[0106] In yet another embodiment, the spatial structure features include topological features, which are extracted through the following steps: Divide the set of fruit locations within the fruit tree cluster into an upper subset and a lower subset according to the height direction; Perform Delaunay triangulation on the fruit positions in the upper and lower subsets respectively to obtain the corresponding upper and lower edge sets; Calculate the Euclidean lengths of each edge in the upper and lower edge sets, and generate the upper edge length distribution histogram and the lower edge length distribution histogram respectively. The upper and lower side length distribution histograms are used as topological features.
[0107] In yet another embodiment, the translation module 804 is specifically used for: Calculate the centroid difference between the matching fruit tree clusters in the previous frame and the next frame, and use it as the initial translation vector; A random sampling consensus algorithm is used to perform robust regression on all initial translation vectors to remove outliers, thus obtaining spatial translation vectors.
[0108] In yet another embodiment, the counting module 805 is specifically used for: For each translated fruit position point, find the nearest neighbor fruit position point in the next frame image and establish candidate matching pairs; Calculate the displacement vectors for all candidate matching pairs; The dominant motion distance is determined based on the distance distribution of the displacement vector, and the dominant motion direction is determined based on the angular distribution of the displacement vector. Select the matching pairs whose displacement vectors satisfy the first and second conditions from the candidate matching pairs, and use them as the target matching pairs; The first condition is that the deviation between the displacement vector and the dominant motion distance is less than a first threshold, and the second condition is that the deviation between the displacement vector and the dominant motion direction is less than a second threshold.
[0109] In yet another embodiment, the counting module 805 is specifically used for: In the next frame image, mark the fruit position information corresponding to the target matching pair as duplicate detection and remove it; The location information of unmatched fruits in the next frame is included in the count as new detection targets; The newly detected targets in each frame are accumulated to obtain the fruit count result.
[0110] Figure 9 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 9 As shown, the electronic device may include a processor 910, a communication interface 920, a memory 930, and a communication bus 940. The processor 910, communication interface 920, and memory 930 communicate with each other via the communication bus 940. The processor 910 can call logical instructions in the memory 930 to execute a fruit recognition and counting method. This method includes: acquiring fruit position information of each frame in a video sequence; clustering the fruit position information according to its spatial distribution characteristics to obtain fruit tree clusters for each frame; extracting the spatial structure features of each fruit tree cluster and matching the fruit tree clusters in adjacent previous and subsequent frames based on these features; calculating a spatial translation vector based on the successfully matched fruit tree clusters and translating the fruit position information of the previous frame according to the spatial translation vector; locally matching the translated fruit position information with the fruit position information of the subsequent frame, removing matched fruit position information, and counting unmatched fruit position information to obtain a fruit counting result.
[0111] Furthermore, the logical instructions in the aforementioned memory 930 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the fruit recognition and counting methods provided by the above methods. The method includes: acquiring fruit position information of each frame in a video sequence image; clustering according to the spatial distribution characteristics of the fruit position information to obtain fruit tree clusters of each frame; extracting the spatial structure features of each fruit tree cluster, and matching the fruit tree clusters in adjacent previous and subsequent frames based on the spatial structure features; calculating a spatial translation vector based on the successfully matched fruit tree clusters, and translating the fruit position information of the previous frame according to the spatial translation vector; locally matching the translated fruit position information with the fruit position information of the subsequent frame, removing the matched fruit position information and counting the unmatched fruit position information to obtain the fruit counting result.
[0113] In another aspect, this application also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the fruit identification and counting methods provided by the above methods. The method includes: acquiring fruit position information of each frame in a video sequence image; clustering the fruit position information according to the spatial distribution characteristics to obtain fruit tree clusters of each frame; extracting the spatial structure features of each fruit tree cluster, and matching the fruit tree clusters in adjacent previous and subsequent frames based on the spatial structure features; calculating a spatial translation vector based on the successfully matched fruit tree clusters, and translating the fruit position information of the previous frame according to the spatial translation vector; locally matching the translated fruit position information with the fruit position information of the subsequent frame, removing the matched fruit position information and counting the unmatched fruit position information to obtain the fruit counting result.
[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0115] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for fruit identification and counting, characterized in that, The method includes: Obtain the fruit position information of each frame in the video sequence; Clustering is performed based on the spatial distribution characteristics of the fruit location information to obtain fruit tree clusters for each frame of the image; Extract the spatial structure features of each fruit tree cluster, and match the fruit tree clusters in adjacent previous and next frame images based on the spatial structure features; Calculate the spatial translation vector based on the successfully matched fruit tree cluster, and translate the fruit position information of the previous frame image according to the spatial translation vector; The translated fruit position information is locally matched with the fruit position information of the next frame image. The matched fruit position information is removed and the unmatched fruit position information is counted to obtain the fruit count result.
2. The fruit identification and counting method according to claim 1, characterized in that, The step of obtaining the fruit position information of each frame in the video sequence includes: The video sequence images are input into a pre-trained fruit detection model, which outputs the fruit location information for each frame. The fruit detection model includes a feature extraction layer, a feature fusion layer, and an output prediction layer; The feature extraction layer includes the first 13 layers of the VGG16 backbone network, and a coordinate attention mechanism module is cascaded after each convolutional block of the VGG16 backbone network to obtain feature maps that retain spatial location information. The feature fusion layer employs a bidirectional feature pyramid network to perform bidirectional information flow weighted fusion of feature maps at different scales, and the bidirectional feature pyramid network internally uses depthwise separable convolution to process the weighted fused feature maps. The output prediction layer includes a classification head and a regression head, and the regression head is configured to output a set of fruit location points using the Huber loss function.
3. The fruit identification and counting method according to claim 1, characterized in that, The step of clustering based on the spatial distribution characteristics of the fruit location information to obtain fruit tree clusters for each frame of the image includes: Calculate the distance between any two fruit locations along the horizontal axis of the image coordinate system; Fruit locations whose horizontal distance is less than the horizontal distance threshold are grouped into the same candidate cluster; Candidate clusters containing more than the minimum cluster size threshold are identified as fruit tree clusters.
4. The fruit identification and counting method according to claim 1, characterized in that, The spatial structure features include topological features, which are extracted through the following steps: The set of fruit location points within the fruit tree cluster is divided into an upper subset and a lower subset according to the height direction; Delaunay triangulation is performed on the fruit position points in the upper subset and the lower subset respectively to obtain the corresponding upper edge set and lower edge set; Calculate the Euclidean length of each edge in the upper edge set and the lower edge set, and generate the upper edge length distribution histogram and the lower edge length distribution histogram respectively; The upper side length distribution histogram and the lower side length distribution histogram are used as the topological features.
5. The fruit identification and counting method according to claim 1, characterized in that, The calculation of the spatial translation vector based on the successfully matched fruit tree cluster includes: Calculate the centroid difference of the matching fruit tree clusters in the previous frame image and the next frame image, and use it as the initial translation vector; A random sampling consensus algorithm is used to perform robust regression on all the initial translation vectors to remove outliers, thus obtaining the spatial translation vector.
6. The fruit identification and counting method according to claim 1, characterized in that, The step of locally matching the translated fruit position information with the fruit position information of the subsequent frame image includes: For each translated fruit position point, find the nearest neighbor fruit position point in the next frame image and establish a candidate matching pair; Calculate the displacement vectors for all the candidate matching pairs; The dominant motion distance is determined based on the distance distribution of the displacement vector, and the dominant motion direction is determined based on the angular distribution of the displacement vector. From the candidate matching pairs, the matching pairs whose displacement vectors satisfy the first condition and the second condition are selected as the target matching pairs; The first condition is that the deviation between the displacement vector and the dominant motion distance is less than a first threshold, and the second condition is that the deviation between the displacement vector and the dominant motion direction is less than a second threshold.
7. The fruit identification and counting method according to claim 6, characterized in that, The obtained fruit counting results include: The fruit position information corresponding to the target matching pair in the next frame image is marked as a duplicate detection and removed; The location information of the unmatched fruit in the next frame image is included in the count as a new detection target; The newly detected targets in each frame of the image are accumulated to obtain the fruit count result.
8. A fruit identification and counting device, characterized in that, include: The acquisition module is used to acquire the fruit position information of each frame in the video sequence image; The clustering module is used to cluster fruits based on the spatial distribution characteristics of the fruit location information to obtain fruit tree clusters for each frame of the image; The feature extraction and matching module is used to extract the spatial structure features of each fruit tree cluster, and to match the fruit tree clusters in adjacent previous and next frame images based on the spatial structure features. The translation module is used to calculate a spatial translation vector based on the successfully matched fruit tree clusters, and to translate the fruit position information of the previous frame image according to the spatial translation vector. The counting module is used to locally match the translated fruit position information with the fruit position information of the next frame image, remove the matched fruit position information and count the unmatched fruit position information to obtain the fruit counting result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the fruit identification and counting method as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fruit identification and counting method as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fruit identification and counting method as described in any one of claims 1 to 7.