A fruit tree image acquisition method and system based on two-dimensional skeleton and viewpoint planning
By using a two-dimensional skeleton and viewpoint planning method, the problems of low resolution and collision risk in fruit tree image acquisition are solved, realizing efficient and complete fruit tree image acquisition and real-time operation, while reducing computational complexity and collision risk.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NORTHWEST A & F UNIV
- Filing Date
- 2026-05-26
- Publication Date
- 2026-06-23
AI Technical Summary
Existing fruit tree image acquisition technologies suffer from limitations in optical resolution, physical occlusion, and poor adaptability based on geometric contouring methods, resulting in insufficient image clarity and the risk of mechanical collisions. Furthermore, 3D reconstruction calculations are slow and time-consuming.
A method based on two-dimensional skeleton and viewpoint planning is adopted. By acquiring tree canopy images with depth information, a binary mask and depth information are generated, which are then refined and pruned. Feature points are identified and the skeleton is reconstructed. Viewpoint planning and autonomous image acquisition are then performed. Autonomous image acquisition is achieved using a multi-degree-of-freedom robotic arm and an image acquisition module.
It solves the problems of low resolution and collision risk in fruit tree acquisition technology, realizes the complete image acquisition of the tree canopy and meets the requirements of real-time operation, reduces the amount of computation, and ensures safety boundaries and image coverage integrity.
Smart Images

Figure CN122265609A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of smart agricultural equipment technology, specifically relating to a method and system for fruit tree image acquisition based on two-dimensional skeleton and viewpoint planning. Background Technology
[0002] With the rapid development of modern smart agriculture, orchard management is shifting from extensive to intensive and intelligent methods. In the production of economic fruit trees such as apples, high-precision image acquisition of the tree canopy is fundamental to subsequent intelligent operations. For example, pest and disease control requires early detection of tiny spots on the underside of leaves, and non-destructive harvesting requires precise location of branches for path planning. These tasks all demand that the acquisition system provide high-resolution, multi-angle, and full-coverage canopy image information.
[0003] Currently, most mainstream orchard inspection robots employ a mobile chassis and a fixed gimbal camera. Their typical operating mode involves the robot traveling in a straight line along the centerline between rows of fruit trees, while the camera captures images of the trees on either side using a fixed pitch angle or horizontal scanning method. Furthermore, existing agricultural automation equipment, such as automatic target sprayers and mechanical pruning arms, typically employs a geometric contouring strategy. This involves using simple distance sensors to detect the distance to the tree canopy surface and controlling the end effector to maintain a constant distance from the canopy's outer contour or to scan along a pre-defined geometric surface. The mainstream technological approach in academia and industry for obtaining plant skeletons primarily relies on high-precision 3D reconstruction techniques, such as point cloud scanning based on LiDAR or dense reconstruction based on multi-view stereo vision. These techniques generate high-precision 3D models through data acquisition, preprocessing, and skeleton extraction.
[0004] However, existing technologies suffer from two significant problems. First, there are limitations in optical resolution and physical occlusion. Cameras are typically positioned far from the tree canopy and use wide-angle lenses, resulting in blurred images of fine branches and tiny lesions deep within the canopy at long distances and large fields of view, making feature extraction difficult. Furthermore, observation from a fixed, external two-dimensional perspective inevitably leads to foreground branches and leaves obscuring the background, preventing the acquisition of complete canopy image information. Second, geometric contouring methods are poorly adapted to complex structures, unable to penetrate deep into the recessed areas of the canopy, and lack prior knowledge of branch orientation. Close-range movements easily cause rigid collisions between the robotic arm and the trunk, forcing existing equipment to increase operating safety distances at the expense of image clarity. On the other hand, technologies relying on high-precision 3D reconstruction suffer from severe computational lag, and data processing is complex and time-consuming. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a method and system for fruit tree image acquisition based on two-dimensional skeleton and viewpoint planning. The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a method for fruit tree image acquisition based on two-dimensional skeleton and viewpoint planning, comprising: Step 1: Obtain a tree canopy image with depth information of the target tree canopy; based on the tree canopy image with depth information, obtain a binarized mask of the foreground branches and the depth information corresponding to the binarized mask. Step 2: The binarized mask is refined and trimmed sequentially to obtain an initial two-dimensional skeleton. Feature points are identified based on the initial two-dimensional skeleton, and artifact removal is performed on the feature points to obtain optimized feature points. The growth direction vector of the optimized feature points is calculated. Step 3: Based on the optimized feature points and the growth direction vector, the broken skeleton fragments in the initial two-dimensional skeleton are recombined according to the preset connection screening conditions and exemption region mechanism to obtain a two-dimensional skeleton diagram with complete topological relationships. Step 4: Based on the two-dimensional skeleton diagram, viewpoint planning is performed to generate a viewpoint sequence on a preset two-dimensional reference depth plane. The viewpoint sequence is then mapped to three-dimensional space to obtain the corresponding pose sequence. Based on the pose sequence, autonomous image acquisition of the target tree canopy is achieved.
[0006] This invention also provides a fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning, comprising: Mounting base; A multi-degree-of-freedom robotic arm, mounted on the mounting base, is used to drive the camera to move according to a planned pose sequence; An image acquisition module, installed at the end of the multi-degree-of-freedom robotic arm, is used to acquire image data and depth information of the fruit tree; The control unit is communicatively connected to the multi-degree-of-freedom robotic arm and the image acquisition module, respectively, and is used to receive the image data and the depth information and execute the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning as described in any of the above embodiments to obtain a pose sequence. Based on the pose sequence, the control unit controls the multi-degree-of-freedom robotic arm to drive the image acquisition module to achieve autonomous image acquisition of the target tree canopy.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention presents a fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning, which solves the problems of "low resolution in long-distance shooting" and "collision-prone close-range shooting based on geometric shape" in existing fruit tree acquisition technologies, as well as the problem of lagging three-dimensional reconstruction calculations. This invention uses a two-dimensional skeleton topology deduction algorithm to logically repair visual breaks caused by branch and leaf occlusion, achieving the perception of branches broken due to canopy self-occlusion; it utilizes a two-dimensional reference depth plane to reduce the dimensionality of complex spatial planning, significantly reducing the computational load of skeleton reconstruction and path planning, meeting the real-time operation requirements of edge devices; and it ensures, from the geometric principles of perspective projection, that the actual physical field of view always covers the planned reference field of view, avoiding missed detections while constructing a collision-avoidance safety boundary.
[0008] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described in detail below with reference to the accompanying drawings. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of a fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning provided by an embodiment of the present invention; Figure 2 This is a flowchart of feature point recognition and two-dimensional skeleton graph topology reconstruction provided in an embodiment of the present invention; Figure 3 This is a flowchart of a viewpoint sequence generated based on a greedy strategy, provided in an embodiment of the present invention. Figure 4 This is a flowchart of a process for controlling a robotic arm to achieve autonomous image acquisition based on a depth probing mechanism, provided in an embodiment of the present invention. Figure 5 This is a schematic diagram of a fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning provided in an embodiment of the present invention. Detailed Implementation
[0010] To further illustrate the technical means and effects adopted by the present invention to achieve the intended purpose, the following describes in detail, with reference to the accompanying drawings and specific embodiments, a method and system for fruit tree image acquisition based on two-dimensional skeleton and viewpoint planning proposed in accordance with the present invention.
[0011] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of specific embodiments in conjunction with the accompanying drawings. Through the description of the specific embodiments, a more in-depth and concrete understanding can be gained of the technical means and effects adopted by the present invention to achieve its intended purpose. However, the accompanying drawings are for reference and illustration only and are not intended to limit the technical solutions of the present invention.
[0012] In a first aspect, embodiments of the present invention provide a method for acquiring fruit tree images based on a two-dimensional skeleton and viewpoint planning. This method can be executed by a multi-degree-of-freedom robotic arm mounted on a mounting base and an image acquisition module mounted at its end effector, in conjunction with a control unit.
[0013] Please see Figure 1 , Figure 1 This is a schematic diagram of a fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning provided by an embodiment of the present invention. Figure 1 As shown, the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning in this embodiment may include the following steps: Step 1: Obtain a tree canopy image with depth information of the target tree canopy. Based on the tree canopy image with depth information, obtain the binarized mask of the foreground branches and the corresponding depth information of the binarized mask.
[0014] Specifically, the multi-degree-of-freedom robotic arm can be controlled to move to a preset global shooting posture, ensuring that the optical axis of the image acquisition module installed at the end of the multi-degree-of-freedom robotic arm is perpendicular to the central plane of the tree canopy, and that the field of view can cover the entire tree canopy. This image acquisition module is then used to acquire a tree canopy image with depth information of the target tree.
[0015] Optionally, the multi-degree-of-freedom robotic arm can be a six-degree-of-freedom robotic arm. For example, the global shooting posture is the joint angle: To simultaneously capture high-resolution texture details and depth information, the image acquisition module can be a wide-band near-infrared (Vis-NIR) camera and a depth camera. This image acquisition module can acquire Vis-NIR grayscale images and depth images with a resolution of 5120×5120. Using pre-calibrated intrinsic and extrinsic parameter matrices, the depth image and the Vis-NIR grayscale image are pixel-level aligned through reprojection transformation to generate a tree canopy image data frame with depth information.
[0016] It is understandable that using a multimodal sensor combination in the image acquisition module is not a necessary condition. In practical applications, depending on the camera's performance specifications and the specific accuracy requirements of the task, a single RGB-D (Red, Green, Blue, and Depth, capturing both color images and depth information) camera with sufficient imaging accuracy required by the segmentation network can be used to directly output aligned image and depth data, in which case there is no need to perform complex cross-modal alignment steps.
[0017] In this embodiment, the process of obtaining the binarized mask of the foreground branches includes: performing semantic segmentation on the tree canopy image to obtain a probability map of the foreground branches; performing binarization on the probability map to obtain an initial mask of the foreground branches; and performing morphological processing and edge smoothing on the initial mask to obtain the binarized mask of the foreground branches.
[0018] For example, the tree canopy image can be input into a pre-trained semantic segmentation network (e.g., YOLOv8l-seg) for semantic segmentation processing, outputting a probability map of the foreground branches. It's worth noting that although YOLOv8l-seg is an instance segmentation network, it can still directly output a semantic segmentation mask that does not contain instance information. The probability map is then binarized, for example, with a threshold set to 0.5, to obtain the initial mask for the foreground branches. Understandably, due to the limitations of model computing power, the output branch mask size is 1600×1600. It should be noted that this invention does not impose specific resolution requirements; the resolution depends on the specific computing power.
[0019] To eliminate segmentation noise and holes, the initial mask needs to undergo morphological processing. For example, two closing operations are performed using elliptical structuring elements with a side length of 11 pixels to effectively fill the holes inside the branches and bridge minor breaks. Subsequently, a 5×5 Gaussian filter kernel is used to smooth the mask edges, eliminating jagged noise and obtaining the final binarized mask for the foreground branches. Simultaneously, based on pixel-level alignment, the depth information corresponding to each pixel in this binarized mask can be obtained.
[0020] Step 2: The binarized mask is refined and trimmed sequentially to obtain an initial two-dimensional skeleton. Feature points are identified based on the initial two-dimensional skeleton, and artifact removal is performed on the feature points to obtain optimized feature points. The growth direction vector of the optimized feature points is then calculated.
[0021] In this embodiment, step 2 includes: Step 2.1: Use a thinning algorithm to thin the binary mask and extract the central skeleton. Use a burr trimming algorithm to trim the burrs on the central skeleton to obtain the initial two-dimensional skeleton.
[0022] For example, the Guo-Hall parallel thinning algorithm can be applied to the binarized mask to extract the central skeleton with a width of one pixel. Then, a burr trimming algorithm is executed to iteratively remove isolated short branches with a length less than a preset threshold (e.g., 15 pixels) to reduce topological complexity and obtain the initial two-dimensional skeleton.
[0023] Step 2.2: Traverse the pixels of the initial two-dimensional skeleton, and determine the feature points according to the 8-neighbor connectivity of the pixels and the preset feature point classification rules. The feature points include endpoints, intersections and inflection points.
[0024] Specifically, the pixels of the initial two-dimensional skeleton are traversed, the 8-neighbor connectivity N of each pixel is calculated, and feature points are determined according to a preset feature point classification rule, which includes: When N=1, the current pixel is determined to be an endpoint; When N≥3, the current pixel is determined to be an intersection point; When N=2 and the local angle change along the skeleton path exceeds the preset angle change threshold, the current pixel is determined to be an inflection point.
[0025] For example, the angle change threshold is 45°. In this embodiment, the angle change threshold adjustment range is 30°-60°.
[0026] Step 2.3: Cluster the intersection points, and merge intersection points with a distance less than the preset clustering threshold into a single physical intersection point to obtain the optimized feature points.
[0027] Specifically, the Euclidean distance between all pairs of intersection points is calculated, and intersection points with a distance less than a preset clustering threshold, such as 4 pixels, are grouped into one class. The geometric center of this cluster is calculated as the unique physical intersection point, and a mapping relationship between the representative point and the original cluster members is established to eliminate topological redundancy and obtain optimized feature points.
[0028] Step 2.4: For the optimized feature points, trace back a preset number of pixels along the skeleton path, and calculate the growth direction vector of each feature point through line fitting.
[0029] Specifically, for the optimized feature points, a preset number of pixels are traced along the skeleton path. For example, 9 pixels. The growth direction vector of each feature point is calculated by least-squares line fitting. For endpoints, the tracing is performed inward along the skeleton; for intersections, the tracing is performed outward along each branch connected to them, resulting in a set of multi-branch direction vectors. If the tracing path length is less than a preset number of pixels, such as 2 pixels, then the direction is marked as invalid.
[0030] Step 3: Based on the optimized feature points and growth direction vectors, the broken skeleton fragments in the initial two-dimensional skeleton are recombined according to the preset connection screening conditions and exemption region mechanism to obtain a two-dimensional skeleton diagram with complete topological relationships.
[0031] This step aims to repair skeleton breaks caused by occlusion, employing a hierarchical connection strategy of "Y-shaped first, then linear." For the detailed 2D skeleton graph topology reconstruction process, please refer to [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart of feature point recognition and two-dimensional skeleton graph topology reconstruction provided in an embodiment of the present invention.
[0032] In this embodiment, step 3 includes: Step 3.1: Traverse all the triples formed by the endpoints of the feature points, and filter the endpoint triples that satisfy the Y-shaped connection conditions to form Y-shaped connection candidates.
[0033] The conditions for a Y-shaped connection include the following three conditions: 1. The Euclidean distance between the endpoints is less than a preset first threshold. For example, the first threshold is a pixel distance of 10%-15% of the image width, approximately 245 pixels.
[0034] 2. The largest interior angle of the triangle formed by the three endpoints is less than a preset first angle threshold. For example, the first angle threshold is 100°.
[0035] 3. The growth direction vectors of the three endpoints point to the geometric center of the triangle and the deviation is less than the preset angle tolerance. For example, the angle tolerance is set to 55°. In this embodiment, the adjustable range of the angle tolerance is 45-60°.
[0036] Step 3.2: Traverse the feature point pairs formed by the remaining feature points, filter out feature point pairs that meet the linear connection condition, and form linear connection candidates.
[0037] In this embodiment, the remaining feature points include endpoints, intersections, and inflection points. The linear connection condition includes the following three conditions: 1. The distance between feature point pairs is less than the first threshold.
[0038] 2. The angle between the direction of the line connecting the feature point pairs and the growth direction of each feature point is less than a preset second angle threshold. For example, the second angle threshold is 35°.
[0039] 3. Feature point pairs belong to different connected regions. Different connected regions are broken segments that are not yet connected by the skeleton.
[0040] Step 3.3: Construct an exemption region mechanism. Use the exemption region mechanism to perform obstacle detection on each candidate connection in the Y-shaped connection candidate and the linear connection candidate, and remove the candidate connections with occlusion.
[0041] Specifically, the exemption zone mechanism includes: performing an expansion operation on the connection points and their neighboring pixels in the candidate connection to generate an exemption zone; if all obstacle pixels on the connection path of the candidate connection are within the exemption zone, it is considered unobstructed and the candidate connection is retained; otherwise, the candidate connection is discarded.
[0042] In this embodiment, obstacle pixels are other skeleton pixels.
[0043] Step 3.4: Sort all remaining candidate connections in ascending order of Euclidean distance, establish virtual connections sequentially using a greedy strategy, and update the state of connected regions until all candidate connections have been processed, resulting in a two-dimensional skeleton graph.
[0044] Specifically, all candidate connections are sorted in ascending order of Euclidean distance. Virtual connections are then attempted sequentially: connecting lines are drawn on the binary graph, and the intersection of newly generated connections with existing skeleton lines is checked in real-time, using an exemption zone mechanism for detection. If no conflict is found, the connection is accepted, and the connected region state is updated, merging two previously independent connected regions into one. This process is repeated until all candidate connections have been processed, ultimately resulting in a two-dimensional skeleton graph with complete topological relationships.
[0045] Step 4: Based on the two-dimensional skeleton graph, viewpoint planning is performed to generate a viewpoint sequence on the preset two-dimensional reference depth plane. The viewpoint sequence is then mapped to three-dimensional space to obtain the corresponding pose sequence. Based on the pose sequence, autonomous image acquisition of the target tree canopy is achieved.
[0046] In this embodiment, step 4 includes: Step 4.1: Based on the camera parameters and preset working distance, calculate the field of view coverage size on the two-dimensional reference depth plane, which serves as the smallest unit for viewpoint planning.
[0047] Specifically, based on the camera sensor's physical size (e.g., 12.5mm), lens focal length (e.g., 12mm), and a preset traversal distance (e.g., 0.25m), combined with the global shooting distance (e.g., 1.5m), the field of view coverage size on the two-dimensional reference depth plane, including width and height, is calculated as the smallest unit for subsequent viewpoint planning. In this embodiment, the two-dimensional reference depth plane is defined as a virtual plane located at the tree trunk depth or the global maximum effective depth.
[0048] Step 4.2: Discretize the two-dimensional skeleton graph and use a greedy iterative algorithm to search for the optimal viewpoint. Use coverage gain and overlap rate as scoring indicators to generate the optimal viewpoint set that satisfies the overlap rate constraint.
[0049] Specifically, the two-dimensional skeleton map obtained in step 3 is discretized and sampled, for example, with a step size of 5 pixels, to construct a set of points to be covered. A sliding window with a field-of-view coverage size of step size as the smallest unit is used on the two-dimensional reference depth plane. The number of uncovered skeleton pixels contained in the current window is calculated as the coverage gain. Simultaneously, the overlap ratio between the current window and the already generated viewpoint windows is calculated and converted into an overlap rate score. According to coverage gain and overlap rate score Calculate the overall score. Select the window with the highest score as the optimal viewpoint, add it to the optimal viewpoint set, and mark the skeleton pixels within that window as covered. Repeat the search until all skeleton points are covered, thus obtaining the optimal viewpoint set.
[0050] In this embodiment, the overlap rate score is converted based on a Gaussian function, with the target overlap rate set to 10%. The formula for calculating the comprehensive score is as follows: ,in and For example, the weighting coefficients are... , Furthermore, the coverage gain in the overall score calculation formula... This is a value that has already been normalized. The normalization method is to divide the total number of uncovered skeleton pixels in the window by the total number of skeleton pixels in the window.
[0051] Step 4.3: Sort the optimal viewpoint set using a path optimization algorithm to obtain the viewpoint sequence.
[0052] Optionally, the generated unordered optimal viewpoint set can be sorted using the nearest neighbor algorithm to obtain the viewpoint sequence, i.e., the robotic arm movement sequence.
[0053] It is understood that in this embodiment, the path optimization algorithm is not limited, and other path optimization algorithms may be used provided that the computation time is controllable.
[0054] The viewpoint sequence generation process in this embodiment is as follows: Figure 3 As shown, Figure 3 This is a flowchart of a viewpoint sequence generated based on a greedy strategy, provided in an embodiment of the present invention.
[0055] Step 4.4: Map each optimal viewpoint in the viewpoint sequence to its corresponding 3D spatial pose to form a pose sequence. Control the robotic arm to perform motion according to the pose sequence, and perform depth exploration or pose reset when the viewpoint is unreachable or an anomaly is triggered.
[0056] In this embodiment, mapping the optimal viewpoint to the corresponding three-dimensional spatial pose includes: extracting the preset quantile values of the depth distribution of all branches within the coverage area of the optimal viewpoint as the safe depth, and mapping the viewpoint coordinates of the optimal viewpoint and the corresponding safe depth to the three-dimensional spatial pose through pose inverse solution and safety buffer compensation mechanism.
[0057] The process of controlling the robotic arm to achieve autonomous image acquisition in this embodiment is as follows: Figure 4 As shown, Figure 4 This is a flowchart illustrating the autonomous image acquisition process achieved by controlling a robotic arm based on a depth probing mechanism, as provided in this embodiment of the invention. Specifically, the depth distribution of all branches within the coverage area of the optimal viewpoint is extracted, and a preset quantile value, such as the 8th percentile, is taken as the safe depth. To filter foreground noise, the viewpoint coordinates of the optimal viewpoint are determined using the camera intrinsic inverse matrix and the hand-eye calibration matrix. With security depth The inverse kinematics solution is obtained by reversing the robot arm's position in the base coordinate system and moving backward along the optical axis by a safe buffer distance (e.g., 0.08m) to obtain the final 3D spatial pose. The robot arm is then controlled to move to this 3D spatial pose. If the inverse kinematics solution fails due to unreachability, a depth exploration mechanism is automatically initiated: attempting up to four suboptimal positions sequentially along the optical axis with a step size of 0.08m. If a feasible solution is found, it is executed; otherwise, the viewpoint is skipped. If a joint limit or singularity error is triggered during movement, the control system automatically drives the robot arm back to the preparatory posture, resets the joint state, and re-attempts to execute the current command, ensuring long-term stable operation of the system. In this embodiment, the preparatory posture is the joint angle: .
[0058] The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning in this invention solves the problems of "low resolution for long-distance shooting" and "easy collision for close-range shooting based on geometric shape" in existing fruit tree acquisition technologies, as well as the problem of lagging three-dimensional reconstruction calculations. This invention uses a two-dimensional skeleton topology deduction algorithm to logically repair visual breaks caused by branch and leaf occlusion, achieving the perception of branches broken due to canopy self-occlusion; by using a two-dimensional reference depth plane to reduce the dimensionality of complex spatial planning, it can significantly reduce the computational load of skeleton reconstruction and path planning, meeting the real-time operation requirements of edge devices; from the geometric principle of perspective projection, it can strictly ensure that the actual physical field of view always covers the planned reference field of view, avoiding missed detections while constructing a strict anti-collision safety boundary.
[0059] Secondly, embodiments of the present invention provide a fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning, applicable to the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning provided in the first aspect. Please refer to... Figure 5 , Figure 5 This is a schematic diagram of a fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning, provided by an embodiment of the present invention. Figure 5 As shown, the fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning in this embodiment includes: a mounting base, a multi-degree-of-freedom robotic arm, an image acquisition module, and a control unit.
[0060] Specifically, the mounting base supports other components of the system. A multi-degree-of-freedom robotic arm is mounted on the mounting base and used to move the camera according to a planned pose sequence. An image acquisition module is mounted at the end of the multi-degree-of-freedom robotic arm and is used to acquire image data and depth information of the fruit tree. The control unit is communicatively connected to both the multi-degree-of-freedom robotic arm and the image acquisition module. It receives image data and depth information and executes the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning provided in the first aspect to acquire a pose sequence. Based on the pose sequence, it controls the multi-degree-of-freedom robotic arm to drive the image acquisition module to achieve autonomous image acquisition of the target tree canopy.
[0061] Optionally, the mounting base can be a fixed platform or a mounting bracket integrated into a mobile vehicle, such as a four-wheel differential mobile chassis, for transporting the system to the work position. The multi-DOF robotic arm can be the JAKA Zu 5 six-DOF collaborative robot, with a repeatability of ±0.03mm, a payload of 7kg, and a working radius of 819mm. The image acquisition module can adopt an "eye-in-hand" configuration to acquire RGB-D image data of the fruit trees. This image acquisition module can include a high-resolution camera (e.g., a Hikvision MV-CH250-90GN near-infrared camera with a 12mm focal length lens) and a depth camera (e.g., an Intel RealSense D435if), achieving pixel-level alignment through pre-calibrated intrinsic and extrinsic parameters; alternatively, an integrated RGB-D camera (e.g., a ZED 2 or Azure Kinect) can be used to directly output aligned images and depth data. The control unit can be a high-performance industrial computer equipped with an NVIDIA GPU (graphics processing unit), communicating with the robot controller via Ethernet.
[0062] In this embodiment, the system needs to have completed high-precision parameter calibration beforehand: the intrinsic parameter matrix of the NIR camera is determined as follows: The pose vector of the tool coordinate system relative to the flange is obtained through hand-eye calibration. (Units: mm, rad). The operating parameters are set as follows: global shooting distance is 1.5m, and close-up traversal distance is 0.25m.
[0063] The workflow of the fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning in this embodiment of the invention is as follows: The control unit first drives the robotic arm to move to the preset global shooting posture, and the image acquisition module acquires the tree canopy image with depth information; the control unit executes steps 1 to 4 to generate a two-dimensional skeleton diagram with complete topological relationships, plans the viewpoint sequence on the two-dimensional reference depth plane and maps it to a three-dimensional pose sequence; finally, the control unit controls the robotic arm to move to each target pose in sequence, and performs adaptive closed-loop control based on real-time depth information during the movement to achieve safe and full-coverage image information acquisition.
[0064] For details regarding the fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning, as well as its corresponding beneficial effects, please refer to the relevant content on the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning provided in the first aspect; it will not be elaborated upon here.
[0065] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations are intended to cover non-exclusive inclusion, such that an article or device comprising a list of elements includes not only those elements but also other elements not expressly listed. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the article or device comprising said element. Terms such as "connected" or "linked" are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect. The orientations or positional relationships indicated by terms such as "upper," "lower," "left," and "right" are based on the orientations or positional relationships shown in the accompanying drawings and are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0066] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. In addition, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.
[0067] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. A method for acquiring fruit tree images based on two-dimensional skeleton and viewpoint planning, characterized in that, include: Step 1: Obtain a tree canopy image with depth information of the target tree canopy; based on the tree canopy image with depth information, obtain a binarized mask of the foreground branches and the depth information corresponding to the binarized mask. Step 2: The binarized mask is refined and trimmed sequentially to obtain an initial two-dimensional skeleton. Feature points are identified based on the initial two-dimensional skeleton, and artifact removal is performed on the feature points to obtain optimized feature points. The growth direction vector of the optimized feature points is calculated. Step 3: Based on the optimized feature points and the growth direction vector, the broken skeleton fragments in the initial two-dimensional skeleton are recombined according to the preset connection screening conditions and exemption region mechanism to obtain a two-dimensional skeleton diagram with complete topological relationships. Step 4: Based on the two-dimensional skeleton diagram, viewpoint planning is performed to generate a viewpoint sequence on a preset two-dimensional reference depth plane. The viewpoint sequence is then mapped to three-dimensional space to obtain the corresponding pose sequence. Based on the pose sequence, autonomous image acquisition of the target tree canopy is achieved.
2. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 1, characterized in that, The process of obtaining the binarization mask for the foreground branches includes: The tree canopy image is semantically segmented to obtain a probability map of the foreground branches. The probability map is then binarized to obtain an initial mask for the foreground branches. The initial mask is then subjected to morphological processing and edge smoothing to obtain a binarized mask for the foreground branches.
3. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 1, characterized in that, Step 2 includes: Step 2.1: Use a thinning algorithm to thin the binary mask and extract the central skeleton. Use a burr trimming algorithm to trim the burrs on the central skeleton to obtain the initial two-dimensional skeleton. Step 2.2: Traverse the pixels of the initial two-dimensional skeleton, and determine the feature points according to the 8-neighbor connectivity of the pixels and the preset feature point classification rules. The feature points include endpoints, intersections and inflection points. Step 2.3: Cluster the intersection points, and merge intersection points with a distance less than a preset clustering threshold into one physical intersection point to obtain the optimized feature points; Step 2.4: For the optimized feature points, trace a preset number of pixels along the skeleton path, and calculate the growth direction vector of each feature point through line fitting.
4. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 3, characterized in that, The feature point classification rules include: When the 8-neighbor connectivity of a pixel is 1, the current pixel is determined to be an endpoint. If the 8-neighbor connectivity of a pixel is greater than or equal to 3, then the current pixel is determined to be an intersection point. When the 8-neighbor connectivity of a pixel is 2 and the local angle change along the skeleton path exceeds a preset angle change threshold, the current pixel is determined to be an inflection point.
5. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 1, characterized in that, Step 3 includes: Step 3.1: Traverse all the triples formed by the endpoints of the feature points, filter the endpoint triples that satisfy the Y-shaped connection conditions, and form Y-shaped connection candidates; Step 3.2: Traverse the feature point pairs formed by the remaining feature points, filter the feature point pairs that satisfy the linear connection condition, and form linear connection candidates; Step 3.3: Construct the exemption region mechanism and use the exemption region mechanism to perform obstacle detection on each candidate connection in the Y-shaped connection candidate and the linear connection candidate, and remove candidate connections with occlusion. Step 3.4: Sort all remaining candidate connections in ascending order of Euclidean distance, establish virtual connections sequentially using a greedy strategy, and update the connected region state until all candidate connections have been processed, thus obtaining the two-dimensional skeleton graph.
6. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 5, characterized in that, The Y-shaped connection conditions include: the Euclidean distance between the endpoints is less than a preset first threshold, the maximum interior angle of the triangle formed by the three endpoints is less than a preset first angle threshold, and the growth direction vectors of the three endpoints point to the geometric center of the triangle and the deviation is less than a preset angle tolerance. The linear connection conditions include: the distance between feature point pairs is less than the first threshold, the angle between the direction of the line connecting the feature point pairs and the growth direction of each feature point is less than a preset second angle threshold, and the feature point pairs belong to different connected regions.
7. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 5, characterized in that, The exemption region mechanism includes: performing an expansion operation on the connection points and their neighboring pixels in the candidate connection to generate an exemption region; if all obstacle pixels on the connection path of the candidate connection are within the exemption region, they are considered unobstructed and the candidate connection is retained; otherwise, the candidate connection is discarded.
8. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 1, characterized in that, Step 4 includes: Step 4.1: Based on the camera parameters and the preset working distance, calculate the field of view coverage size on the two-dimensional reference depth plane, which serves as the smallest unit for viewpoint planning; Step 4.2: Discretize the two-dimensional skeleton graph and use a greedy iterative algorithm to search for the optimal viewpoint. Use coverage gain and overlap rate as scoring indicators to generate the optimal viewpoint set that satisfies the overlap rate constraint. Step 4.3: Sort the optimal viewpoint set using a path optimization algorithm to obtain the viewpoint sequence; Step 4.4: Map each optimal viewpoint in the viewpoint sequence to its corresponding three-dimensional spatial pose in sequence to form the pose sequence. Control the robotic arm to perform movement according to the pose sequence, and perform depth exploration or pose reset when the viewpoint is unreachable or an anomaly is triggered.
9. The fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning according to claim 8, characterized in that, In step 4.4, mapping the optimal viewpoint to the corresponding 3D spatial pose includes: The preset quantile values of the depth distribution of all branches within the coverage area of the optimal viewpoint are extracted as the safe depth. The viewpoint coordinates of the optimal viewpoint and the corresponding safe depth are mapped to the three-dimensional spatial pose through pose inverse solution and safety buffer compensation mechanism.
10. A fruit tree image acquisition system based on two-dimensional skeleton and viewpoint planning, characterized in that, include: Mounting base; A multi-degree-of-freedom robotic arm, mounted on the mounting base, is used to drive the camera to move according to a planned pose sequence; An image acquisition module, installed at the end of the multi-degree-of-freedom robotic arm, is used to acquire image data and depth information of the fruit tree; The control unit is communicatively connected to the multi-degree-of-freedom robotic arm and the image acquisition module, respectively, and is used to receive the image data and the depth information and execute the fruit tree image acquisition method based on two-dimensional skeleton and viewpoint planning as described in any one of claims 1 to 9, obtain the pose sequence, and control the multi-degree-of-freedom robotic arm to drive the image acquisition module to achieve autonomous image acquisition of the target tree canopy based on the pose sequence.