Apple picking robot collision-free grabbing method and system

CN117162113BActive Publication Date: 2026-08-21SHANDONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311183017.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-13
Publication Date
2026-08-21
Estimated Expiration
2043-09-13

AI Technical Summary

Technical Problem

该方法在抓取充分暴露于摄像机视野中的果实时是十分有效的,但是在非结构化且复杂的实际果园环境时,许多果实被枝叶包围,通过末端执行器简单抓取后回拉有可能对树木与机械臂造成损坏,同时,也有可能导致目标果实的意外移动,降低识别与检测效率

Benefits of technology

[0038]本发明采用双深度摄像头协同工作,通过两个摄像头远近景模式的切换可以快速、精准实现采摘机器人移动抓取。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117162113B_ABST
    Figure CN117162113B_ABST
Patent Text Reader

Abstract

The present application provides a kind of apple picking robot collision-free grabbing method and system, it is related to picking robot field, based on the long-range image of long-range camera acquisition, the detection of long-range apple tree is carried out, and the navigation path of robot is generated, guide robot to move to picking range;Based on the close-range image of hand camera acquisition, through collision-free grabbing identification network, accurately identify and locate target apple in picking range, and generate the collision-free grabbing posture of mechanical arm, carry out the obstacle avoidance of apple and grab;Wherein, the collision-free grabbing posture includes grabbing direction and grabbing path, the collision-free grabbing identification network, on the basis of semantic segmentation, generate a obstacle avoidance leaf grabbing direction, and carry out the path planning of mechanical arm, obtain grabbing path;The present application uses double depth camera to work cooperatively, based on point cloud obstacle avoidance and long-short range switching method, carries out collision-free grabbing, improves the efficiency of grabbing and identification, reduces the risk of fruit tree damage.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of harvesting robots, and in particular relates to a collision-free grasping method and system for apple harvesting robots. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Currently, the visual recognition systems of apple-picking robots mostly employ deep learning methods based on RGB images to detect, segment, and locate fruits. After detecting the target, the robotic arm plans a path to approach and separate the fruit from the tree. This method is very effective when grasping fruits that are fully exposed to the camera's field of view. However, in unstructured and complex real-world orchard environments, many fruits are surrounded by branches and leaves. Simply grabbing and pulling back the fruit with an end effector may damage the tree and the robotic arm. Furthermore, it may cause the target fruit to move unexpectedly, reducing recognition and detection efficiency.

[0004] Therefore, the existing solution has problems with obstacle avoidance control, resulting in low accuracy and performance in apple picking. Summary of the Invention

[0005] To overcome the shortcomings of the prior art, this invention provides a collision-free grasping method and system for apple picking robots. It employs dual depth cameras working in tandem and uses point cloud obstacle avoidance and near-far view switching methods to perform collision-free grasping, thereby improving grasping and recognition efficiency and reducing the risk of damage to fruit trees.

[0006] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of this invention provides a collision-free grasping method for apple picking robots.

[0008] A collision-free grasping method for an apple-picking robot includes:

[0009] Based on distant images captured by a distant camera, the robot detects distant apple trees and generates a navigation path to guide the robot to the picking area.

[0010] Based on close-up images captured by a hand camera, the system uses a collision-free grasping recognition network to accurately identify and locate target apples within the picking range, and generates a collision-free grasping posture for the robotic arm to perform obstacle-avoidance grasping of the apples.

[0011] The collision-free grasping posture includes the grasping direction and the grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on semantic segmentation and performs path planning for the robotic arm to obtain the grasping path.

[0012] Furthermore, the distant camera is mounted on the chassis of the harvesting robot, and the hand camera is mounted on the end of the robotic arm;

[0013] Both cameras capture two types of images: color images and depth images.

[0014] Furthermore, a template matching algorithm is used to detect remote apple trees, and a motion planning module is used to generate the robot's navigation path.

[0015] Furthermore, the collision-free grasping and recognition network includes a feature recognition network, a point cloud construction module, and a grasping estimation network;

[0016] The feature recognition network detects the apple mask image from the color image captured by the hand camera;

[0017] The point cloud construction module combines the apple mask image with the depth image to generate a mask depth map, and then converts it into the corresponding point cloud.

[0018] The grasping estimation network generates a collision-free grasping posture for the robotic arm based on point clouds.

[0019] Furthermore, the feature recognition network utilizes the YOLACT network to detect the apple mask image, specifically as follows:

[0020] Image features are extracted using ResNet-50;

[0021] Based on the extracted features, a set of k prototype masks for the entire image are predicted by ProtoNet. The Prediction Head predicts the output as detection boxes, categories and mask coefficients respectively. Then, duplicate detection boxes are removed by fast non-maximum suppression.

[0022] The Protonet and mask coefficient branches are combined using linear weighting, and then the final apple mask image is generated using sigmoid nonlinear activation.

[0023] Furthermore, the specific steps of the grasping estimation network are as follows:

[0024] Based on the results of point cloud semantic segmentation, a grasping direction for an obstacle-avoiding leaf is generated using a multilayer perceptron.

[0025] The path planning of the robotic arm is performed using the RRT sampling algorithm to obtain the grasping path, and finally the collision-free grasping posture of the robotic arm is generated.

[0026] Furthermore, the point cloud semantic segmentation specifically includes:

[0027] Based on the apple and the surrounding branches and leaves, the point cloud is divided into target point cloud and non-target point cloud;

[0028] The target point cloud and the non-target point cloud are input into PointNet respectively to obtain target features and non-target features;

[0029] The two fused features are output as the result of point cloud semantic segmentation.

[0030] The second aspect of the present invention provides a collision-free grasping system for apple picking robots.

[0031] A collision-free grasping system for apple picking robots includes a range navigation module and a posture generation module:

[0032] The range navigation module is configured to: detect distant apple trees based on distant images captured by a distant camera, generate a navigation path for the robot, and guide the robot to move to the picking area;

[0033] The posture generation module is configured to: accurately identify and locate the target apple within the picking range based on the close-up image captured by the hand camera, through a collision-free grasping recognition network, and generate a collision-free grasping posture for the robotic arm to perform obstacle avoidance grasping of the apple.

[0034] The collision-free grasping posture includes the grasping direction and the grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on semantic segmentation and performs path planning for the robotic arm to obtain the grasping path.

[0035] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps of a collision-free grasping method for an apple-picking robot as described in the first aspect of the present invention.

[0036] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a collision-free grasping method for an apple picking robot as described in the first aspect of the present invention.

[0037] The above one or more technical solutions have the following beneficial effects:

[0038] This invention employs dual depth cameras working in tandem. By switching between near and far-view modes of the two cameras, the harvesting robot can quickly and accurately move and grasp objects.

[0039] This invention constructs a collision-free grasping and recognition network. First, it accurately identifies and locates the target apple, then constructs point cloud data, and finally generates a collision-free grasping posture based on the point cloud data. By using the collision-free grasping posture, obstacle avoidance control is performed on the point cloud, which improves the grasping and recognition efficiency and reduces the risk of damage to the fruit tree.

[0040] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0041] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0042] Figure 1 This is a flowchart of the method in the first embodiment.

[0043] Figure 2 This is a schematic diagram showing the camera placement in the first embodiment.

[0044] Figure 3 This is a schematic diagram of the moving and grasping process in the first embodiment.

[0045] Figure 4 This is a structural diagram of the collision-free grasping and recognition network in the first embodiment.

[0046] Figure 5 This is a structural diagram of the feature extraction network in the first embodiment.

[0047] Figure 6 The structure diagram of the estimation network is captured for the first embodiment. Detailed Implementation

[0048] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0049] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Example 1

[0051] In one or more embodiments, a collision-free grasping method for an apple-picking robot is disclosed, such as... Figure 1 As shown, it includes the following steps:

[0052] Step S1: Based on the distant view image captured by the distant camera, detect the distant apple tree and generate the robot's navigation path to guide the robot to the picking area.

[0053] Step S2: Based on the close-up images captured by the hand camera, the target apple is accurately identified and located within the picking range through a collision-free grasping recognition network, and a collision-free grasping posture of the robotic arm is generated to perform obstacle avoidance grasping of the apple.

[0054] The collision-free grasping posture includes the grasping direction and the grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on semantic segmentation and performs path planning for the robotic arm to obtain the grasping path.

[0055] Furthermore, the distant camera is mounted on the chassis of the harvesting robot, and the hand camera is mounted on the end of the robotic arm;

[0056] Both cameras capture two types of images: color images and depth images.

[0057] Furthermore, a template matching algorithm is used to detect remote apple trees, and a motion planning module is used to generate the robot's navigation path.

[0058] Furthermore, the collision-free grasping and recognition network includes a feature recognition network, a point cloud construction module, and a grasping estimation network;

[0059] The feature recognition network detects the apple mask image from the color image captured by the hand camera;

[0060] The point cloud construction module combines the apple mask image with the depth image to generate a mask depth map, and then converts it into the corresponding point cloud.

[0061] The grasping estimation network generates a collision-free grasping posture for the robotic arm based on point clouds.

[0062] Furthermore, the feature recognition network utilizes the YOLACT network to detect the apple mask image, specifically as follows:

[0063] Image features are extracted using ResNet-50;

[0064] Based on the extracted features, a set of k prototype masks for the entire image are predicted by ProtoNet. The Prediction Head predicts the output as detection boxes, categories and mask coefficients respectively. Then, duplicate detection boxes are removed by fast non-maximum suppression.

[0065] The Protonet and mask coefficient branches are combined using linear weighting, and then the final apple mask image is generated using sigmoid nonlinear activation.

[0066] Furthermore, the specific steps of the grasping estimation network are as follows:

[0067] Based on the results of point cloud semantic segmentation, a grasping direction for an obstacle-avoiding leaf is generated using a multilayer perceptron.

[0068] The path planning of the robotic arm is performed using the RRT sampling algorithm to obtain the grasping path, and finally the collision-free grasping posture of the robotic arm is generated.

[0069] Furthermore, the point cloud semantic segmentation specifically includes:

[0070] Based on the apple and the surrounding branches and leaves, the point cloud is divided into target point cloud and non-target point cloud;

[0071] The target point cloud and the non-target point cloud are input into PointNet respectively to obtain target features and non-target features;

[0072] The two fused features are output as the result of point cloud semantic segmentation.

[0073] The implementation process of a collision-free grasping method for an apple-picking robot in this embodiment will be described in detail below.

[0074] This embodiment provides a collision-free grasping method for apple-picking robots based on a combination of point cloud obstacle avoidance and near-far view switching. The hardware requirement for this method is to mount two depth cameras on the picking robot, which has a moving chassis and a robotic arm. Figure 2 As shown.

[0075] 1) The base camera (long-range camera) is installed on the chassis of the picking robot. It has a large field of view and is responsible for recognizing distant scenery, that is, detecting distant apple trees and generating the robot's navigation path to guide the robot to the picking area.

[0076] 2) The hand camera is installed at the end of the robotic arm. With a small field of view, it is responsible for identifying close-up objects, that is, accurately identifying and locating the target apple within the picking range, and generating a collision-free grasping posture for the robotic arm to perform obstacle avoidance grasping of the apple.

[0077] By switching between near and far-view modes using two cameras, the harvesting robot can quickly and accurately move and grasp objects. The specific process is as follows: Figure 3 As shown:

[0078] The system acquires distant color and depth images captured by the base camera. When a distant apple tree is detected using a template matching algorithm, the motion planning module uses the A* algorithm to generate a navigation path and controls the chassis to move to a distance from the fruit tree to the picking range. In this embodiment, the template matching algorithm uses OpenCV, and the motion planning module uses the A* algorithm.

[0079] After the robot moves into the picking area, it acquires close-up color and depth images from the hand camera, accurately identifies and locates the target apple through the YOLACT network, and generates a collision-free grasping posture through a collision-free grasping recognition network.

[0080] Finally, by generating a collision-free grasping posture, the robot's robotic arm is controlled to perform collision-free grasping. The robot has a built-in controller with two functions: one is to control the movement of the robot's chassis based on the navigation path, and the other is to control the robotic arm to grasp the target apple based on the grasping posture.

[0081] The collision-free grasping and recognition network mainly consists of three modules: a feature recognition network, a point cloud construction module, and a grasping estimation network. Figure 4 As shown:

[0082] 1) The feature recognition network uses the YOLACT network model to detect the apple's mask information, i.e., the apple mask image, from the color image captured by the hand camera, such as... Figure 5 As shown, specifically:

[0083] Image feature extraction is performed using ResNet-50: the input color image is processed through a convolutional module to output corresponding feature maps, and each feature map is defined to correspond to layers C3 to C5 in the image.

[0084] In PANet, further convolutions are performed on layers C3 to C5, and the resulting image features are fed into the input Protonet module and the Prediction Head module respectively using a branch-parallel operation.

[0085] The Protonet module predicts and outputs a set of k prototype masks for the entire image (k is defined empirically by the user).

[0086] The Prediction Head module predicts the output bounding box, class, and masking coefficients, respectively.

[0087] Detection 1 and Detection 2 are the masking coefficients obtained by the algorithm after detecting the first and second apples in the image, respectively.

[0088] The Protonet and mask coefficient branches are combined using linear weighting, and then the final apple mask image is generated using sigmoid nonlinear activation.

[0089] 2) The point cloud construction module combines the apple mask image on the color image with the depth image at the corresponding position to generate an environment perception map. Then, the point cloud data corresponding to the environment perception map is fed into the grasping estimation network.

[0090] 3) The grasping estimation network takes the point cloud as input and generates a collision-free grasping pose including the grasping direction and grasping path.

[0091] The generation of the grasping direction is achieved by using the PointNet model to semantically segment the fruit and surrounding branches and leaves, thereby enabling environmental perception analysis of the picking scene, obtaining semantic segmentation results of the picking environment, and using a multilayer perceptron to generate a grasping direction for the obstacle-avoiding leaf. The semantic segmentation results of the picking environment can be transmitted to the ROS visualization interface, making it easier for robot controllers to intuitively observe the picking process.

[0092] The grasping path is generated by using the semantic segmentation results of the picking environment obtained by the grasping estimation network. In the robot's motion planning and control module, the RRT sampling algorithm is used to generate an initial point at the current position of the robotic arm. Then, a random sampling point is generated near the initial point. If the line connecting the initial point and the sampling point does not intersect with the obstacle, the above steps are repeated to generate a new sampling point until a picking path that does not intersect with the picking target is found, thus achieving obstacle avoidance and apple grasping.

[0093] Example 2

[0094] In one or more embodiments, a collision-free grasping system for an apple-picking robot is disclosed, including a range navigation module and a posture generation module:

[0095] The range navigation module is configured to: detect distant apple trees based on distant images captured by a distant camera, generate a navigation path for the robot, and guide the robot to move to the picking area;

[0096] The posture generation module is configured to: accurately identify and locate the target apple within the picking range based on the close-up image captured by the hand camera, through a collision-free grasping recognition network, and generate a collision-free grasping posture for the robotic arm to perform obstacle avoidance grasping of the apple.

[0097] The collision-free grasping posture includes the grasping direction and the grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on semantic segmentation and performs path planning for the robotic arm to obtain the grasping path.

[0098] Example 3

[0099] The purpose of this embodiment is to provide a computer-readable storage medium.

[0100] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a collision-free grasping method for an apple-picking robot as described in Embodiment 1 of this disclosure.

[0101] Example 4

[0102] The purpose of this embodiment is to provide an electronic device.

[0103] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of a collision-free grasping method for an apple picking robot as described in Embodiment 1 of this disclosure.

[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A collision-free grasping method for an apple-picking robot, characterized in that, include: Based on distant images captured by a distant camera, the robot detects distant apple trees and generates a navigation path to guide the robot to the picking area. Based on close-up images captured by a hand camera, the system uses a collision-free grasping recognition network to accurately identify and locate target apples within the picking range, and generates a collision-free grasping posture for the robotic arm to perform obstacle-avoidance grasping of the apples. The collision-free grasping posture includes grasping direction and grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on point cloud semantic segmentation and performs path planning for the robotic arm to obtain the grasping path. The collision-free grasping and recognition network includes a feature recognition network, a point cloud construction module, and a grasping estimation network; the point cloud semantic segmentation specifically includes: Based on the apple and the surrounding branches and leaves, the point cloud is divided into target point cloud and non-target point cloud; The target point cloud and the non-target point cloud are input into PointNet respectively to obtain target features and non-target features; the two features after fusion are output as the result of point cloud semantic segmentation. The specific steps of the capture estimation network are as follows: Based on the results of point cloud semantic segmentation, a grasping direction for an obstacle-avoiding leaf is generated using a multilayer perceptron. The path planning of the robotic arm is performed using the RRT sampling algorithm to obtain the grasping path, and finally the collision-free grasping posture of the robotic arm is generated.

2. The collision-free grasping method for an apple-picking robot as described in claim 1, characterized in that, The distant camera is mounted on the chassis of the harvesting robot, and the hand camera is mounted on the end of the robotic arm; both cameras capture two types of images: color images and depth images.

3. The collision-free grasping method for an apple-picking robot as described in claim 1, characterized in that, The remote apple tree is detected using a template matching algorithm, and the robot's navigation path is generated using a motion planning module.

4. The collision-free grasping method for an apple-picking robot as described in claim 1, characterized in that, The feature recognition network detects the apple mask image from the color image captured by the hand camera; The point cloud construction module combines the apple mask image with the depth image to generate a mask depth map, and then converts it into the corresponding point cloud. The grasping estimation network generates a collision-free grasping posture for the robotic arm based on point clouds.

5. The collision-free grasping method for an apple-picking robot as described in claim 1, characterized in that, The feature recognition network uses the YOLACT network to detect apple mask images, specifically as follows: Image features are extracted using ResNet-50; Based on the extracted features, a set of k prototype masks for the entire image are predicted by ProtoNet. The Prediction Head predicts the output as detection boxes, categories and mask coefficients respectively. Then, duplicate detection boxes are removed by fast non-maximum suppression. The Protonet and mask coefficient branches are combined using linear weighting, and then the final apple mask image is generated using sigmoid nonlinear activation.

6. A collision-free grasping system for apple picking robots, characterized in that, Includes a range navigation module and a pose generation module: The range navigation module is configured to: detect distant apple trees based on distant images captured by a distant camera, generate a navigation path for the robot, and guide the robot to move to the picking area; The posture generation module is configured to: accurately identify and locate the target apple within the picking range based on the close-up image captured by the hand camera, through a collision-free grasping recognition network, and generate a collision-free grasping posture for the robotic arm to perform obstacle avoidance grasping of the apple. The collision-free grasping posture includes grasping direction and grasping path. The collision-free grasping recognition network generates a grasping direction for an obstacle-avoiding leaf based on point cloud semantic segmentation and performs path planning for the robotic arm to obtain the grasping path. The collision-free grasping and recognition network includes a feature recognition network, a point cloud construction module, and a grasping estimation network; the point cloud semantic segmentation specifically includes: Based on the apple and the surrounding branches and leaves, the point cloud is divided into target point cloud and non-target point cloud; The target point cloud and the non-target point cloud are input into PointNet respectively to obtain target features and non-target features; The two fused features are used as the output of point cloud semantic segmentation. The specific steps of the capture estimation network are as follows: Based on the results of point cloud semantic segmentation, a grasping direction for an obstacle-avoiding leaf is generated using a multilayer perceptron. The path planning of the robotic arm is performed using the RRT sampling algorithm to obtain the grasping path, and finally the collision-free grasping posture of the robotic arm is generated.

7. An electronic device, characterized in that it comprises: Memory is used to store computer-readable instructions in a non-transitory manner. as well as Processor, for executing the computer-readable instructions, When the computer-readable instructions are executed by the processor, they perform the method described in any one of claims 1-5.

8. A storage medium, characterized in that, The computer-readable instructions are stored non-transitory, wherein when the non-transitory computer-readable instructions are executed by a computer, the instructions of the method according to any one of claims 1-5 are executed.

Citation Information

Patent Citations

  • Disordered workpiece three-dimensional visual pose estimation method based on deep learning

    CN114140526A

  • Picking mechanical arm grabbing motion planning method

    CN115139315A

  • Seven-degree-of-freedom fruit picking robot and picking method thereof

    CN116439018A