Picking Manipulator RL Control for Collision-Free Fruit Harvesting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Robots used for picking fruits are prone to grazing fruits and branches, leading to damage and picking failures due to the complex morphology of fruit stalks in unstructured environments.

Innovation Solution

An autonomous operation decision-making method for a picking manipulator involves constructing virtual scenes with fruit, branch, and manipulator models, determining target picking points and planes, and using reinforcement learning to optimize picking actions through a reward function, ensuring collision avoidance and desired posture.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the manipulator operates in a highly unstructured environment with complex fruit stalk morphology, then the robot can perform picking tasks, but it is prone to grazing fruits and branches causing damage

Engineering Contradiction:
Improvepicking capability in unstructured environmentVSAvoiddamage to fruits and branches
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The system performs preliminary planning of the picking trajectory and identifies collision-free paths before executing the picking action. The manipulator plans its motion path in advance to avoid grazing fruits and branches, thereby preventing damage before it occurs

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses visual feedback from cameras to detect the real-time position of fruits and branches, and adjusts the manipulator's trajectory dynamically to avoid collisions. The feedback loop enables the robot to adapt to the complex unstructured environment while preventing damage to delicate objects

Inventive Principle:
Principle #23Feedback

2Device complexity

If the manipulator uses traditional control methods for picking, then the system structure is simple, but the picking success rate is low due to grazing incidents

Engineering Contradiction:
Improvecontrol system structureVSAvoidpicking success rate
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The system replaces traditional mechanical control methods with intelligent algorithms including reinforcement learning and trajectory optimization. These computational approaches enable the manipulator to automatically learn optimal picking strategies and avoid collisions without complex mechanical modifications

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system dynamically adjusts motion parameters such as velocity, acceleration, and trajectory points based on real-time environmental perception. By changing these parameters adaptively, the manipulator achieves high picking success rates while maintaining relatively simple system structure

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12564948B2Autonomous operation decision-making method of picking manipulator
Publication Date: 2026.03.03 INTELLIGENT EQUIPMENT RESEARCH CENTER BEIJING ACADEMY OF AGRICULTURE AND FORESTRY SCIENCES
  • US12564948B2 patent drawing
  • US12564948B2 patent drawing
  • US12564948B2 patent drawing

AI summary

The present application relates to the technical field of manipulators and provides an autonomous operation decision-making method of a picking manipulator. The autonomous operation decision-making method of a picking manipulator includes: acquiring sample images of fruits and branches and constructing a plurality of virtual scenes, where each virtual scene includes a picking manipulator model, a fruit model, and a branch model; in the virtual scene, determining azimuths of a target picking point and a target picking plane of an end effector model as parameters and inputting the parameters to a reward function; performing reinforcement learning training on the picking manipulator in the plurality of virtual scenes according to the reward function to determine an optimal picking action function; and controlling the picking manipulator to execute a picking task in an actual environment according to the optimal picking action function.