一种面向开放环境的视觉语言引导机器人抓取方法
By using multimodal data processing and spatial topological adjacency matrix construction, the robotic arm is driven to move occluded objects, iteratively update the point cloud state, and generate the target 6D grasping pose. This solves the problem of target objects being occluded in open environments and improves the grasping success rate and safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUBEI AUTOMATIZATION RES INST
- Filing Date
- 2026-06-12
- Publication Date
- 2026-07-17
AI Technical Summary
Existing visual language-guided grasping methods suffer from low success rates and safety in complex open environments due to object occlusion leading to missing target features, lack of topological and force-compliant interaction planning between multiple objects, and lack of state closed-loop updates after obstacle removal.
By acquiring multimodal data and performing cross-modal feature space alignment, a two-dimensional instance mask and a three-dimensional target point cloud are generated. The command visual grounding confidence is calculated, a spatial topological adjacency matrix is constructed, and the robotic arm is driven to push the occluded object point cloud cluster. The point cloud state is iteratively updated to generate the target 6D grasping pose, and a pre-trained visual language model is used for compliant interactive control.
It improves the robot's success rate in grasping complex and disordered scenarios, ensures the safety and accuracy of the grasping process, and prevents damage to the target object and the environment.
Smart Images

Figure CN122401439A_ABST