基于视觉语言大模型的机械臂语音交互抓取方法及系统
By using a robotic arm voice interaction grasping method based on a large visual language model, the problem of difficulty in matching natural language target descriptions with visual detection labels is solved, enabling stable object handling in multi-target desktop scenarios and improving the system's lightweight nature and ease of deployment.
CN122401432APending Publication Date: 2026-07-17QINGDAO IND SOFTWARE RES INST QINGDAO BRANCH OF SOFTWARE RES INST CAS
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- QINGDAO IND SOFTWARE RES INST QINGDAO BRANCH OF SOFTWARE RES INST CAS
- Filing Date
- 2026-06-10
- Publication Date
- 2026-07-17
Smart Images

Figure CN122401432A_ABST
Abstract
本发明涉及机器人控制、桌面视觉引导抓取与自然语言交互技术领域,具体公开了一种基于视觉语言大模型的机械臂语音交互抓取方法及系统。该方法包括:获取搬运指令;对搬运指令进行解析,输出结构化任务信息;利用语言模型或视觉语言模型对起始物体描述和终止物体描述进行规范化处理,转换为用于视觉检测识别的标签;将桌面图像与标签输入目标检测模型,获取候选边界框及置信度,提取中心点像素坐标;求取工作平面坐标换算矩阵,并将起始物体中心点像素坐标和终止物体中心点像素坐标换算为工作平面参考坐标;补全预设动作流程中的可变参数,根据预设动作流程输出搬运动作执行指令至机械臂。
Need to check novelty before this filing date? Find Prior Art