Robot control method and device based on pose condition anchor point attention, and equipment

By using the posture-conditional anchor attention mechanism, the problem of unstable attention in vision-language-action models in complex environments is solved, generating more accurate and stable motion trajectories and improving robot execution efficiency and system efficiency.

CN122425696APending Publication Date: 2026-07-21ZHIPING (SHENZHEN) TECH CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHIPING (SHENZHEN) TECH CO LTD
Filing Date
2026-05-13
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing vision-language-action models lack spatial selectivity in complex environments, making attention easily distracted by task-irrelevant objects or backgrounds, resulting in redundant, unstable, or erroneous motion trajectories.

Method used

A pose-conditional anchor attention mechanism is introduced. By extracting multimodal features and generating pose-conditional anchor attention weights, dense visual features are weighted using these weights and combined with text features and robot state to generate action sequences. A flow matching Transformer model is then used for action planning.

Benefits of technology

Maintain a high success rate in complex environments, generate more accurate and stable motion trajectories, reduce redundant actions, improve execution efficiency, enhance the ability to execute long-term tasks, and reduce system complexity and latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122425696A_ABST
    Figure CN122425696A_ABST
Patent Text Reader

Abstract

The application discloses a robot control method and device based on posture condition anchor point attention and equipment. Including: extracting the multimodal features of the input data, obtaining global visual features, text features and dense visual features; based on the multimodal features and the posture of the robot end effector, generating posture condition anchor point attention weight; using the posture condition anchor point attention weight to weight the dense visual features, obtaining weighted visual features; fusing the text features, weighted visual features and current state of the robot, obtaining multimodal observation representation, generating action sequence based on the multimodal observation representation. The application introduces posture condition anchor point attention, anchors perception in task-related areas, and can still maintain high success rate and generate more accurate and stable action trajectory in complex environments with background, light changes and interference.
Need to check novelty before this filing date? Find Prior Art