The invention belongs to the technical field of
teleoperation and
robot control, and particularly relates to an auxiliary
teleoperation method based on a vision-language-
action model. Based on a few-sample strong generalization auxiliary
teleoperation framework, the key technology of the method is divided into two core stages, namely a data preprocessing stage and a strategy learning and reasoning stage, and the method comprises the following steps: S1, injecting
random noise into a supervised track, and constructing intention disturbance distribution; s2, extracting a
key frame of a supervised track, and constructing an intention representation of geometric
perception; and S3, encoding the processed trajectory as potential embedding, and providing conditions for a vision-language-
action model controller. According to the method, the visual information, the language instruction and the action strategy are fused, rapid
adaptation of the teleoperation task is achieved, good generalization ability is achieved among different operators, cross-operator migration and
robust control are supported, and accurate cross-operator intention recognition and
strategy execution are achieved.