The invention relates to the technical field of
robot control, and discloses a master-slave cooperation-based control method, device, equipment and medium, and the method comprises the steps: collecting a spatial
pose, a
contact force and a dual-view image, and constructing a multi-
modal demonstration
data set; executing
time alignment and normalization to form a standardized input sequence; action features, force features and visual features are extracted through the action
modal coding network, the mechanical
modal coding network and the visual coding network; fusing to form a multi-modal
tensor sequence, and inputting the multi-modal
tensor sequence into a Transform decoder to generate a joint
time sequence representation; outputting a
training action prediction vector from the joint
time sequence representation through an action predictor, constructing a supervision
loss function in combination with expert demonstration
annotation to update each network, and obtaining an optimization control model; and
processing the real-time
input control driving execution
mechanism based on the optimization control model. According to the method, the control model is trained and optimized through multi-modal sensing and
time sequence modeling, the joint control instruction is generated in deployment, and stable execution of a complex interaction task is achieved.