A multi-task prompt decision
transformer construction method and device, equipment and storage medium, relate to the field of
robot control
simulation environment, including constructing MPDT, MPDT includes multiple
decision maker layers and triple prompt module constructed based on reward, state and action;Control target
decision maker layer and triple prompt module are in
active state, MPDT is trained based on offline training sample, and multi-task prompt fine-tuning is carried out based on triple prompt, specific task prompt and cross-task prompt are generated;After training is completed, the combination of cross-task prompt and
test sample is used as the input of MPDT for testing, the hidden state test mean and hidden state test variance corresponding to the
test sample are calculated, and the mean and variance of the offline training sample are aligned and calculated;Based on the alignment result, the cross-task prompt is updated, the target multi-task prompt decision
transformer is generated, and the generalization ability of the
transformer in the zero sample and multi-task scene is improved.