The invention discloses a visual language
action model-oriented target injection type
fine tuning method, which comprises the following steps of: firstly, constructing any existing visual language
action model, introducing a condition
image generation model, and generating a target image with consistent
semantics and vision according to an initial observation image and a task target instruction; secondly, target image features are injected into observation input through zero-initialization
convolution, parameters are gradually increased from zero, it is ensured that interference
noise is not introduced in the initial stage of fine adjustment to destroy a pre-training strategy, and in the training process, along with gradual optimization of the parameters, target image feature information is gradually fused into
model representation, and the target image feature information is obtained; therefore, the understanding ability and the execution performance of the task target are improved. According to the method, through a lightweight target image injection mechanism and an efficient fine adjustment process, the performance of the model on various reference tasks can be remarkably improved in few training rounds, and the problem that an existing visual language
action model cannot systematically introduce target
image guidance is effectively solved.