The invention discloses a visual language navigation method based on adaptive memory refinement and state action
fine tuning, which comprises the following steps: collecting and marking expert path data, integrating an
open source instruction and a path subset to extract an expert track, and forming an expert path training sample; constructing a
navigation model framework comprising a self-adaptive memory refining module, screening an optimal frame through a multi-dimensional memory optimization equation, and forming and updating a
memory bank; utilizing an expert path training sample to supervise and finely adjust the
navigation model; a fine-grained correction
data set based on a state-action pair is generated and screened by quantifying the deviation degree of an
intelligent agent navigation state and an expert track and combining
trust region constraints; and performing secondary
fine tuning on the fine-tuned model by utilizing the fine-grained correction
data set and combining with the multi-
modal video-language question and answer data, endowing the model with an error correction capability, avoiding disastrous forgetting and realizing intelligent navigation. According to the method, the long-
sight-distance
spatial memory and reasoning capability of the model in a complex scene is enhanced.