The application provides a dynamic memory enhancement method and
system for a visual-language-
action model and a storage medium, and solves the problems of short-sighted memory, memory
pollution and inability to dynamically adjust the fusion weight of the existing VLA model in a long-distance task. The application comprises: constructing a
working memory and retrieving historical memory from a
perception-
cognition-reward
memory bank; using a gating network with trainable parameters to adaptively weight and fuse the current memory and the historical memory to generate an enhanced
working memory; using a progress evaluator to calculate a task progress difference value based on the enhanced
working memory to generate a single-step dense reward; using the reward as a
merit function to perform online policy optimization on the gating network parameters; and when the
memory bank is full, merging adjacent entries based on the weighted indicators of feature similarity and reward similarity; and finally outputting a motion sequence by a motion expert network. The application can be used for
robot long-
time sequence task learning and
adaptive control in a dynamic environment.