The application provides a visual-
language model prompting method and
system based on a
coupling prompt field, and belongs to the fields of
artificial intelligence and
computer vision. The method comprises the following steps: mapping a
base class task and a new class task to a shared feature space, defining a
coupling prompt field, and enabling the
base class and the new class task to form mutual constraints in the shared feature space; performing scale constraints on the
coupling prompt field through projection layer norm alignment and coding layer layer-by-layer norm alignment; integrating the projection layer norm alignment loss and the coding layer layer-by-layer norm alignment loss, fusing the alignment loss and a task loss of a visual-
language model, forming a model overall training target, and performing training; performing reasoning on the
base class task and the new class task based on the trained model, and outputting a classification prediction result. The application solves the problems of existing end-to-end, decoupling prompt learning, base class-new class isolated optimization, easy falling into
local optimum, norm drift and entanglement collapse, and improves the cross-task generalization and robustness of the model.