The invention relates to the technical field of visual position recognition, and particularly discloses a multi-
modal visual position recognition reordering method and
system based on guidance, and the method comprises the steps: obtaining a query image, and retrieving a plurality of candidate images based on a pre-trained
visual basic model and the query image; constructing a composite multi-
modal prompt object, wherein the composite multi-
modal prompt object comprises an
image pair formed by the query image and the current
candidate image, and an instruction text used for guiding a multi-modal large
language model to perform
visual comparison; outputting a structured similarity judgment result, wherein the result comprises a quantitative similarity
score; and sorting based on the similarity scores corresponding to all the candidate images, and determining the
candidate image with the highest
score as an
optimal matching result. Through combination of guiding type prompt
engineering and structured output, an intermediate
text generation link is avoided fundamentally, and the calculation efficiency is improved while the fidelity of all original visual information is reserved.