The invention discloses a text and image bidirectional alignment method and
system based on multi-hop parallel reasoning, and the method comprises the steps: analyzing and marking a text, and obtaining a multi-
granularity text feature group;
image generation is processed, grids and target output are combined into a multi-scale visual feature group, an alignment index table is established, a sub-target chain is constructed based on visual features, the position and the dependency relationship are recorded, and explicit constraints are injected into five elements; recalling the candidate area, generating an
anchor point through
verification and fusion, writing the
anchor point into a cross-
modal index, locally completing three-layer decoupling scoring based on
anchor point geometry and the cross-
modal index, and outputting a three-state evidence in combination with a threshold value; the method comprises the steps of generating an evidence set through geometric and
semantic consistency calibration, generating cross-hop evidence according to a rule, weighting and
pooling the evidence set, generating a shared vector and traceable meta-information, inputting a task head output result and generating an evidence
list and an auditing path, and according to the method, local alignment, constraint
verification and evidence accumulation are completed hop by hop. And high-precision, traceable and interpretable cross-
modal correspondence is realized.