The application discloses a pointing object text positioning method based on iterative semantic visual association, which comprises the following steps: obtaining a document image containing a pointing object and a
natural language instruction; extracting visual and language features, predicting a pointing object boundary box and generating a pointing
mask to extract pointing context features; splicing the pointing context features and the instruction
semantics to generate channel scaling coefficients, and modulating the visual features in space and channel; using the instruction
semantics to generate FiLM parameters to modulate the enhanced visual features, and fusing the features through cross-
modal attention to predict an initial target boundary box; starting an
iterative refinement process, generating a target
mask according to the current predicted box and calculating a semantic-visual consistency
score, combining geometric difference and
semantic information to modulate and fuse features to predict a better boundary box, and outputting the result with the highest
score after iteration until the termination condition is met. The application enhances the fine-grained spatial semantic understanding and high-precision iterative positioning capability.