The invention discloses a pointing object text positioning method based on iterative semantic visual association. The method comprises the following steps: acquiring a document image containing a pointing object and a
natural language instruction; visual and language features are extracted, a pointing object bounding box is predicted, and a pointing
mask is generated to extract pointing context features; the pointing context features and instruction
semantics are spliced to generate a channel scaling coefficient, and space and
channel modulation is performed on the visual features; using instruction
semantics to generate FiLM parameters to modulate the enhanced visual features, and predicting an initial target bounding box through cross-
modal attention fusion features; and starting an
iterative refinement process, generating a target
mask according to the current prediction frame and calculating a semantic-visual consistency
score, modulating fusion features in combination with geometric difference and
semantic information to predict a better bounding box, and outputting a result with the highest
score after iteration is performed until a termination condition is met. According to the method, fine-grained spatial semantic understanding and high-precision iterative positioning capabilities are enhanced.