The invention relates to the technical field of
computer vision, and discloses a figure interaction detection method based on spatial fine-grained context interaction
feature fusion, which comprises the following steps of: firstly, performing target detection on an input image to obtain a target detection
result set and a person-object
pairing feature; performing gridding projection on the image to obtain image global features, and inputting the image global features into a spatial fine-grained
feature learning module to obtain spatial fine-grained features; inputting the spatial fine-grained features and the human-object
pairing features into a spatial context interaction
feature fusion module to obtain spatial context interaction features, and inputting the spatial context interaction features and the spatial fine-grained features into a visual
encoder to obtain enhanced human-object
pairing features; obtaining an interaction category
score according to the feature and a text embedding feature of a character interaction category; and iteratively optimizing the character interaction detection model until convergence. According to the method, the description of local details of a
visual space is enhanced, the context relationship between a person and an object is enhanced, and the accuracy of person interaction detection is improved.