The invention provides an image
subtitle algorithm based on structured semantic extraction and geometric
feature fusion. The image
subtitle algorithm comprises a visual
semantic feature extraction module and a
subtitle generation module, the visual
semantic feature extraction module comprises a regional feature, a network feature, a structured
semantic feature and the subtitle generation module, firstly, the most similar text
sentence is retrieved from an image through a CLIP model, concept
semantics and attribute features are extracted, and after Top-K text features are retrieved through
cosine similarity, multi-level clustering is carried out, and finally, the subtitle generation module is used for generating subtitles; and the isolation problem of
semantic information extraction is fundamentally solved. And secondly, gradually fusing the grid features, the regional features and the structured semantic features by designing a geometric
perception semantic enhancement encoder, and introducing geometric coordinate information in the fusion process, thereby enhancing the expression ability of the structured semantic features. According to the process, the modeling capability of the model for the image space relationship is remarkably enhanced, and the understanding of global
semantics is improved.