Scene text video question answering method and system based on selection and focus mechanism
By using a selection and focusing mechanism to filter keyframes and model the spatiotemporal evolution of scene text, the problems of keyframe omission and high computational cost in video text understanding are solved, thus improving the accuracy and robustness of video question answering.
CN122157282BActive Publication Date: 2026-07-21JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS +1
View PDF 2 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-21
Smart Images

Figure CN122157282B_ABST
Abstract
The application proposes a scene text video question answering method and system based on a selection and focusing mechanism, which first extracts scene text and corresponding video frames in the video; decomposes the original question into coarse-grained sub-questions and fine-grained sub-questions; screens out candidate scene text related to the coarse-grained question according to the coarse-grained sub-questions, and filters to obtain reserved scene text; screens out key frames according to the original question and the reserved scene text; captures the spatial dynamics and time existence of the scene text in the cross frame by using the attention mechanism to obtain local features; inputs the key frames and the local features into the visual language model together with the original question to generate the final output result. The application adopts the "coarse-fine" granularity question guiding selection strategy, discards the traditional uniform frame sampling method, and automatically filters the redundant frames and locks the key frames containing the answers according to the question semantics, so that the question related frames are not missed.
Need to check novelty before this filing date? Find Prior Art