The invention discloses an Agent-based
video retrieval method, which comprises two major steps of
video storage and retrieval: during
video storage, an original video is subjected to
time sequence segmentation to obtain video clips, a comprehensive description containing a picture description and an audio description is generated through a video description module, and then the overall description of the original video is obtained through summarization of a large
language model; respectively storing the two types of descriptions into a
database and generating a vector index; during retrieval, the Query
processing Agent analyzes user query to clarify a retrieval intention, and the
database retrieval Agent optimizes the query and obtains a
retrieval result based on the vector index. In the video description module, a cue word Agent generates a targeted cue word according to picture description, an audio understanding
large model is guided to extract audio information associated with a picture, and tight combination of the audio information and the picture information is achieved. The problems that in the prior art, sound and picture information association is insufficient, and user intentions are difficult to distinguish are solved, the accuracy of
video retrieval is effectively improved, interaction obstacles between the user and the
system are reduced, and the user experience is improved.