The invention provides double-
granularity alignment efficient partial correlation
video retrieval based on implicit fragment modeling and semantic
decomposition, and aims to solve the problems of
information redundancy and low efficiency in an existing video modeling method, the problem of
granularity mismatching between
sentence representation and video frame features and the problem that alignment between a text and a video is not refined enough. According to the method, the expression and modeling of video data are optimized by introducing a structure combining a
Gaussian mixture model and Transform, and multi-scale local details with short time span are adaptively integrated by introducing a window attention mechanism and a cross attention mechanism, so that finer features are obtained, and cross-
modal similarity
score calculation of texts and videos is facilitated. The method comprises the following steps: data preprocessing and frame segmentation: preprocessing and segmenting an input video into frames, and preparing for
feature extraction after each frame of image is subjected to
standardization processing; implicit modeling is carried out on the fragment-level features, uniformly sampled frame-level visual features are input into a
Gaussian mixture modeling module, adjacent frame focusing modeling is carried out through a multi-scale
Gaussian attention mechanism, and fragment-level representation of different receptive fields is implicitly formed. Enhancing local details of the frame-level features, performing
convolution operation on each frame by adopting windows with different scales, and calculating a feature relationship in each local window; local features of all scales are fused through a cross attention mechanism, semantic expression of frame features is enhanced, and it is ensured that each video frame can capture important local details.