The invention discloses a video cross-view alignment retrieval method based on a high-dynamic scene, belongs to the field of
image processing, and particularly relates to a cross-view alignment retrieval method for solving the problem of two-way action migration from
third person demonstration to first person execution in a real scene. The method comprises the following steps: sampling an input video through a scene dynamic enhancement frame extraction mechanism to extract double-density effective frames, comparing and screening inter-frame visual feature similarity, abandoning redundant similar frames, retaining key frames with remarkable dynamic change, and accurately capturing high-dynamic action features; and learning uniform significant feature representation by combining a dual-
encoder architecture and a parameter efficient
adaptation technology, and solving the core problem that common features are difficult to extract in cross-view retrieval through complementary strategies such as multi-positive-example aggregation, bidirectional sorting consistency constraint and intra-
modal cohesion, so that the method is suitable for the cross-view retrieval of the multi-view retrieval. Visual representation, language anchoring and cross-view alignment performance are synchronously optimized, finally fine-grained efficient retrieval is achieved in a real scene, and lightweight deployment cross-view bidirectional action migration and retrieval are supported.