The invention discloses an unsupervised semi-
pairing cross-
modal retrieval method and
system based on
deep learning, relates to the field of
artificial intelligence, and is used for solving the problems of
annotation data dependence, asymmetric
semantic association and high-dimensional
storage efficiency. According to the method, a double-
branch visual
encoder and a dynamic prompt text
encoder are combined, dynamic weighting of visual-text features is achieved through gating cross attention, and
modal redundancy interference is restrained. An enhancement strategy is generated through low-frequency semantic guidance, and the long-
tail word coverage rate is increased; a dual-stage quantitative hierarchical index is constructed, coarse-grained clustering and fine-grained product quantitative compression feature storage is adopted, and million-
level data real-time retrieval is supported. A degradation aware increment maintenance mechanism monitors data distribution offset through a KL
divergence threshold, and triggers index reconstruction to maintain long-term update precision. According to the method, limitation of a traditional strong
pairing model is broken through, cross-
modal sensitive content second-level positioning is achieved, asymmetric
semantic alignment is effectively solved, and retrieval efficiency is improved.