The invention discloses an unsupervised cross-
modal retrieval method based on robust
consensus learning, and the method comprises the following steps: extracting image and text features of input data through a pre-training
backbone network, and obtaining multi-level feature representation through a three-layer
projector; generating a prototype as a pseudo
label based on top layer feature clustering, calculating a
sample weight by using contour coefficients of three levels, performing optimization through layered
consensus prototype comparison loss, and finally obtaining a high-reliability pseudo
label; the
noise tolerance triple alignment loss and the
modal consistency loss are jointly optimized, and the robustness of the model to
noise is enhanced; and cross-
modal retrieval is realized in the optimized shared
semantic space. According to the method, a two-stage training strategy is adopted, the pseudo-labels are generated through top-layer clustering, the
sample weight is optimized through multi-layer
semantic information, robust learning is carried out based on the reliable pseudo-labels, the problems of pseudo-
label noise and semantic
granularity mismatch in an unsupervised environment are solved, and the accuracy and robustness of cross-modal retrieval are remarkably improved.