The invention belongs to the field of
computer vision and cross-
modal retrieval, and particularly relates to an image-text
pedestrian retrieval method based on image block replacement and cross-
modal identity alignment, which comprises the following steps: acquiring a public
data set of image and text description, and constructing a
pedestrian re-recognition model PRCIA; inputting the
data set into a PRCIA model for training and
verification, and performing iterative updating on a training weight file through forward and
backward propagation to obtain a trained PRCIA model; constructing a reasoning stage model, reserving double encoders and fusing global features in a reasoning stage, and ensuring the calculation efficiency; and inputting the
test set into the reasoning stage model to obtain a detection result, thereby realizing text-based
pedestrian detection. According to the method, fine-grained association between the image blocks and the text phrases is established through the PR module, cross-
modal identity feature expression is enhanced through the CIA module, the accuracy and robustness of text-to-image pedestrian retrieval are remarkably improved, and the requirements of actual scenes are more easily met.