Text-guided person re-identification model construction method and device based on confidence perception, and retrieval method
By employing a confidence-based perception mechanism and cross-modal mask image modeling, the problems of noise interference in image-text matching and insufficient cross-modal alignment capability are solved, thereby improving the accuracy and robustness of text-guided person re-identification.
CN122369064APending Publication Date: 2026-07-10XINJIANG UNIVERSITY
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-07-10
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
The present application relates to the field of computer vision and multi-modal retrieval technology, and particularly relates to a text-guided person re-identification model construction method and retrieval method and device based on confidence perception. The model construction method comprises: obtaining a text-image pair training sample composed of a person image and a corresponding text description; constructing a text-guided person re-identification baseline model comprising an image encoder, a text encoder and a cross-modal alignment module; fitting a Gaussian mixture model according to the single-sample cross-modal alignment loss of each text-image pair sample in a batch to obtain a sample confidence weight, and weighting the cross-modal alignment loss based on the confidence weight; performing image block division and random masking on the input image, and combining the text features to perform cross-modal reconstruction on the masked image blocks; jointly optimizing the model according to the weighted cross-modal alignment loss, the reconstruction loss and the baseline loss to obtain a trained text-guided person re-identification model. The retrieval method comprises: using the trained text-guided person re-identification model to respectively encode the query text and the candidate person image, calculating the similarity and outputting the retrieval result. The present application can reduce the interference of noisy text-image pair samples on model training, enhance the fine-grained alignment capability between text semantics and image local regions, and thus improve the accuracy and robustness of text-guided person re-identification.
Need to check novelty before this filing date? Find Prior Art