Text-based image pedestrian retrieval model construction method and image pedestrian retrieval method
By introducing the accurate image feature enhancement module and the text interference suppression module in the image pedestrian retrieval model, the challenges in cross-modal alignment are solved, the retrieval performance and noise resistance are improved, and more accurate feature alignment is achieved.
Patent Information
- Application Number
- CN202510210952.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-25
AI Technical Summary
The prior art faces the challenge of cross-modal alignment in text-based image pedestrian retrieval, especially due to the influence of image background noise and mismatched words in the text, resulting in inaccurate fine-grained image features and inalignment of text image features.
A framework containing an accurate image feature enhancement module and a text interference suppression module is proposed. By extracting foreground information and removing background noise, the accuracy of local image features is improved, and text interference is reduced by identifying and masking inconsistent words, achieving more accurate text feature extraction.
The search performance and noise resistance of the image pedestrian search model are improved, more accurate image and text feature alignment is achieved, and the generalization ability of the model is enhanced.
Smart Images

Figure CN120216713A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field related to image-text cross-modal retrieval, and more specifically, relates to a method for constructing a text-based image pedestrian retrieval model and an image pedestrian retrieval method. Background Art
[0002] Text-based pedestrian retrieval is a sub-task of image-text retrieval and person re-identification, aiming to retrieve images of specific individuals from a large database according to text descriptions. Its importance is obvious when only text information is available, making it indispensable for applications such as video surveillance and assisted security. However, the inherent differences between the visual modality and the text modality pose significant challenges to this task.
[0003] Current methods usually adopt global feature and local feature alignment to bridge the modality gap. Global-level methods strive to learn global features shared by the two modalities in the embedding space, but they may ignore the correlation relationships related to fine-grained features. Local-level methods focus on extracting local features of a specific modality and achieving precise alignment. However, these methods face significant challenges, mainly related to the noise in the text and the limitation of small-scale training data. The noise in the text and the image may lead to a mismatch between the text query and the visual representation of the individual. In addition, small-scale training data hinders the model's ability to learn robust and precise features.
[0004] Previous global-level methods often overlooked the refinement of extracting image features. This oversight may cause problems because images usually contain regions unrelated to the paired text, which may introduce noise and interfere with the alignment of image and text features. Although some local-level methods have made some progress in feature refinement, they usually divide text and visual features into equally sized blocks and calculate the similarity between the blocks in sequence. However, this block-to-block comparison lacks strict semantic correspondence or precise word-to-region alignment, which may lead to unintentional misalignment. In addition, in addition to the noise in the image, previous studies ignored a key fact that some words in the text description may not exactly match the paired image. The mismatch between the local text description and its corresponding image region will cause differences. When using local feature alignment techniques, this difference may be amplified. Tan et al. proposed a noise-aware masking method to mask the mismatched words in the text, but their main focus was on the text generated by large language models, and their method ignored the difference between individual tokens and the whole text, resulting in some differences and increasing the difficulty of achieving accurate alignment. Summary of the Invention
[0005] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides a method for constructing a text-based image pedestrian retrieval model and an image pedestrian retrieval method, aiming to retrieve images of specific individuals according to text descriptions, and solve the challenges in cross-modal alignment through an accurate feature enhancement module and a text interference reduction module, improving the retrieval performance and anti-noise ability.
[0006] To achieve the above object, according to one aspect of the present invention, there is provided a method for constructing a text-based image pedestrian retrieval model, including:
[0007] Construct an image pedestrian retrieval network, including an image feature enhancement module, a text noise reduction module, and a pre-trained bimodal feature extraction backbone network; wherein, the bimodal feature extraction backbone network is used to encode and obtain the global image feature and the local image feature vector of the input image, and the image feature enhancement module is used to process the input image to obtain a binarized block image vector, perform a vector dot product on the block image vector and the local image feature vector to obtain a local image feature vector with background information removed, take the average of all features in the vector and map it to obtain the pedestrian enhancement feature of the input image;
[0008] The text noise reduction module is used to obtain the local text feature of the input text corresponding to the input image based on the bimodal feature extraction backbone network, judge the degree of deviation of the similarity between each local text feature and each feature in the local image feature vector from the similarity between the average value of all local text features and each feature in the local image feature vector, mask the local text corresponding to the local text feature with a quantization value of the deviation degree greater than the threshold in the input text, and finally obtain a new input text;
[0009] Construct an image-text pair training sample set, and based on the global image feature, pedestrian enhancement feature of each sample, and the global text feature encoded by the bimodal feature extraction backbone network based on the new input text, iteratively adjust the parameters of the bimodal feature extraction backbone network and the operation parameters of the mapping through the image-text feature similarity loss to realize the construction of the image pedestrian retrieval model.
[0010] Further, the implementation manner of the image feature enhancement module when processing to obtain the binarized block image vector is: perform human parsing processing on the pedestrians in the input image to obtain a pedestrian contour mask map; perform binarization processing and block division on the mask map to obtain a binarized block image vector.
[0011] Further, the implementation manner of the text noise reduction module when determining the quantization value of the deviation degree is:
[0012] Calculate the similarity between each local text feature and each feature in the local image feature vector with background information removed respectively to obtain the first similarity vector corresponding to the local text feature; average all local text features, and calculate the similarity between the average local text feature and each feature in the local image feature vector with background information removed respectively to obtain the second similarity vector corresponding to the average local text feature; subtract each first similarity vector from the second similarity vector, and the resulting vector is the quantization value of the deviation degree corresponding to the local text feature.
[0013] Further, the implementation manner of the text denoising module for the local text mask based on the quantization value of the deviation degree is as follows:
[0014] Average all elements in the subtracted vector, add the result to the preset balance parameter value as the probability of local text error corresponding to the first similarity vector, and mask the local text corresponding to the probability that meets the threshold in the input text to obtain a new input text.
[0015] Further, the image-text feature similarity loss includes ID loss and SDM loss.
[0016] Further, the method further includes: in each iteration, store all global image features, pedestrian enhancement features corresponding to the current batch, and global text features encoded by the dual-modal feature extraction backbone network based on the new input text into the image feature memory bank and the text feature memory bank respectively, and combine the momentum update model based on exponential moving average to stabilize the feature extraction process of the dual-modal feature extraction backbone network;
[0017] Then the image-text feature similarity loss further includes MITC loss, and its loss function is expressed as:
[0018]
[0019] In the formula, represents the MITC loss from the global image feature of the image to the global text feature encoded based on the new input text, represents the MITC loss from the global text feature encoded based on the new input text to the global image feature of the image, represents the MITC loss from the pedestrian enhancement feature of the image to the global text feature encoded based on the new input text, The MITC loss from the global text feature encoded based on the new input text to the pedestrian enhancement feature of the image;
[0020] Among them, In the formula, τ is a temperature hyperparameter, B$B$ is the current batch size during training, and $N$ is the total number of image-text pairs; denotes the true matching label of the $j$-th image-text pair in the $i$-th batch. If it means that the current image-text pair is a matching pair from the same pedestrian identity. If it means that the current image-text pair does not match; is the global image feature of all images in the $i$-th batch, is the global text feature encoded based on the new input text corresponding to the $j$-th image-text pair in the text memory bank $T$, is the global text feature encoded based on the new input text corresponding to the $n$-th image-text pair in the text memory bank; denotes and the similarity between; denotes and the similarity between;
[0021] In the formula, replace with the corresponding image feature, replace and with the corresponding image features, which is the expression of; In the formula denotes the pedestrian enhanced feature of all images in the $i$-th batch, which is the expression of; In the formula, replace with the corresponding image feature, and with the corresponding image features, which is the expression of.
[0022] Furthermore, the pre-trained bimodal feature extraction backbone network is composed of the image encoder and the text encoder in the CLIP large model.
[0023] According to another aspect of the present invention, a text-based image pedestrian retrieval method is provided, including:
[0024] Obtain the pedestrian image set and the target text of the image to be searched, and input them into the bimodal feature extraction backbone network in the image pedestrian retrieval model respectively, and determine the pedestrian image corresponding to the target text from the image set to implement text-based image pedestrian retrieval;
[0025] wherein, the image pedestrian retrieval model is constructed by using the construction method described above.
[0026] According to another aspect of the present invention, there is provided an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method described above are implemented.
[0027] According to another aspect of the present invention, there is provided a computer-readable storage medium, which includes a stored computer program. Wherein, when the computer program is run by a processor, the device where the storage medium is located is controlled to execute the steps of the method described above.
[0028] Generally speaking, compared with the prior art through the above technical solutions conceived by the present invention, the technical solutions provided by the present invention mainly have the following beneficial effects:
[0029] 1. The present invention considers that cross-modal alignment is mainly interfered by the image background and the influence of mismatched words in the text, resulting in inaccurate fine-grained image features and misalignment of text and picture features. To solve this challenging task, a novel precise image feature enhancement and text interference suppression framework is proposed. This framework contains two key components: a precise image feature enhancement module, which improves the accuracy of local image features by extracting foreground information and removing background noise; a text interference suppression module (i.e., text noise reduction module) effectively reduces the impact of irrelevant or misleading text interference by identifying and masking inconsistent words with high probability, achieving the effect of text noise reduction. Experimental results show that this method has advantages such as strong generalization ability and good retrieval performance. Therefore, the present invention can retrieve images of specific individuals according to text descriptions, and through the precise feature enhancement module and text interference reduction module, solve the challenges in cross-modal alignment, and improve the retrieval performance and anti-noise ability.
[0030] 2. The present invention further proposes that the dual-modal feature extraction backbone network adopts the image and text encoders in the pre-trained CLIP large model as the basic visual and text feature extractors in the entire image pedestrian retrieval model. This method effectively transfers the knowledge of CLIP to text-based person retrieval, and extracts relatively accurate image features and text features. Description of the Drawings
[0031] Figure 1 It is a schematic diagram of a method for constructing a text-based image pedestrian retrieval model provided by an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the problems existing in text-based image pedestrian retrieval provided by an embodiment of the present invention;
[0033] Figure 3 It is a schematic diagram of another method for constructing a text-based image pedestrian retrieval model provided by an embodiment of the present invention. Detailed implementation manners
[0034] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0035] Embodiment 1
[0036] A method for constructing an image pedestrian retrieval model based on text, as Figure 1 shown, includes:
[0037] Construct an image pedestrian retrieval network, including an image feature enhancement module, a text denoising module, and a pre-trained bimodal feature extraction backbone network;
[0038] The bimodal feature extraction backbone network is used to encode and obtain the global image feature and the local image feature vector of the input image. The image feature enhancement module is used to process the input image to obtain a binarized block image vector, perform a dot product of the block image vector and the local image feature vector to obtain a local image feature vector (also called the main body embedding feature) with background information removed, take the average of all features in the vector and map it to obtain the pedestrian enhancement feature of the input image. The text denoising module is used to obtain the local text feature of the input text corresponding to the input image based on the bimodal feature extraction backbone network, judge the degree of deviation of the similarity between each local text feature and each feature in the local image feature vector from the similarity between the average value of all local text features and each feature in the local image feature vector, mask the local text corresponding to the local text feature with a quantization value of the deviation degree greater than the threshold in the input text, and finally obtain a new input text;
[0039] Construct an image-text pair training sample set;
[0040] Based on the global image feature, pedestrian enhancement feature of each sample, and the global text feature encoded by the bimodal feature extraction backbone network based on the new input text, through the image-text feature similarity loss, iteratively adjust the parameters of the bimodal feature extraction backbone network and the operation parameters of the mapping to realize the construction of the image pedestrian retrieval model.
[0041] The image pedestrian retrieval model constructed by the method of this embodiment includes a bimodal feature extraction backbone network (including an image encoder and a text encoder) and two implicit alignment strategies (i.e., the fine-grained image feature enhancement module and the text interference suppression module). The bimodal feature extraction backbone network is a unified vision-text network, which helps to obtain a joint embedding space in which visual and text features can be harmoniously aligned. The fine-grained image feature enhancement module aims to improve the accuracy of local image features by extracting foreground information and removing background information, and further obtain enhanced features from the image by carefully integrating the global features from the image encoder, which helps to establish a closer association with the text description. The text denoising module aims to identify and mask inconsistent words with high probability based on the local image feature vector to reduce the impact of irrelevant or misleading text information.
[0042] As Figure 2 shown, in the text-pedestrian retrieval method, the image usually contains regions irrelevant to the paired text, which may introduce noise and interfere with the alignment of image and text features. Figure 2 (a) on the left in Figure 2 vividly shows this problem, showing that the regions marked in red fail to align with the text description and may confuse the feature alignment process. The fine-grained image feature enhancement module uses a self-correcting human parsing algorithm to selectively extract foreground information and filter out background noise, as shown in Figure 2 (a) on the right in
[0043] to reduce semantic errors and improve the accuracy of feature localization. In addition, in addition to the noise in the image, some words in the text description may not exactly match the paired image. For example, in Figure 2 (b), there is a difference between the content described by the red-colored word in the right text and the actual clothing features of the pedestrian shown in the left figure.
[0043] The task of text-based person image retrieval aims to retrieve images of specific individuals according to text descriptions. However, this task faces the challenge of cross-modal alignment. This embodiment considers that it mainly stems from the interference of the image background and the influence of unmatched words in the text, resulting in inaccurate fine-grained image features and misalignment of text and image features. To solve this challenging task, a novel framework for precise image feature enhancement and text interference suppression is proposed. This framework contains two key components: the precise image feature enhancement module, which improves the accuracy of local image features by extracting foreground information and removing background noise; the text interference suppression module (i.e., the text denoising module) effectively reduces the impact of irrelevant or misleading text interference by identifying and masking inconsistent words with high probability, thereby enhancing the ability of the retrieval model to focus on more accurate text features during the learning process and achieving the effect of text denoising. Experimental results show that this method has advantages such as strong generalization ability and good retrieval performance.
[0044] Generally speaking, this embodiment provides a fine-grained feature refinement and text noise suppression framework to improve the retrieval performance of a text-based person retrieval model.
[0045] Preferably, the cross-modal feature extraction backbone network adopts the image encoder and text encoder in the pre-trained CLIP large model and serves as the basic visual and text feature extractor in the entire image pedestrian retrieval model. This method effectively transfers the knowledge of CLIP to text-based person retrieval and extracts relatively accurate image features and text features.
[0046] Specifically, for image feature extraction. Given an image \(I\in\mathbb{R}\) W×H×C , first, it is divided into a series of non-overlapping image patches, with a total of \(M\) patches. Subsequently, the image patches are mapped into one-dimensional tokens, and a learnable [CLS] token is added at the beginning of the token sequence. Finally, these linear tokens are fed into the image encoder to extract a set of image features, denoted as \(F\) v =\(\{f\) cls v , \(f_1\) v , \(\cdots\), \(f\) M v \}\). Here, \(f\) cls v represents the global-level image feature, and \(F\) vl =\(\{f_1\) v , \(\cdots\), \(f\) M v \}\) represents the local-level image feature.
[0047] For text feature extraction. Given a text \(T\) o , the text encoder of CLIP is used to extract text tokens. Subsequently, the Byte Pair Encoding (BPE) method is used to tokenize the text tokens, and [SOS] (Start of Sequence) and [EOS] (End of Sequence) tokens are added at the beginning and end of the text token sequence respectively. This tokenized sequence is then processed through the text Transformer block to generate a set of text features, denoted as \(F\) t =\(\{f\) sos t , \(f_1\) t , \(\cdots\), \(f\) eos t \}\), where \(f\) eos t represents the global-level text feature, and \(F\) tl =\(\{f_1\) t , \(\cdots\), \(f\) N t \}\) represents the local-level text feature.
[0048] As a preferred implementation, the implementation manner of the image feature enhancement module when processing and obtaining the binary segmented image vector is as follows: perform human parsing processing on the pedestrians in the input image to obtain a pedestrian contour mask image; perform binary processing and segmentation on the mask image to obtain a binary segmented image vector.
[0049] As Figure 3 shown, in order to obtain more accurate image features and filter out background noise, a human parsing technique is used to generate a mask I of the main body parts for each image m . The main body parts include the head, limbs, and torso. For a single image, obtain its local image feature F vl and the main body part image mask I m .
[0050] The main body part image mask is binary processed, setting the pixel values of the body area to 255 and the pixel values of the background area to 0. Then, a summation operation is performed on all three color channels, setting any non-zero value to 1, thereby obtaining an intermediate feature representation I v . This process effectively captures human features while discarding irrelevant background information. Subsequently, we use the nearest neighbor interpolation algorithm to resize I v and then perform an element-wise multiplication operation with F vl to obtain the main body embedding vector F′ vl . After an averaging operation, F′ vl is input into the projection layer to generate the main body feature fmain of the image. This series of operations can robustly extract the main human features and promote the alignment of image and text features.
[0051] The fine-grained image feature enhancement module uses a self-correcting human parsing algorithm to selectively extract foreground information and filter out background noise, thereby reducing semantic errors and improving the accuracy of feature localization. By carefully integrating the global features from the image encoder, enhanced features are further obtained from the image, which helps to establish a closer association with the text description. The fine-grained image feature enhancement module can proficiently analyze the key components in the image, thereby refining the feature extraction stage and improving the retrieval accuracy.
[0052] As a preferred implementation, the implementation manner of the text noise reduction module when determining the quantization value of the deviation degree is as follows:
[0053] Calculate the similarity between each local text feature and each feature in the local image feature vector with background information removed, to obtain the first similarity vector corresponding to this local text feature; average all local text features, calculate the similarity between the average local text feature and each feature in the local image feature vector with background information removed, to obtain the second similarity vector corresponding to the average local text feature; subtract each first similarity vector from the second similarity vector, and the resulting vector is the quantization value of the deviation degree corresponding to this local text feature.
[0054] Furthermore, as a further preferred implementation, the implementation manner of the text denoising module for the local text mask based on the quantization value of the deviation degree is as follows:
[0055] Average all elements in the subtracted vector, add it to the preset balance parameter value, as the probability of local text error corresponding to this first similarity vector, mask the local text corresponding to the probability that meets the threshold in the input text, to obtain a new input text.
[0056] The method of this embodiment introduces a balance parameter value, so that the average value is within a set numerical range, forming an operation standard, which is convenient for comparison with the set threshold.
[0057] Specifically, as Figure 3 shown, during training, one input image corresponds to one input text for training. For an image-text pair, obtain the local-level text feature F tl through the text encoder. Calculate the difference between the average of all image tokens corresponding to each text token (local text) and the average of all tokens corresponding to image tokens, to determine whether a certain text token is noise. Also obtain the local image feature vector F' vl from the fine-grained image feature enhancement module. Calculate the similarity between the average text token feature and each feature in F' vl , which can be expressed as: Where represents the average value of the degree of consistency between each text token and the image feature.
[0058] Then, for each text token feature tl in F calculate and the similarity S vl between F' i : Where represents the cosine similarity between the i-th text token embedding and the j-th image token embedding. Next, calculate the deviation between S i and : Subsequently, calculate S'i The mean value, and define a noise level P i is defined as: where p is a balance parameter. After optimization, P i represents the probability that the text token is masked. Use the text noise suppression module to convert the initial text T o into a new input text, called the refined text T r , where words with low noise or correlation are selectively masked. This masked text T r is then input into the text encoder to obtain more accurate global text features and participate in model training. The text noise suppression module improves the robustness and effectiveness of the model by focusing on valid text tokens and reducing the impact of irrelevant tokens.
[0059] As a preferred implementation, the image-text feature similarity loss includes the ID loss and the SDM loss.
[0060] The ID loss L id is defined by the formula:
[0061]
[0062] where N represents the number of images per batch, y i represents the pedestrian ID corresponding to the global feature f i , C represents the number of pedestrian categories included, both represent parameters in the training.
[0063] For the SDM loss, it is introduced as follows:
[0064] Given a mini-batch of data containing N image-text pairs, set the global feature representation of each image as Construct a set of image-text representation pairs where y i,j is a true matching label. If y i,j = 1, it means that is a matching pair from the same identity; if y i,j = 0, it means that this pair is unmatched. Let represent the dot product (i.e., cosine similarity) between u and v after L2 normalization. Then, the probability of the matching pair can be simply calculated by the following softmax function:
[0065]
[0066] In the formula, τ is a temperature hyperparameter used to control the peak of the probability distribution. The matching probability p i,j can be regarded as in the mini-batch The cosine similarity score between and accounts for the proportion of the total cosine similarity score between and
[0067]
[0068] where ∈ is a small value to avoid the problem of a zero denominator, represents the true matching probability.
[0069] Similarly, the SDM loss from the global text features encoded from the new input text to the global image features of the image, L t2i , can be formulated by swapping f v and f t in equations (1) and (2). The SDM loss from the pedestrian-enhanced features of the image to the global text features encoded from the new input text from the global text features encoded from the new input text to the pedestrian-enhanced features of the image is similar.
[0070] The total SDM loss is calculated as follows:
[0071] As a preferred implementation, the method further includes: in each iteration, storing the corresponding global image features, pedestrian-enhanced features, and global text features encoded by the bimodal feature extraction backbone network based on the new input text into the image feature memory bank and the text feature memory bank, and combining with the momentum update model based on the exponential moving average to stabilize the feature extraction process of the bimodal feature extraction backbone network;
[0072] Then the image-text feature similarity loss further includes the MITC loss, and its loss function is expressed as:
[0073]
[0074] where represents the MITC loss from the global image features of the image to the global text features encoded from the new input text, represents the MITC loss from the global text features encoded from the new input text to the global image features of the image, represents the MITC loss from the pedestrian-enhanced features of the image to the global text features encoded from the new input text, the MITC loss from the global text features encoded from the new input text to the pedestrian-enhanced features of the image;
[0075] Among them, where τ is a temperature hyperparameter, B is the current batch size during training, and N is the total number of image-text pairs; represents the true matching label of the j-th image-text pair in the i-th batch. If then it means that the current image-text pair is a matching pair from the same pedestrian identity. If then it means that the current image-text pair does not match; is the global image feature of all images in the i-th batch, is the global text feature encoded based on the new input text corresponding to the j-th image-text pair in the text memory bank T, is the global text feature encoded based on the new input text corresponding to the n-th image-text pair in the text memory bank; represents and the similarity between; represents and the similarity between;
[0076] In the formula, replace with the corresponding image feature, replace and with the corresponding image features, and it is the expression of ; In the formula, represents the pedestrian enhancement feature of all images in the i-th batch, which is the expression of ; In the formula, replace with the corresponding image feature, and with the corresponding image features, and it is the expression of .
[0077] Considering that the limitation of the small batch size may not reflect all details comprehensively, an image memory bank and a text memory bank are proposed. By storing all the image and text features in the current dataset, these memory banks can provide a more detailed and comprehensive understanding of the main part of the image details, thereby enhancing the model's ability to learn and utilize from a large amount of relational information.
[0078] Specifically, using the current volatile model to construct the memory bank may introduce instability. To address this issue, a momentum update framework is introduced. Specifically, the exponential moving average (EMA) is adopted to stabilize and optimize the main part of the image and text encoders, ensuring consistent and reliable memory updates throughout the training process. EMA can be expressed as:
[0079]
[0080] where only the online model and are updated through backpropagation, and m ∈ [0, 1] is the momentum coefficient. The encoder using momentum update is used to generate the features stored in the memory bank, resulting in a more stable representation.
[0081] Subsequently, a comprehensive measure of bimodal similarity is introduced:
[0082]
[0083] where B is the current batch size, is the enhanced feature of the image pedestrian, is the j-th feature in the text memory bank. Correspondingly, the loss from text to image can also be obtained. Similarly, the momentum update model is also applied to the global image features and text features. Therefore, the contrastive loss can be expressed as:
[0084]
[0085] main2t represents the loss from the enhanced feature of the image pedestrian to the text feature, t2main represents the loss from the text feature to the enhanced feature of the image pedestrian, i2t represents the loss from the global image feature to the text feature, and t2i represents the loss from the text feature to the global image feature. 2 represents the meaning of "to" in English, and the two have the same pronunciation.
[0086] Due to the limited number of samples in the batch, the relationship between positive and negative samples becomes simple. Therefore, it is urgent to expand the sampling range to find more difficult negative samples. This preferred embodiment constructs an image memory bank and a text memory bank to expand the range of each batch. This expansion enables the training to use more diverse and challenging positive and negative samples, enhancing the robustness of the model representation.
[0087] Considering the limitations of small batch sizes, image and text memory banks are proposed to store all image and text features in the current dataset. This helps to understand the main body image details in more detail and enhance the learning effect of the model. Specifically, to mitigate the impact of the unstable current model, a momentum update framework using Exponential Moving Average (EMA) is introduced to stabilize and optimize the image and text encoders, ensuring consistent and reliable memory updates during training.
[0088] Embodiment 2
[0089] A text-based image pedestrian retrieval method, comprising:
[0090] Obtain a pedestrian image set and a target text of the image to be searched, and input them into the dual-modal feature extraction backbone network in the image pedestrian retrieval model respectively, and determine the pedestrian image corresponding to the target text from the image set, so as to realize text-based image pedestrian retrieval;
[0091] Wherein, the image pedestrian retrieval model is constructed by using the construction method described above.
[0092] In order to prove that the text-pedestrian retrieval method combining feature enhancement and text denoising has advantages in both performance and adaptability, the present invention is verified and analyzed through the following experiments:
[0093] A series of experiments were carried out on three text-pedestrian retrieval datasets (including CUHK-PEDES, ICFG-PEDES, and RSTPReid). The mean Average Precision and the probability that the top (highest confidence) image in the search results has a correct result (rank-n) were used as the accuracy metrics for pedestrian re-identification. For image processing, the CLIP-VIT-B / 16 model was used; for text input, the CLIP text encoder was used. In the training stage, a variety of image enhancement techniques were applied, including random cropping, flipping, and erasing, to enhance data diversity. Similarly, the text enhancement strategy involved randomly deleting, masking, and replacing words to enhance text input. The images were resized to a resolution of 128*384 pixels, and the length limit of the text input was 77 characters. The Adam optimizer was selected as the optimizer, with an initial learning rate of 1*10^-5, and the learning rate decay factor was 10 at the 20th and 40th epochs, and a total of 60 epochs were trained. In addition, the temperature parameter τ was set to 0.02.
[0094] The present invention was compared with other methods in the field of text-pedestrian retrieval, and the best performance and a large improvement were achieved on the three datasets. The experimental results are shown in the following table:
[0095] Table 1 Performance Comparison
[0096]
[0097] As can be seen from the above two tables, under the same experimental environment, compared with the current advanced algorithms, our method shows superior performance.
[0098] The related technical solutions are the same as those in Embodiment 1 and will not be elaborated here.
[0099] Embodiment 3
[0100] This application also relates to an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.
[0101] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory.
[0102] The related technical solutions are the same as above and will not be elaborated here.
[0103] Embodiment 4
[0104] This application also relates to a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0105] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0106] The related technical solutions are the same as above and will not be elaborated here.
[0107] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for constructing a text-based image pedestrian retrieval model, characterized in that: include: An image pedestrian retrieval network is constructed, comprising an image feature enhancement module, a text interference suppression module and a pre-trained dual-modal feature extraction backbone network; wherein the dual-modal feature extraction backbone network is used to encode and obtain the global image features and local image feature vectors of the input image; the image feature enhancement module is used to process the input image to obtain a binary block image vector, perform vector dot multiplication on the block image vector and the local image feature vector to obtain a local image feature vector with background information removed, average and map all features in the vector to obtain the pedestrian enhancement features of the input image; the text denoising module is used to obtain the local text features of the input text corresponding to the input image based on the dual-modal feature extraction backbone network, determine the degree of deviation between the similarity between each local text feature and each feature in the local image feature vector from the average value of all local text features and the similarity between each feature in the local image feature vector, mask the local text corresponding to the local text feature whose quantized value of the deviation is greater than a threshold in the input text, and obtain a new input text; An image-text pair training sample set is constructed. Based on the global image features, pedestrian enhancement features and global text features encoded by the bimodal feature extraction backbone network based on new input text of each sample, the parameters of the bimodal feature extraction backbone network and the operating parameters of the mapping are iteratively adjusted through image-text feature similarity loss to realize the construction of an image pedestrian retrieval model.
2. The construction method according to claim 1, characterized in that: The image feature enhancement module is implemented in the following manner when processing to obtain a binary block image vector: performing human body analysis processing on pedestrians in the input image to obtain a pedestrian contour mask map; dividing the mask map into blocks and binarizing it to obtain a binary block image vector.
3. The construction method according to claim 1, characterized in that: The text noise reduction module is implemented in the following manner when determining the quantitative value of the deviation degree: The similarity between each local text feature and each feature in the local image feature vector with the background information removed is calculated respectively, so as to obtain a first similarity vector corresponding to the local text feature; all local text features are averaged, and the similarity between the average local text feature and each feature in the local image feature vector with the background information removed is calculated respectively, so as to obtain a second similarity vector corresponding to the average local text feature; each first similarity vector is subtracted from the second similarity vector, and the vector after the subtraction is the quantized value of the degree of deviation corresponding to the local text feature.
4. The construction method according to claim 1, characterized in that: The text denoising module implements local text masking based on the quantized value of the deviation degree as follows: All elements in the differenced vector are averaged and added to the preset balance parameter value as the probability that the local text corresponding to the first similarity vector is wrong. The local text corresponding to the probability that meets the threshold is masked in the input text to obtain a new input text.
5. The construction method according to claim 1, characterized in that: The image-text feature similarity loss includes ID loss and SDM loss.
6. The construction method according to claim 5, characterized in that: The method further includes: in each iteration, storing all global image features, pedestrian enhancement features corresponding to the current batch, and global text features encoded by the dual-modal feature extraction backbone network based on new input text into an image feature memory bank and a text feature memory bank, and stabilizing the feature extraction process of the dual-modal feature extraction backbone network by combining a momentum update model based on an exponential moving average; Then the image-text feature similarity loss also includes MITC loss, and its loss function is expressed as: In the formula, Represents the MITC loss from the global image features of the image to the global text features encoded based on the new input text, represents the MITC loss from the global text features encoded based on the new input text to the global image features of the image, The MITC loss represents the pedestrian enhancement feature of the image to the global text feature encoded based on the new input text, MITC loss based on the global text features encoded from the new input text to the pedestrian enhancement features of the image; in, Where τ is a temperature hyperparameter, B is the current batch size in training, N is the total number of image-text pairs; represents the true matching label of the jth image-text pair in the i-th batch, if It means that the current image-text pair is a matching pair from the same pedestrian identity. If It means that the current image-text pair does not match; f i v is the global image feature of all images in the i-th batch, is the global text feature encoded based on the new input text corresponding to the jth image-text pair in the text memory T, is the global text feature encoded based on the new input text corresponding to the nth image-text pair in the text memory; represents f i v and The similarity between represents f i v and The similarity between exist In the formula, f i v and Swap position, f i v and Swap the position, that is The expression of In the formula, f i v Represents the pedestrian enhancement features of all images in the i-th batch, which is The expression of In the formula, f i v and Swap position, f i v and Swap the position, that is expression.
7. The construction method according to claim 1, characterized in that: The pre-trained bimodal feature extraction backbone network is composed of the image encoder and text encoder in the CLIP large model.
8. A text-based image pedestrian retrieval method, characterized in that: include: Obtain a pedestrian image set and a target text of the image to be searched, and input them into a bimodal feature extraction backbone network in the image pedestrian retrieval model respectively, determine the pedestrian image corresponding to the target text from the image set, and realize text-based image pedestrian retrieval; Wherein, the image pedestrian retrieval model is constructed using the construction method described in any one of claims 1 to 7.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method steps according to any one of claims 1 to 7 or the method steps according to claim 8 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored computer program, wherein when the computer program is executed by a processor, the device where the storage medium is located is controlled to perform the method steps according to any one of claims 1 to 7 or the method steps according to claim 8.
Citation Information
Patent Citations
Text pedestrian retrieval method based on text dynamic guidance visual feature extraction
CN113221680A
Text image cross-modal pedestrian retrieval method and system based on implicit relation reasoning alignment
CN116383671A
Cross-modal pedestrian re-identification method based on inter-modal common semantic learning
CN118711217A
Text-based pedestrian image retrieval model establishment method, retrieval method and system
CN118797092A
Image text pedestrian re-identification method and system based on double-flow Transform framework
CN118865442A
Cited By
Text-to-image pedestrian retrieval method and device based on mask denoising and medium
CN120804351A