Characteristic learning model based on CLIP large model and reloading pedestrian re-identification method in monitoring scene
Through the feature learning model based on the CLIP large model, the feature decomposition and dependency reduction module are used to extract the non-garment features of dressed pedestrians in the monitoring scenario, solving the problem of insufficient recognition accuracy of dressed pedestrians in the prior art, and achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510207870.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the monitoring scenario, after pedestrians change their clothing, it is difficult for the prior art to effectively extract robust non-garment features, resulting in the accuracy of pedestrians re-identification when changing clothing.
The feature learning model based on the CLIP model is adopted, and the feature decoupling and separation of required features and unnecessary features are decoupled and separated by the feature decomposition module, and the model's dependence on clothing information is reduced by relying on the reduction module, thereby extracting robust required features.
It effectively improves the performance and robustness of the model, improves the accuracy of pedestrian re-identification under dressing conditions, and reduces the dependence on unnecessary features.
Smart Images

Figure CN120047973A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision, image clustering and deep learning, and more specifically, is a method for re-identifying pedestrians who have changed clothes in a monitoring scenario. Background Art
[0002] Person re-identification is an important task in the field of computer vision. Its main goal is to determine whether pedestrian images belong to the same person by comparing them in cross-camera scenarios. This task has a wide range of applications in real life, such as intelligent monitoring, intelligent transportation, public safety, etc., and has important research value and practical significance. In the real world, the appearance of pedestrians usually changes significantly due to changes in clothing, especially in cross-day or long-term monitoring scenarios. In intelligent monitoring, criminal suspects may try to evade monitoring by changing clothes. The task of pedestrian re-identification with clothing change came into being.
[0003] Clothing-changing pedestrian re-identification forces researchers to focus on more robust feature extraction methods, such as human structure features and shape features, and further promotes the transition of pedestrian re-identification from relying on explicit appearance features to deep semantic features. In long-term monitoring tasks (such as missing persons search and criminal tracking), clothing-changing pedestrian re-identification technology can effectively deal with scenarios where suspects try to evade tracking by changing clothes, thereby improving tracking efficiency.
[0004] For pedestrian re-identification, since the clothing part of the pedestrian image occupies most of the pixels in the image, and the features of the clothing part are more significant, it is easier for the model to use it to identify the pedestrian's identity. However, in the re-identification of pedestrians with changing clothes, the clothing features will change significantly when pedestrians change clothes, so clothing cannot be used as a stable and reliable feature to distinguish the pedestrian's identity. In order to make the features extracted by the model more robust, we must let the model extract the features of the non-clothing part of the pedestrian and reduce the dependence on the clothing part as much as possible.
[0005] However, in most previous methods, the reduction of attention to clothing is mainly achieved by, on the one hand, constraining the encoder to extract the clothing part, but it cannot effectively solve the problem without separating clothing elements from non-clothing elements; on the other hand, it is mainly achieved by distancing and removing clothing elements, which may affect non-clothing elements without separating them.
[0006] Therefore, there is an urgent need for a pedestrian re-identification method that retains the robustness of the required features, so as to effectively extract rich and independent required features to effectively re-identify pedestrians in the case of changing clothes. Summary of the invention
[0007] In view of the shortcomings of the prior art, the present invention provides a feature learning model based on the CLIP large model and a method for re-identifying pedestrians in a surveillance scenario. Based on the image-text comparison of CLIP, the method effectively decouples and separates the required features from the unnecessary features through the feature decomposition module, and ensures that the two are relatively independent while containing all the information of the corresponding features; on the other hand, the feature dependency reduction module is used to reduce the dependency of the encoder on unnecessary elements during the training process. Thus, an encoder capable of extracting robust required features is obtained.
[0008] A CLIP-based pedestrian re-identification method with robustness while retaining required features includes: a text extraction module, an image-text comparison module, a feature decomposition module and a dependency reduction module.
[0009] The text extraction module is a frozen model that uses a preset text template to train text prompts that are aligned with the image through image-to-text comparison and text-to-image comparison.
[0010] The image-text comparison module uses the CLIP model to compare images to texts, so that images with the same ID are aligned with the trained text prompts in the text dimension.
[0011] As an embodiment of the present invention, the feature decomposition module uses the feature to project two sub-features, and trains the independence and robustness of the two sub-features through the orthogonality of the two sub-features and alignment with the corresponding real features.
[0012] As an embodiment of the present invention, the dependency reduction module reduces the model's reliance on unnecessary feature information by attempting to remove the reliance on clothing information in the original image generated by the image encoder when the original image passes through the image encoder during the training process.
[0013] Through the above modules, the effect of decoupling the required robust features can be achieved.
[0014] The beneficial effects of the present invention are:
[0015] The present invention provides a method for re-identifying pedestrians with robustness and required features based on CLIP, which effectively improves the performance and robustness of the model through feature decomposition and dependency reduction strategies. The feature decomposition module decouples image features into required features and unnecessary features, and ensures their independence through orthogonality constraints, while maintaining the integrity of each feature, effectively avoiding the problem of feature mixing; the dependency reduction module dynamically reduces the image encoder's dependence on clothing information during training, so that the model focuses more on the extraction of pedestrian identity features, thereby improving the recognition ability under the condition of changing clothes. In addition, through the text extraction and image-text comparison module, the preset text template and the text prompts trained by the CLIP model are used to achieve effective alignment and multi-modal collaborative optimization of image and text information, thereby enhancing the dynamic feature learning ability of the model. The synergy of the above-mentioned innovative modules not only realizes the robust extraction of required features, but also effectively solves the performance bottleneck of traditional methods in re-identification with changed clothes, greatly improving the task performance.
[0016] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0017] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0019] Figure 1 This is a flow chart of a feature learning model based on the CLIP large model and a method for re-identifying pedestrians in a surveillance scenario in accordance with an embodiment of the present invention;
[0020] Figure 2 It is a general model diagram of a feature learning model based on the CLIP large model and a method for re-identifying pedestrians in a surveillance scenario in an embodiment of the present invention;
[0021] Figure 3 It is a feature learning model based on the CLIP large model and a feature decomposition module in a method for re-identifying pedestrians in a surveillance scenario in an embodiment of the present invention;
[0022] Figure 4 The present invention is an overall flow chart of a feature learning model based on the CLIP large model and a method for re-identifying pedestrians in a surveillance scenario according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0024] See also Figure 1 , the embodiment of the present invention provides a feature learning model based on the CLIP large model and a flow chart of a method for re-identifying a pedestrian in a surveillance scene, including: S101, a text extraction module; S102, an image text comparison module; S103, a feature decomposition module; S104, a dependency reduction module;
[0025] The working principle of the above technical solution is as follows: the text extraction module freezes the model, and uses the preset text template to train the text prompts aligned with the image through image-to-text comparison and text-to-image comparison. Then the image-text comparison module uses the CLIP model to align the trained text prompts of the images with the same ID in the text dimension through image-to-text comparison. At the same time, the feature decomposition module uses the feature to project two sub-features, and trains the independence and robustness of the two sub-features through the orthogonality of the two sub-features and alignment with the corresponding real features. At the same time, the dependency reduction module reduces the model's dependence on unnecessary feature information by trying to remove the dependence of the image encoder on the clothing part information in the original image when the original image passes through the image encoder during the training process. Through this process, a model that can effectively extract the required robust features can be obtained.
[0026] The beneficial effects of the above technical solution are: the performance and robustness of the model are effectively improved through feature decomposition and dependency reduction strategies. The feature decomposition module decouples the image features into two parts: required features and unnecessary features, and ensures their independence through orthogonality constraints, while maintaining the integrity of each feature, effectively avoiding the problem of feature mixing; the dependency reduction module dynamically reduces the image encoder's dependence on clothing information during training, so that the model focuses more on the extraction of pedestrian identity features, thereby improving the recognition ability under changing clothes conditions. In addition, through the text extraction and image text comparison module, the preset text template and the text prompts trained by the CLIP model are used to achieve effective alignment and multimodal collaborative optimization of image and text information, thereby enhancing the dynamic feature learning ability of the model. The synergy of the above innovative modules not only realizes the robust extraction of the required features, but also effectively solves the performance bottleneck of traditional methods in changing clothes re-identification, greatly improving the task performance.
[0027] See also Figure 2 , this is the overall model diagram of the present invention.
[0028] The working principle of the above technical solution: Considering that when the CLIP model is used for pedestrian re-identification after changing clothes, the model pays too much attention to clothing features during the pre-training process, and usually relies on these clothing features with large changes, resulting in pedestrians after changing clothes being mistakenly identified as different identities. In order to solve this problem, it is necessary to explore how to adjust the attention paid to clothing-related features and other identity features in image features. One possible way is to fine-tune the CLIP model to strengthen its attention to the stable identity features of pedestrians and reduce its dependence on clothing features. Of course, due to the entanglement between different features, simple partial extraction may not obtain perfect features, so we try to separate the required features from the unnecessary features as much as possible, so that the specified features can be better learned. To this end, the present invention proposes a feature decomposition module and a dependency reduction module, which realizes the decoupling and separation of clothing-related and clothing-irrelevant features through the bidirectional separation of sub-features of feature decomposition, and at the same time attempts to remove the dependence of the image encoder on the clothing part information in the image during the training process. The loss function of the feature decomposition module is roughly as follows:
[0029]
[0030] Among them, it corresponds to a batch of pedestrian image dataset, the size is N; W cr is the projection matrix from the original features to the clothing-related features, and the projection matrix from the original features to the clothing-related features and are the real clothing features and the non-clothing features, which are trained by the extracted clothing images and the images with clothing removed, respectively. and They are respectively two sub-features obtained by decomposing the original features.
[0031] The dependency reduction module loss function is roughly as follows:
[0032]
[0033] where p ID is the predicted classification result, f i is the original image feature, f i xi It is the cross feature of the original image features and the real clothing features.
[0034] The beneficial effects of the above technical solution are as follows: the feature decomposition module decouples the required features from the unnecessary features by decomposing the features into two independent sub-features, and at the same time ensures that the decomposed sub-features contain all the information of the corresponding part of the original image through feature alignment. At the same time, the dependency reduction module is used to reduce the model's dependence on unnecessary features during the training process. Therefore, the present invention achieves the effect of feature decoupling and information retention through feature decomposition and dependency reduction, thereby greatly improving the accuracy of pedestrian re-identification in the case of changing clothes.
[0035] See also Figure 3 In one embodiment, the feature decomposition module can be used to decompose the feature into sub-features of a specific part and separate them, thereby effectively making the features of the two parts of the feature relatively independent, and at the same time ensuring that its sub-features retain all information of the specific part of the original feature through alignment.
[0036] The working principle of the above technical solution is: In order to explore richer feature information in the image, for each image x i We use the proposed pre-trained self-corrected human parsing model to segment pedestrians and clothing from the background, thereby extracting the clothing part in the image and the pedestrian part outside the clothing to obtain and We decompose the features of each original image through two projection matrices, and then constrain the two sub-features to be relatively independent by constraining them to be orthogonal to each other (Formula 3):
[0037]
[0038] Among them, it corresponds to a batch of pedestrian image dataset, the size is N; W cr is the projection matrix from the original features to the clothing-related features, and the projection matrix from the original features to the clothing-related features
[0039] This can ensure that the two sub-features decomposed from the original feature are independent of each other, but it is not enough to just make the two sub-features independent of each other. We also need to make them contain all the information of the corresponding parts in the original image, that is, the clothing-related feature can contain all the feature information of the clothing part in the original image; the clothing-independent feature can contain all the feature information of the pedestrian in the original image except the clothing. To this end, the present invention constrains through feature comparison:
[0040]
[0041] where f i c and f i ncare the actual clothing features and the non-clothing features, which are trained by the extracted clothing images and the images without clothing, respectively. i cr and f i ce They are respectively two sub-features obtained by decomposing the original features.
[0042] See also Figure 4 , a feature learning model based on the CLIP large model and the overall implementation process of the pedestrian re-identification method in the surveillance scenario. First, the image data set is input, and then the multi-modal text prompt is obtained by comparing the image and text. Then, the image and text comparison is used to align all the images of pedestrians with the same identity to the text prompts of the corresponding identity, and feature decomposition is performed to decouple and separate the two features in the feature. At the same time, the model reduces its dependence on the features of the unnecessary parts through dependency reduction. Through the parallel constraints of the above three parts, the model can effectively distinguish between the required features and the unnecessary features, and identify the pedestrians in two different scenes through the required features, and output the re-identification results.
[0043] The working principle of the above technical solution is: before using the pedestrian image dataset for training, it is pre-processed to extract the images of the required part and the images of the unnecessary part of the pedestrian images in the training set. Next, the text prompts of the required part and the text prompts of the parts to be excluded are trained respectively through two text templates. Next, the original features obtained by the pedestrian image encoder of the original image are matched to the text prompts of the required part through image text comparison, and the clothing features obtained by the clothing image encoder of the image of the unnecessary part are matched to the text prompts of the unnecessary part. And the original features are decomposed into two independent sub-features, and the two sub-features are aligned with the features of the required part image and the features of the unnecessary part image respectively, so that we can let the model effectively identify the features of the required part and the unnecessary part, and reduce the loss of its feature information, and effectively extract all the feature information of each part. At the same time, in order to further reduce the model's possible dependence on the clothing part when training the original image, the present invention trains the cross-features of the original features and the clothing features, and trains by subtracting the cross-features from the original features in the middle.
[0044] The beneficial effects of the above technical solution are: through feature decomposition, bidirectional separation and alignment, the two specific features are effectively decoupled and all the information of the corresponding parts are retained respectively, so that the model can extract all the information of the required parts for subsequent pedestrian re-identification work, thereby greatly improving the performance of pedestrian re-identification.
[0045] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of the claims of the present invention.
Claims
1. A feature learning model based on the CLIP large model and a pedestrian re-identification method in a surveillance scenario, the main features of which include: Text extraction module, image-text comparison module, feature decomposition module and dependency reduction module. First, it goes through the text extraction module, and then the parallel image-text comparison module, feature decomposition module and dependency reduction module. This process first trains the required multi-modal text prompts through the text extraction module, and then simultaneously trains an image encoder that can extract the required features in a generalized manner through the supervision of the image-text comparison module, feature decomposition module and dependency reduction module. The text extraction module is a pre-trained model based on CLIP. Through text-to-image comparison and image-to-text comparison, since the model is frozen, the pre-set text prompts are trained, thereby training the text prompts corresponding to the image. The image-text comparison module is based on image-to-text comparison loss, which constrains image features to be effectively aligned with text features through CLIP, so that pedestrian images with the same identity are aligned to corresponding text prompts. The feature decomposition module described above, since the original features of an image are complex, that is, the features of each part are entangled with each other, effectively decoupling and separating the two is a necessary requirement to ensure high performance. There are many effective methods for decoupling and separation, such as feature selection, feature projection, or simple feature separation. However, on the one hand, these methods have a certain impact on the required features, and on the other hand, they do not effectively separate the required features from the unnecessary features. The present invention takes into account the required feature part in the original feature without affecting it, and separates it from the unnecessary feature part in both directions through feature decomposition and alignment. The dependency reduction module is trained using images, which may generate dependencies on clothing in the images during training. In order to reduce the dependencies, we remove the influence of clothing during training, that is, we need to reduce the influence of clothing during training by removing the cross-features of original image features and clothing features.
2. The feature learning model based on the CLIP large model and the method for re-identifying pedestrians in a surveillance scenario according to claim 1 are characterized in that: The features of the pedestrian image dataset are decomposed, and by constraining the two sub-features to be independent of each other and aligned with the corresponding real features, the features of the two parts are decoupled and separated without losing information, so that the features of the two parts do not affect each other and all the corresponding information is retained. At the same time, this feature decoupling method is also adaptable to the decoupling and separation of a wider range of required features and unnecessary features.
3. According to the CLIP large model-based feature learning model and the method for re-identifying pedestrians in a surveillance scenario, the method is characterized in that: The original features and clothing features are used to obtain the cross-features of the two, and then the cross-features are used to reduce the influence of the original feature clothing part on the model during training.
4. According to claim 2, a feature learning model based on CLIP large model and a method for re-identifying pedestrians in a surveillance scene after changing clothes, characterized in that: The features of the pedestrian image dataset are decomposed using two projection matrices, and the two projection matrices are used to constrain the two to be as orthogonal as possible to approximately simulate their mutual independence.
5. According to claim 4, a feature learning model based on CLIP large model and a method for re-identifying pedestrians in a surveillance scenario after changing clothes, characterized in that: By training two reality branches simultaneously, the feature decomposition is supervised by the features of the two reality branches at the same time. While ensuring the information contained in the two sub-features, all the information of the corresponding part of the original feature is retained as much as possible, thereby ensuring the accuracy and robustness of the model in extracting the required features.
Citation Information
Patent Citations
Generative adversarial network system for pedestrian recognition data set enhancement training
CN111382675A