Shielding pedestrian retrieval method and system based on shielding data simulation and feature enhancement
Through the method of occlusion data simulation and feature enhancement, the pedestrian data set is expanded and the graphic model CLIP-ReID is used for training. Combined with feature extraction and key point detection of ViT and CNN, the problem of insufficient robustness of pedestrian retrieval in occlusion scenarios is solved and the accuracy of pedestrian retrieval is improved.
Patent Information
- Application Number
- CN202510638959.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-22
AI Technical Summary
The existing pedestrian search technology is not robust enough in occlusion scenarios. The public data set is small in scale and the occlusion scenarios are single. The model does not have good occlusion resistance for occlusion pedestrians.
Through the method of occlusion data simulation and feature enhancement, the original pedestrian data set is expanded, and the graphic model CLIP-ReID is used for training, combining feature extraction and key point detection of ViT and CNN, and using multivariate feature fusion and feature enhancement modules to optimize the model's anti-occlusion capability.
Improves the accuracy and robustness of pedestrian search, and the model performs well on multiple data sets and can effectively deal with pedestrian search challenges in occlusion scenarios.
Smart Images

Figure CN120526481A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision technology, and in particular relates to a method and system for retrieving occluded pedestrians based on occlusion data simulation and feature enhancement. Background Art
[0002] Pedestrian retrieval is a person retrieval task, which is usually used to retrieve specific pedestrians from large-scale image or video datasets. The goal is to find the pedestrian image or video clip that is most similar to the pedestrian in the query image by matching the appearance features of the person in the image. Pedestrian retrieval is widely used in security monitoring, intelligent transportation, video surveillance, public safety and other fields. Currently, pedestrian retrieval faces many challenges. Although the clarity of cameras is getting higher and higher, many challenges still exist, such as occlusion, pedestrian changing clothes, lighting changes, and posture changes. In the direction of occluded pedestrian retrieval, the public dataset is small in scale, the occlusion scene is single, the information learned by the model is limited, and it is not very robust to occluded pedestrians.
[0003] Therefore, to address the above problems, a pedestrian retrieval method based on occlusion data simulation and feature enhancement is proposed. It can effectively simulate non-target pedestrian occlusion and object occlusion in reality, and add feature enhancement modules and key point detection to make the model's anti-occlusion ability stronger. Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a method and system for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement.
[0005] To achieve the above-mentioned purpose, the present invention is implemented through the following technical solutions: In a first aspect, the present invention provides a method for retrieval of occluded pedestrians based on occlusion data simulation and feature enhancement, comprising the following steps: S1. Obtain the original pedestrian dataset; S2. The original pedestrian dataset is expanded by occlusion simulation, wherein the occlusion simulation operation includes upper and lower pasting and occlusion patching; the original pedestrian dataset is subjected to the upper and lower pasting operation to obtain the upper and lower pasting dataset, and the occlusion patching operation to obtain the occlusion patch dataset; S3. Construct an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. The model uses the image-text model CLIP-ReID as a baseline model and includes a first-stage training module, a second-stage training module, and a loss calculation module. The first-stage training module includes a ViT-based image encoder, a text generation module, and a Transformer-based text encoder. The text description extracts text features through a Transformer-based text encoder. The second-stage training module includes a multi-feature fusion module and a Transformer-based text encoder. The multi-feature fusion module includes a ViT-based image encoder, a CNN-based image encoder, a key point detection model HRNet, and a feature enhancement module. S4. Optimize the model through the loss function in the loss calculation module to obtain a trained model; S5. Input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder trained in the second stage to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, sort the retrieval results in descending order of similarity scores, and finally obtain the retrieval results of the occluded pedestrian.
[0006] Furthermore, step S2 specifically includes: In the top-bottom pasting operation, the batch size of each original pedestrian dataset is set to 64, with 16 pedestrians and 4 images for each pedestrian. For each pedestrian, the upper half of the first image and the lower half of the second image are combined into a new image, with the pedestrian ID unchanged and corresponding to the first image. The upper half of the second image and the lower half of the first image are combined into a new image, with the pedestrian ID unchanged and corresponding to the second image. The operation between the third and fourth images is the same as that between the first and second images. Finally, the top-bottom pasting training set is obtained. In the occlusion patching operation, the occlusions are cropped from the original pedestrian dataset and then randomly pasted to the top, bottom, left, and right positions of the original pedestrian dataset, finally obtaining the occlusion patch dataset.
[0007] Furthermore, the first stage training module of step S3 specifically includes: The original pedestrian dataset and the occlusion patch dataset images are input into the ViT-based image encoder to obtain image features The images in the original pedestrian dataset and the occlusion patch dataset are input into the text generation module to create a text description for the identity ID of the corresponding pedestrian in each image; the text description is subjected to text feature extraction by the Transformer-based text encoder to obtain text features. .
[0008] Furthermore, the second stage training module of step S3 specifically includes: The images in the original pedestrian dataset, the upper and lower patch dataset, and the occluder patch dataset are input into the ViT-based image encoder for feature extraction to obtain the first extracted feature , the first extracted feature Including the first original pedestrian image feature , the first up-and-down image feature And the first occluder patch image features ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the CNN-based image encoder for feature extraction, and obtain the second extracted feature ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the key point detection model HRNet for feature extraction, and obtain the third extracted feature ; The second feature extraction And the third extracted features Perform feature fusion to obtain the first fusion feature ; Extract the first feature And the third extracted features Perform feature fusion to obtain the second fusion feature ; The formula is as follows: , , in, Represents element-wise multiplication operation; The first original pedestrian image feature and the first up-down image feature Input into the feature enhancement module, the specific process of the feature enhancement module is as follows: Extract the first original pedestrian image features The mean and variance , according to the mean and variance Get the first noise vector with the same shape as the input target feature and the second noise vector , through the first noise vector and the second noise vector For the first up-and-down image feature Perform feature enhancement to obtain the first upper and lower inter-pasted image enhancement feature , the formula is as follows , , , in, Indicates the whitening feature, represents the dot product operation, ; express Obey the mean of 1 and is the normal distribution with variance, Represents the first original pedestrian image feature The population standard deviation of represents a normal distribution; express Obey is the mean, is the normal distribution with variance, Represents the first original pedestrian image feature The overall mean of The first original pedestrian image feature , the first up-and-down image enhancement feature And the first occluder patch image features Integrate together to get the integrated features , the formula is as follows: .
[0009] Furthermore, the loss function of the first-stage training module in step S4 includes text-to-image contrast loss and image-to-text contrast loss: The text-to-image contrast loss function The formula is as follows: , in, Indicates that all labels in the input image scale are The image index set, express The cardinality, , B is the size of the input image during training; represents the product of image feature embedding and text feature embedding, represents the image feature embedding of the p-th picture, represents text feature embedding, represents the image feature embedding of the a-th picture, The base of is the natural logarithm e; The formula is as follows: , in, represents the operation of a linear layer that projects image features into a cross-modal embedding space, represents the operation of a linear layer that projects text features into a cross-modal embedding space, and Represents the image features and labels of the p-th picture respectively. The text features of the picture, the base of is the natural logarithm e; Contrastive loss function for image to text The formula is as follows: , in, and Represent the image feature embedding and text feature embedding of the o-th picture respectively; The total loss function of the first stage training module for: .
[0010] Furthermore, the loss function of the second stage training module in step S4 includes the mean square error loss ID loss , triplet loss , contrastive learning loss : The mean square error loss The formula is as follows: , Among them, n represents the number of images when calculating the mean square error loss, represents the first fusion feature of the i-th image, represents the second fusion feature of the i-th image; Integration features Calculating ID loss and triplet loss , the formula is as follows: , , Where N represents the number of images; represents the value of category z in the target distribution, which is The distribution of each feature in ; Represents the probability value predicted as category z; and They represent the feature distance of positive pairs and the feature distance of negative pairs respectively; represents the interval parameter, set to 0.3; The base of is the natural logarithm e; Indicates the maximum value operation; In the feature enhancement module, the first up-and-down image enhancement feature and the first original pedestrian image feature Do contrastive learning loss : Where N represents the number of images when calculating the contrastive learning loss, represents the features of the first original pedestrian image of the i-th image, represents the enhanced features of the first top-down image of the i-th image; j represents the image sequence number, which is not the same image pair index as i; represents the similarity measure between feature a and feature b, Represents the temperature parameter.
[0011] In a second aspect, the present invention provides an occluded pedestrian retrieval system based on occlusion data simulation and feature enhancement, which executes the occluded pedestrian retrieval method based on occlusion data simulation and feature enhancement, including: Data acquisition module: used to obtain the original pedestrian dataset; Data processing module: used to expand the original pedestrian dataset through occlusion simulation to obtain the upper and lower patch dataset and the occlusion patch dataset; Model construction module: used to build an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. The model uses the image-text model CLIP-ReID as a baseline model, including the first-stage training module, the second-stage training module, and the loss calculation module; Model training module: used to train and optimize the model through the loss function in the loss calculation module to obtain a trained model; Pedestrian retrieval module: It is used to input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder obtained in the second stage of training to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, and sort the retrieval results in descending order according to the similarity score, and finally obtain the occluded pedestrian retrieval results.
[0012] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the occluded pedestrian retrieval method based on occlusion data simulation and feature enhancement as described in the first aspect.
[0013] In a fourth aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the occluded pedestrian retrieval method based on occlusion data simulation and feature enhancement as described in the first aspect are implemented.
[0014] The advantages of the present invention are: The present invention makes full use of the features of images and texts with the help of the powerful performance of the multimodal model CLIP-ReID. Based on the original baseline model, two methods of occlusion data simulation are proposed, and the model can learn the invariance of pedestrians in more occlusion scenarios. Through multi-feature fusion, the model can combine the respective advantages of Transformer and CNN to extract more robust pedestrian features, and use key point detection to guide the model to focus on the parts of the pedestrian's body. Through feature enhancement, the features of the original image and the simulated image can be made similar, preventing the distortion of the pedestrian's body and background confusion caused by the upper and lower pasting. The model can achieve excellent performance on multiple data sets, improving the accuracy of pedestrian retrieval. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide further understanding of the present invention and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention and do not constitute a limitation of the present invention.
[0016] Figure 1 is a flow chart of the steps of the method of the present invention; Figure 2 Comparison of test results between the proposed model and the benchmark model on the same two query images. DETAILED DESCRIPTION
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments derived by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0018] Example 1 In this embodiment, Figure 1 As shown, the present invention provides a method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement, the specific steps of which include: S1. Obtain the original pedestrian dataset; S2. The original pedestrian dataset is expanded by occlusion simulation, wherein the occlusion simulation operation includes upper and lower pasting and occlusion patching; the original pedestrian dataset is subjected to the upper and lower pasting operation to obtain the upper and lower pasting dataset, and the occlusion patching operation to obtain the occlusion patch dataset; Specifically, in the up-down pasting operation, the batch size of each original pedestrian dataset is set to 64, with 16 pedestrians and 4 images for each pedestrian. For each pedestrian, the upper half of the first image and the lower half of the second image are combined into a new image, and the pedestrian ID remains unchanged and corresponds to the first image. The upper half of the second image and the lower half of the first image are combined into a new image, and the pedestrian ID remains unchanged and corresponds to the second image. The operation between the third and fourth images is the same as that between the first and second images. Finally, the up-down pasting training set is obtained. In the occlusion patching operation, the occlusions are cropped from the original pedestrian dataset and then randomly pasted to the top, bottom, left, and right positions of the original pedestrian dataset, finally obtaining the occlusion patch dataset.
[0019] S3. Build an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. This model uses the image-text model CLIP-ReID as a baseline model and includes a first-stage training module, a second-stage training module, and a loss calculation module. The first-stage training module includes a ViT-based image encoder, a text generation module, and a Transformer-based text encoder. In this stage, the ViT-based image encoder extracts image features from the original image and the image with the occluder patch added. Simultaneously, the text generation module creates a text description for the corresponding pedestrian ID in each image, "A photo of a X1 X2... X M person", translated as: a X1 X2... X M Pedestrian photos. Identity ID refers to the serial number annotation for each pedestrian identity in the original dataset image. Each X M is a learnable text token, M represents the number of learnable text tags, and X is the pedestrian's identity tag ID. In the first stage, the initial value of X is the pedestrian's identity tag ID, and then this text description is optimized through subsequent training in the first stage; the text description is extracted using a Transformer-based text encoder to extract text features; The second-stage training module includes a multi-feature fusion module and a Transformer-based text encoder. The multi-feature fusion module includes a ViT-based image encoder, a CNN-based image encoder, a key point detection model HRNet, and a feature enhancement module. Specifically, the first phase training modules include: The original pedestrian dataset and the occlusion patch dataset images are input into the ViT-based image encoder to obtain image features The images in the original pedestrian dataset and the occlusion patch dataset are input into the text generation module to create a text description for the identity ID of the corresponding pedestrian in each image; the text description is subjected to text feature extraction by the Transformer-based text encoder to obtain text features. .
[0020] Specifically, the second phase training modules include: The images in the original pedestrian dataset, the upper and lower patch dataset, and the occluder patch dataset are input into the ViT-based image encoder for feature extraction to obtain the first extracted feature , the first extracted feature Including the first original pedestrian image feature , the first up-and-down image feature And the first occluder patch image features ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the CNN-based image encoder for feature extraction, and obtain the second extracted feature ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the key point detection model HRNet for feature extraction, and obtain the third extracted feature ; The second feature extraction And the third extracted features Perform feature fusion to obtain the first fusion feature ; The first extracted feature And the third extracted features Perform feature fusion to obtain the second fusion feature ; The formula is as follows: , , in, Represents element-wise multiplication operation; The first original pedestrian image feature and the first up-down image feature Input into the feature enhancement module, the specific process of the feature enhancement module is as follows: Extract the first original pedestrian image features The mean and variance , according to the mean and variance Get the first noise vector with the same shape as the input target feature and the second noise vector , through the first noise vector and the second noise vector For the first up-and-down image feature Perform feature enhancement to obtain the first upper and lower inter-pasted image enhancement feature , the formula is as follows , , , in, Indicates the whitening feature, represents the dot product operation, ; express Obey the mean of 1, is the normal distribution with variance, Represents the first original pedestrian image feature The population standard deviation of represents a normal distribution; express Obey is the mean, is the normal distribution with variance, Represents the first original pedestrian image feature The overall mean of The first original pedestrian image feature , the first up and down mutual image enhancement feature And the first occluder patch image features Integrate together to get the integrated features , the formula is as follows: .
[0021] The ViT-based image encoder trained in the first phase and the ViT-based image encoder trained in the second phase are the same encoder, with different internal parameters. This image encoder was trained in the second phase, but not in the first phase. Secondly, the Transformer-based text encoder trained in the first phase and the Transformer-based text encoder trained in the second phase are also the same encoder, with the same parameters. This text encoder was not trained in either phase. Finally, the two-phase training of this algorithm is serial, with the first phase performed first and the second phase performed after the first phase. Both the ViT-based image encoder and the Transformer-based text encoder refer to the same encoder in both phases, but with different internal parameters due to different training settings. The training settings for these two encoders are as follows: In the first phase, neither the ViT-based image encoder nor the Transformer-based text encoder participates in training. In the second phase, the ViT-based image encoder participates in training, while the Transformer-based text encoder does not. If the encoder participates in training, the internal parameters will change. From the perspective of the entire encoder, the encoders involved in the two modules are the same encoder. From the perspective of parameters, the second stage of training is based on the completion of the first stage of training to optimize the parameters. Therefore, if the same encoder participates in training, the parameters will change.
[0022] After the first phase of training, the second phase of training is conducted. The first phase provides semantic constraints for the second phase, as it learns text tokens associated with person IDs, providing important semantic information for optimizing image features in the second phase. These text tokens are used as fixed constraints in the second phase to help the image encoder generate more accurate feature representations. The image features V and text features T obtained from the first phase of training are used to calculate the alignment loss to narrow the feature distance between the image and text. Through the first phase of training, the model is able to learn a set of text feature representations associated with pedestrian IDs (i.e., the text tokens are fine-tuned to adapt to the pedestrian ID representation and image features in preparation for the second phase). These text feature representations associated with pedestrian IDs will be used to optimize the ViT-based image encoder in the second phase to extract more accurate pedestrian representations.
[0023] S4. Optimize the model through the loss function in the loss calculation module to obtain a trained model; Specifically, the loss function of the first stage training module includes text-to-image contrast loss and image-to-text contrast loss: The text-to-image contrast loss function The formula is as follows: , in, Indicates that all labels in the input image scale are The image index set, express The cardinality, , B is the size of the input image during training; Represents the product of image feature embedding and text feature embedding. This product is used as a measure of image-text similarity to constrain the Transformer-based text encoder and increase the similarity between image features and their text features. represents the image feature embedding of the p-th picture, represents text feature embedding, represents the image feature embedding of the a-th picture, The base of is the natural logarithm e; The formula is as follows: , in, represents the operation of a linear layer that projects image features into a cross-modal embedding space, represents the operation of a linear layer that projects text features into a cross-modal embedding space, and Represents the image features and labels of the p-th picture respectively. The text features of the picture, the base of is the natural logarithm e; Contrastive loss function for image to text The formula is as follows: , in, and Represent the image feature embedding and text feature embedding of the o-th picture respectively; The total loss function of the first stage training module for: .
[0024] Specifically, the loss function of the second stage training module includes the mean square error loss ID loss , triplet loss , contrastive learning loss : calculate and The mean squared error loss is used to bring the features based on VitTransformer and CNN closer together, allowing the image encoder based on VitTransformer to combine the advantages of the image encoder based on CNN. This also shortens the feature distance between the original dataset and the dataset simulated by occlusion: The mean square error loss The formula is as follows: , Among them, n represents the number of images when calculating the mean square error loss, represents the first fusion feature of the i-th image, represents the second fusion feature of the i-th image; Integration features Calculating ID loss and triplet loss , the formula is as follows: , , Where N represents the number of images; represents the value of category z in the target distribution, which is The distribution of each feature in ; Represents the probability value predicted as category z; and Represents the characteristic distance of positive and negative pairs respectively (in the calculation When also used The features in , where the opposite is: Two features that are consistent with the pedestrian label ID, Two features that are consistent with the pedestrian label ID, Two features that are consistent with the pedestrian label ID, That is, the feature distance between the two features; negative pair: Two features of inconsistent pedestrian label IDs, Two features of inconsistent pedestrian label IDs, There are two features where the pedestrian tag IDs are inconsistent. That is, the feature distance between these two features. ); represents the interval parameter, set to 0.3; The base of is the natural logarithm e; Indicates the maximum value operation; In the feature enhancement module, the first up-and-down image enhancement feature and the first original pedestrian image feature Do contrastive learning loss : Among them, N represents the number of images when calculating the contrastive learning loss, represents the features of the first original pedestrian image of the i-th image, represents the enhanced features of the first top-down image of the i-th image; j represents the image sequence number, which is not the same image pair index as i; represents the similarity measure between feature a and feature b, Represents the temperature parameter.
[0025] S5. Input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder trained in the second stage to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, sort the retrieval results in descending order of similarity scores, and finally obtain the retrieval results of the occluded pedestrian.
[0026] Example 2 During the training process of this model, text description training is performed as the first stage, with a duration of 120 epochs. Image encoder training is performed as the second stage, with a duration of 360 epochs. To enhance the robustness of the model, the training set undergoes random cropping, flipping, and normalization during dataset preprocessing. This model was trained and tested on the occlusion datasets Occluded-Duke and Occluded-ReID, as well as the full-body datasets Market-1501 and DukeMTMC.
[0027] Tables 1, 2, 3, and 4 show the training and testing results of this experiment on the Occluded-Duke dataset, Occluded-ReID dataset, Market-1501 dataset, and DukeMTMC dataset. The results are compared with the benchmark model CLIP-ReID used in this invention. The method of this invention improves the accuracy and first-place hit rate.
[0028] Table 1 Performance comparison of the proposed model and the benchmark model CLIP-ReID on the Occluded-Duke dataset Table 2 Performance comparison of the proposed model and the benchmark model CLIP-ReID on the Occluded-ReID dataset Table 3 Performance comparison of the proposed model and the benchmark model CLIP-ReID on the Market-1501 dataset Table 4 Performance comparison of the proposed model and the benchmark model CLIP-ReID on the DukeMTMC dataset Example 3 In this embodiment, Figure 2 The figure shows a comparison of the test results of the proposed model and a baseline model on the same two query images, both from the Occluded-Duke dataset. The query is the image of the pedestrian to be queried (the pedestrian is occluded), and the gallery is the image in the gallery. The goal is to search the gallery for pedestrian images with the same ID as the query, based on the query. The proposed model shows the first seven pedestrian images in the gallery. Images marked with a light-colored frame indicate that the pedestrian in this image has the same ID as the query image, and the query is successful; images marked with a dark-colored frame indicate that the pedestrian in this image does not have the same ID as the query image, and the query fails. This figure shows that for the same query, the proposed model generates significantly more successful images than the baseline model, demonstrating the model's stronger occlusion resistance.
[0029] Example 4 This embodiment provides an occluded pedestrian retrieval system based on occlusion data simulation and feature enhancement, including: Data acquisition module: used to obtain the original pedestrian dataset; Data processing module: used to expand the original pedestrian dataset through occlusion simulation to obtain the upper and lower patch dataset and the occlusion patch dataset; Model construction module: used to build an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. The model uses the image-text model CLIP-ReID as a baseline model, including the first-stage training module, the second-stage training module, and the loss calculation module; Model training module: used to train and optimize the model through the loss function in the loss calculation module to obtain a trained model; Pedestrian retrieval module: It is used to input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder obtained in the second stage of training to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, and sort the retrieval results in descending order according to the similarity score, and finally obtain the occluded pedestrian retrieval results.
[0030] Example 5 This embodiment provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement as described in Example 1; Storage media include: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk and other media that can store programs.
[0031] Example 6 This embodiment provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the obscured pedestrian retrieval method based on occlusion data simulation and feature enhancement as described in Example 1 are implemented. Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for retrieval of occluded pedestrians based on occlusion data simulation and feature enhancement, characterized in that: The following steps are involved: S1. Obtain the original pedestrian dataset; S2. The original pedestrian dataset is expanded by occlusion simulation, wherein the occlusion simulation operation includes upper and lower pasting and occlusion patching; the original pedestrian dataset is subjected to the upper and lower pasting operation to obtain the upper and lower pasting dataset, and the occlusion patching operation to obtain the occlusion patch dataset; S3. Construct an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. The model uses the image-text model CLIP-ReID as a baseline model and includes a first-stage training module, a second-stage training module, and a loss calculation module. The first-stage training module includes a ViT-based image encoder, a text generation module, and a Transformer-based text encoder. The text description is extracted using a Transformer-based text encoder to extract text features. The second-stage training module includes a multi-feature fusion module and a Transformer-based text encoder. The multi-feature fusion module includes a ViT-based image encoder, a CNN-based image encoder, a key point detection model HRNet, and a feature enhancement module. S4. Optimize the model through the loss function in the loss calculation module to obtain a trained model; S5. Input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder trained in the second stage to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, sort the retrieval results in descending order of similarity scores, and finally obtain the retrieval results of the occluded pedestrian.
2. The method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement according to claim 1, characterized in that: Step S2 specifically includes: In the top-bottom pasting operation, the batch size of each original pedestrian dataset is set to 64, with 16 pedestrians and 4 images for each pedestrian. For each pedestrian, the upper half of the first image and the lower half of the second image are combined into a new image, with the pedestrian ID unchanged and corresponding to the first image. The upper half of the second image and the lower half of the first image are combined into a new image, with the pedestrian ID unchanged and corresponding to the second image. The operation between the third and fourth images is the same as that between the first and second images. Finally, the top-bottom pasting training set is obtained. In the occlusion patching operation, the occlusions are cropped from the original pedestrian dataset and then randomly pasted to the top, bottom, left, and right positions of the original pedestrian dataset, finally obtaining the occlusion patch dataset.
3. The method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement according to claim 2, characterized in that: The first stage training module of step S3 specifically includes: The original pedestrian dataset and the occlusion patch dataset images are input into the ViT-based image encoder to obtain image features The images in the original pedestrian dataset and the occlusion patch dataset are input into the text generation module to create a text description for the identity ID of the corresponding pedestrian in each image; the text description is subjected to text feature extraction by the Transformer-based text encoder to obtain text features. .
4. The method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement according to claim 3, characterized in that: The second stage training module of step S3 specifically includes: The images in the original pedestrian dataset, the upper and lower patch dataset, and the occluder patch dataset are input into the ViT-based image encoder for feature extraction to obtain the first extracted feature , the first extracted feature Including the first original pedestrian image feature , the first up-and-down image feature And the first occluder patch image features ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the CNN-based image encoder for feature extraction, and obtain the second extracted feature ; Input the images in the original pedestrian dataset, the upper and lower patch dataset, and the occlusion patch dataset into the key point detection model HRNet for feature extraction, and obtain the third extracted feature ; The second feature extraction And the third extracted features Perform feature fusion to obtain the first fusion feature ; The first extracted feature And the third extracted features Perform feature fusion to obtain the second fusion feature ; The formula is as follows: , , in, Represents element-wise multiplication operation; The first original pedestrian image feature and the first up-down image feature Input into the feature enhancement module, the specific process of the feature enhancement module is as follows: Extract the first original pedestrian image features The mean and variance , according to the mean and variance Get the first noise vector with the same shape as the input target feature and the second noise vector , through the first noise vector and the second noise vector For the first up-and-down image feature Perform feature enhancement to obtain the first upper and lower inter-pasted image enhancement feature , the formula is as follows , , , in, Indicates the whitening feature, represents the dot product operation, ; express Obey the mean of 1, is the normal distribution with variance, Represents the first original pedestrian image feature The population standard deviation of represents a normal distribution; express Obey is the mean, is the normal distribution with variance, Represents the first original pedestrian image feature The overall mean of The first original pedestrian image feature , the first up and down mutual image enhancement feature And the first occluder patch image features Integrate together to get the integrated features , the formula is: .
5. The method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement according to claim 4, characterized in that: The loss function of the first stage training module in step S4 includes the contrastive loss from text to image and the contrastive loss from image to text: The text-to-image contrast loss function The formula is as follows: , in, Indicates that all labels in the input image scale are The image index set, express The cardinality, , B is the size of the input image during training; represents the product of image feature embedding and text feature embedding, represents the image feature embedding of the p-th picture, represents text feature embedding, represents the image feature embedding of the a-th picture, The base of is the natural logarithm e; The formula is as follows: , in, represents the operation of a linear layer that projects image features into a cross-modal embedding space, represents the operation of a linear layer that projects text features into a cross-modal embedding space, and Represents the image features and labels of the p-th picture respectively. The text features of the picture, the base of is the natural logarithm e; Contrastive loss function for image to text The formula is as follows: , in, and Represent the image feature embedding and text feature embedding of the o-th picture respectively; The total loss function of the first stage training module for: .
6. The method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement according to claim 5, characterized in that: The loss function of the second stage training module in step S4 includes the mean square error loss ID loss , triplet loss , contrastive learning loss : The mean square error loss The formula is as follows: , Among them, n represents the number of images when calculating the mean square error loss, represents the first fusion feature of the i-th image, represents the second fusion feature of the i-th image; Integration features Calculating ID loss and triplet loss , the formula is as follows: , , Where N represents the number of images; represents the value of category z in the target distribution, which is The distribution of each feature in ; Represents the probability value predicted as category z; and They represent the characteristic distance of positive pairs and the characteristic distance of negative pairs respectively; represents the interval parameter, set to 0.3; The base of is the natural logarithm e; Indicates the maximum value operation; In the feature enhancement module, the first up-and-down image enhancement feature and the first original pedestrian image feature Do contrastive learning loss : Where N represents the number of images when calculating the contrastive learning loss, represents the features of the first original pedestrian image of the i-th image, represents the enhanced features of the first top-down image of the i-th image; j represents the image sequence number, which is not the same image pair index as i; represents the similarity measure between feature a and feature b, Represents the temperature parameter.
7. A system for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement, which executes the method for retrieval of obscured pedestrians based on occlusion data simulation and feature enhancement as claimed in claim 1, characterized in that: include: Data acquisition module: used to obtain the original pedestrian dataset; Data processing module: used to expand the original pedestrian dataset through occlusion simulation to obtain the upper and lower patch dataset and the occlusion patch dataset; Model construction module: used to build an occluded pedestrian retrieval model based on occlusion data simulation and feature enhancement. The model uses the image-text model CLIP-ReID as a baseline model, including the first-stage training module, the second-stage training module, and the loss calculation module; Model training module: used to train and optimize the model through the loss function in the loss calculation module to obtain a trained model; Pedestrian retrieval module: It is used to input the image of the occluded pedestrian to be detected into the trained model, use the ViT-based image encoder obtained in the second stage of training to extract image features, compare the extracted image features with the pedestrian images in the retrieval library for similarity, and sort the retrieval results in descending order according to the similarity score, and finally obtain the occluded pedestrian retrieval results.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the occluded pedestrian retrieval method based on occlusion data simulation and feature enhancement as described in any one of claims 1 to 6 are implemented.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the occluded pedestrian retrieval method based on occlusion data simulation and feature enhancement as described in any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Model training method and device for identifying seal characters, equipment and medium
CN121170810A