A method and system for re-identification of occluded pedestrians based on implicit representation decoupling network

By using the implicit representation of the Deep Self Attention Transformer architecture in pedestrian recognition, the pedestrian component features are automatically decoupled and the occlusion features are separated, and the problems of occlusion interference and high time complexity in the prior art are solved, and robust pedestrian feature extraction and matching are achieved.

CN113901922BActive Publication Date: 2025-05-23SHENZHEN RABBIT PREMISE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111180384.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-11
Publication Date
2025-05-23
Estimated Expiration
2041-10-11

AI Technical Summary

Technical Problem

The existing pedestrian re-identification algorithm with occlusion faces the problems of occlusion and background interference and spatial misalignment of pedestrian body parts, and requires strict and cumbersome alignment of pedestrian body parts, with high time complexity and cannot effectively deal with severe occlusion situations.

Method used

The implicit representation decoupling network based on the deep self-attention transformation network (Transformer) architecture is adopted to automatically decouple pedestrian component features with different semantics by global inference of the local features of the occluded pedestrian image, and the occluded features and target pedestrian features are separated using contrasting feature learning technology.

Benefits of technology

It realizes automatic and robust extraction of pedestrian features and matching in occlusion scenarios, avoids strict pedestrian body parts alignment process, reduces time complexity, and effectively reduces occlusion and noise interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113901922B_ABST
    Figure CN113901922B_ABST
Patent Text Reader

Abstract

A method for re-identifying occluded pedestrians based on an implicit representation decoupling network includes: inputting pedestrian images, enhancing occluded samples, and preprocessing pedestrian images; extracting and decoupling pedestrian features: extracting compact global features of pedestrian images using a convolutional neural network, and using Transformer to decouple the input pedestrian features under the guidance of semantic preference object queries to obtain pedestrian ID-related features and ID-irrelevant features; contrast feature learning: performing opposite discriminative constraints on pedestrian ID-related features and ID-irrelevant features, separating occluders and background noise from pedestrian features, and suppressing the interference of occlusion on pedestrian matching; pedestrian image retrieval: using pedestrian ID-related features to calculate the similarity matrix between the query image and the images in the image library and sorting them, and outputting the sorting results. The method of the present invention can automatically decouple pedestrian semantic features while eliminating occlusion noise interference, and achieve robust pedestrian feature extraction and matching in occluded scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital image processing, and in particular to a method and system for re-identifying occluded pedestrians based on an implicit representation decoupling network. Background Art

[0002] Person Re-identification is a technology that searches for pedestrians that match the query and target in an image or video sequence. Specifically, given a surveillance image of a specific pedestrian as the query target, the pedestrian re-identification system needs to search for other images of the same person taken across cameras in a massive amount of surveillance pedestrian images. With the rapid development of cities and the improvement of traffic camera networks, pedestrian re-identification technology has great application prospects in urban management and public security. In real surveillance scenarios, pedestrian images are often blocked by obstacles, which causes serious difficulties for pedestrian matching. Therefore, it is of great practical significance to carry out research on accurate pedestrian re-identification algorithms with occlusion.

[0003] The main challenges faced by occluded pedestrian re-identification algorithms are the interference of occlusions and backgrounds and the spatial misalignment of pedestrian body parts. Existing occluded pedestrian re-identification algorithms can be roughly divided into two categories. The first type of method uses external models pre-trained based on different data sources, such as human foreground segmentation models, human semantic parsing models, and human pose estimation models to pre-process pedestrian images and generate additional human body part annotations to distinguish between human body parts and occlusions, and accurately match the visible body parts of pedestrians. This type of method relies on the supervision information provided by external models, which is sensitive to occlusions and background noise and prone to errors, and it takes a lot of time to generate labels. The second type of method aligns pedestrian body parts based on the similarity of local images and then measures the similarity. This method is based on strict and cumbersome pedestrian part alignment, has high time complexity, and cannot handle severe occlusion situations. Summary of the invention

[0004] In view of the deficiency of existing methods that require strict, cumbersome and time-consuming alignment of pedestrian body parts, the present invention provides a method and system for re-identification of occluded pedestrians based on implicit representation decoupling network, which utilizes deep self-attention transformer network (Transformer) architecture and contrast feature learning technology, and automatically decouples pedestrian component features with different semantics by performing global reasoning on local features of occluded pedestrian images. Meanwhile, separation of occlusion features and features of target pedestrians can realize re-identification of occluded pedestrians, which overcomes the deficiency of existing methods that require strict, cumbersome and time-consuming alignment of pedestrian body parts, and solves the problem of interference of occlusions on pedestrian feature extraction.

[0005] The technical solution of the present invention is as follows:

[0006] According to one aspect of the present invention, a method for re-identifying occluded pedestrians based on an implicit representation decoupling network is provided, comprising the following steps: S1. inputting pedestrian images, enhancing occluded samples, and preprocessing pedestrian images; S2. extracting and decoupling pedestrian features: extracting compact pedestrian global features of pedestrian images using a convolutional neural network, and decoupling the input pedestrian features under the guidance of semantic preference object queries using a deep self-attention transformer network (Transformer) to obtain pedestrian ID-related features and ID-irrelevant features; S3. contrastive feature learning: performing opposite discriminative constraints on pedestrian ID-related features and ID-irrelevant features, separating occluders and background noise from pedestrian features, and suppressing the interference of occlusion on pedestrian matching; and S4. pedestrian image retrieval: using pedestrian ID-related features to calculate and sort the similarity matrix between the query image and images in the image library, and outputting the sorting result.

[0007] Preferably, in the above-mentioned occluded pedestrian re-identification method based on implicit representation decoupling network, step S1 includes the following sub-steps: D1. Sampling and synthesis of occlusion data: selecting a part of occluders from the training set to construct an occluder set; and D2. Preprocessing the image data input to the network, the preprocessing includes scale normalization and random horizontal flipping, random cropping and random erasing.

[0008] Preferably, in the above-mentioned occluded pedestrian re-identification method based on implicit representation decoupling network, in sub-step D1, during the training phase, a set of occluders is used to perform random occlusion data enhancement on each batch of training data, and the occlusion enhanced data and the original data are used together as the network input of the current batch.

[0009] Preferably, in the above-mentioned occluded pedestrian re-identification method based on implicit representation decoupling network, step S2 includes the following sub-steps: D3. The preprocessed image is input into the convolutional neural network to extract compact pedestrian global features, and then the compact pedestrian global features are flattened into a one-dimensional sequence and assisted by learnable position encoding, and input into the encoder and decoder of the deep self-attention transformer network (Transformer); and D4. Under the guidance of learnable semantic object queries, the decoder of the deep self-attention transformer network (Transformer) decouples the input pedestrian features to obtain pedestrian ID-related features and ID-irrelevant features.

[0010] Preferably, in the above-mentioned occluded pedestrian re-identification method based on implicit representation decoupling network, step S3 also includes the following sub-steps: D5. Using the semantic preference contrast feature learning method, opposite discriminative constraints are imposed on the pedestrian ID-related features and the ID-irrelevant features, so as to separate the occlusions and background noise from the pedestrian features and suppress the interference of occlusion on pedestrian matching; and D6. During the training process of the model, the extracted pedestrian ID-related features are constrained using cross entropy loss and triplet contrast loss, and the pedestrian ID-irrelevant features are constrained using reverse triplet contrast loss.

[0011] Preferably, in the above-mentioned occluded pedestrian re-identification method based on implicit representation decoupling network, in step S4, pedestrian ID-related features output by the model are used to perform pedestrian image retrieval. In the test phase, the similarity matrix between the query image and the image features in the image library is calculated, and the cumulative matching feature curve (CMC) and mean average precision (mAP) are calculated according to the pedestrian re-identification evaluation index.

[0012] According to another aspect of the present invention, an occluded pedestrian re-identification system based on an implicit representation decoupling network is provided, which includes an occluded sample enhancement (OSA) module, a pedestrian feature extraction and semantic decoupling module, and a semantic preference guided contrast feature learning module, wherein the occluded sample enhancement (OSA) module is used to process data to enhance the diversity of occluded samples in each batch of training data; the pedestrian feature extraction and semantic decoupling module is used to first input the preprocessed image into a convolutional neural network to extract compact pedestrian global features, and then flatten the compact pedestrian global features into a one-dimensional sequence and supplemented with a learnable position Encoding, the flattened features of the pedestrian image are input into the deep self-attention transformer network (Transformer) with position encoding, and then the decoder of the deep self-attention transformer network (Transformer) decouples the input pedestrian features under the guidance of learnable semantic object queries to obtain pedestrian ID-related features and ID-irrelevant features; and the semantic preference guided contrast feature learning module is used to perform opposite discriminative constraints on pedestrian ID-related features and ID-irrelevant features, separate occlusions and background noise from pedestrian features, and suppress the interference of occlusion on pedestrian matching.

[0013] According to the technical solution of the present invention, the beneficial effects produced are:

[0014] The present invention proposes a representation decoupling network for pedestrian re-identification and a system, which can automatically decouple the semantic features of pedestrians while eliminating occlusion noise interference, and realize robust pedestrian feature extraction and matching in occlusion scenarios; the present invention uses an implicit representation learning network based on a deep self-attention transformer network (Transformer), which does not require additional semantic supervision information and a complex semantic pre-alignment process to solve the problem of pedestrian re-identification with occlusion; and the method of the present invention designs a contrast feature learning technology and a corresponding data enhancement strategy for the implicit representation decoupling network (DRL-Net), which effectively reduces the interference of occlusion and noise in the pedestrian re-identification task.

[0015] In order to better understand and illustrate the concept, working principle and effect of the present invention, the present invention is described in detail below through specific embodiments in conjunction with the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the specific implementation of the present invention or the technical solution in the prior art, the drawings required for use in the specific implementation or the description of the prior art are briefly introduced below.

[0017] Figure 1 is a flow chart of the occluded pedestrian re-identification method based on implicit representation decoupling network of the present invention; and

[0018] Figure 2 This is a model framework diagram of the occluded pedestrian re-identification system based on the implicit representation decoupling network of the present invention. DETAILED DESCRIPTION

[0019] In order to make the purpose, technical method and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific examples. These examples are only illustrative and not limiting of the present invention.

[0020] like Figure 1 As shown, the occluded pedestrian re-identification method based on implicit representation decoupling network of the present invention includes the following steps:

[0021] S1. Pedestrian image input and occlusion sample enhancement, as well as pedestrian image preprocessing, including sub-steps D1 and D2.

[0022] D1. Sampling and synthesis of occlusion data: Select a part of occluders from the training set and construct an occluder set. During the training phase, use the occluder set to perform random occlusion data enhancement on each batch of training data. The occlusion enhancement data and the original data are used as the network input for the current batch.

[0023] In order to enhance the diversity of occluded samples in each batch of training data, the proposed occluded sample enhancement (OSA) method is used to process the data. The occluded sample enhancement (OSA) method includes:

[0024] 1.1 Before training begins, start with the training set x train Select the occluders from the set x abstacle ;

[0025] 1.2 During the training phase, each training batch From the training set x train P pedestrians with different IDs are selected from the occluders, M image samples are selected for each pedestrian, and k occluders are randomly selected from the occluder set [o 1 , ..., o k |∈x abstacle ;

[0026] 1.3 Use the original training data and the occluder to synthesize the occlusion enhancement data, and use the occlusion enhancement data and the original data as the network input of the current batch. Specifically, for each training batch with a label y i Image Sample Use the selected k occluders to synthesize the occlusion enhancement data [x i,1 , ..., x i,k ] and shares the same label y as the original image i , the occlusion enhanced data and the original data are used together as the network input for the current batch.

[0027] D2. Preprocess the image data (i.e., pedestrian images) input to the network, including scale normalization and data enhancement methods such as random horizontal flipping, random cropping, and random erasing.

[0028] S2. Pedestrian feature extraction and decoupling: Use convolutional neural network to extract compact pedestrian global features from pedestrian images, and use Transformer to decouple the input pedestrian features under the guidance of semantic preference object query to obtain pedestrian ID-related features and ID-irrelevant features. This includes sub-steps D3 and D4.

[0029] D3. Input the preprocessed image into the convolutional neural network to extract the compact global features of pedestrians. Then, flatten the compact global features of pedestrians into a one-dimensional sequence and supplement it with learnable position encoding, and input it into the encoder and decoder of the Transformer.

[0030] In D3, the feature extractor used contains a convolutional neural network and a Transformer encoder-decoder layer. The convolutional neural network is used to extract compact global features of pedestrians, and then the Transformer is used to decouple human body features to generate features of different semantic components.

[0031] Among them, the convolutional neural network is the ResNet-50 residual network, and the Transformer network architecture adopts the same DETR [1] The same standard structure is used, but the prediction head of the predicted class label and bounding box in DETR is removed, that is, the prediction head of Transformer is removed. The number of encoding layers, decoder layers, and multi-head attention are set to 2, 2, and 8 respectively, and learnable position encoding is used.

[0032] The specific operation of using convolutional neural network to extract compact pedestrian global features is as follows: for an input pedestrian image x, the convolutional neural network (ResNet-50) extracts the feature map f = CNN(x)∈R C×H×W , C, H, W represent the channel size, height and width of the feature map respectively. The feature map is obtained by the nonlinear activation function Sigmoidσ(·) a=σ(f)∈R C×H×W , and use a 1×1 convolution to reduce the dimension to d dimensions; flatten the feature map along the last two spatial dimensions, and finally get g∈R d×HW .

[0033] D4.Transformer's decoder, guided by learnable semantic object queries, decouples the input pedestrian features to obtain pedestrian ID-related features and ID-independent features.

[0034] In D4, the encoder-decoder layer of the Transformer follows the standard structure [1] , where Transformer is used to decouple human features, specifically: learnable position encoding is used to encode spatial information, and the position encoding and the feature g extracted by the convolutional neural network are added to the input of each encoder attention layer; in order to generate the features of the semantic components, a set of semantic preference object queries are defined This is a set of learnable input embeddings for the decoding layer, with N q -1 human semantic object query and 1 occlusion semantic query; semantic preference object query is added to the input of the attention layer of the decoder and guides the decoder to decouple the input pedestrian image features into the features of the corresponding semantic components Where N q -1 The features of the semantic parts related to the human body are spliced ​​into the ID related features 1 ID-independent feature generated using occlusion semantic query guidance

[0035] In order to enable Transformer to decouple the features of different semantic components without external supervision, an object query decorrelation constraint is proposed to make object queries orthogonal to each other and encourage object queries to have different semantic preferences. The calculation formula for object query decorrelation constraint loss is as follows:

[0036]

[0037] where abs(·) represents the absolute value function, <·,·> represents the inner product, ||·|| represents the modulus length, α is the penalty factor for the decorrelation constraint loss, and q m and q n Represents the different semantic preference object queries mentioned above.

[0038] S3. Contrastive feature learning: Perform opposite discriminative constraints on pedestrian ID-related features and ID-irrelevant features, separate occlusions and background noise from pedestrian features, and suppress the interference of occlusions on pedestrian matching, including sub-steps D5 and D6.

[0039] D5. Using the semantic preference contrast feature learning method, opposite discriminative constraints are imposed on ID-related features and ID-irrelevant features, occlusions and background noise are separated from pedestrian features, and the interference of occlusion on pedestrian matching is suppressed.

[0040] In D5, the proposed semantic preference guides contrast feature learning, and it is expected that the model can separate occlusion features and pedestrian ID features in an unsupervised manner, eliminating the interference of occlusion noise on pedestrian re-identification.

[0041] Semantic preference guides contrast feature learning as follows: For a given pedestrian image x n , using the occluded sample enhancement (OSA) method proposed in sub-step D1, we can construct x n The comparison triples, including x n Itself is used as an anchor, a pedestrian picture with the same ID but different occlusions is used as a positive sample, and a pedestrian picture with a different ID but the same obstacle is used as a negative sample. For the anchor image x n ID-related features f n , and x n ID-related features f of pedestrian images with the same / different IDs n+ / f n- , the contrastive triplet loss is used to constrain its discriminativeness; the contrastive triplet loss used for ID-related feature constraints is expressed as:

[0042]

[0043] where f n is the anchor image x n ID-related features, f n+ / f n- Respectively represent and x n ID-related features of pedestrian images with the same / different IDs, is a function for calculating feature distance, and δ is a boundary parameter.

[0044] For the anchor image x n ID-independent features and x n ID-independent features of pedestrian images with different / same occlusions The proposed reverse contrast triplet loss is used to impose the opposite discriminative constraint on it, that is, a reverse contrast triplet loss is proposed for ID-independent feature constraint, so that the ID-independent features focus on occlusion and noise. For the same anchor image x n , the positive and negative samples of the reverse contrast triplet loss are opposite to those of the triplet loss. The positive samples are pedestrian images with different IDs but the same obstacles, and the negative samples are pedestrian images with the same ID but different occluders. The reverse contrast triplet loss for ID-independent feature constraints is expressed as:

[0045]

[0046] in is the anchor image x n The ID has nothing to do with the feature, Respectively represent and x n ID-independent features of pedestrian images with different / same occlusions, is a function for calculating feature distance, and δ is a boundary parameter.

[0047] D6. During the model training process, the commonly used cross entropy loss and triple contrast loss are used to constrain the pedestrian ID-related features extracted by the model, and the proposed reverse triple contrast loss is used to constrain the pedestrian ID-irrelevant features.

[0048] In step D6, a label smoothing strategy is used for the adopted cross entropy loss to prevent the model from overfitting the classification training set ID. The label smoothed cross entropy loss formula is expressed as:

[0049]

[0050]

[0051] Where N is the number of training samples, M is the number of pedestrian IDs in the training set, is the feature f n The predicted probability of belonging to IDm, y n Yes n Tags, q m About tag y n is a smooth label for , where ∈ is a small constant.

[0052] The total loss function of the final model is defined as:

[0053]

[0054] S4. Pedestrian image retrieval, using pedestrian ID related features to calculate the similarity matrix between the query image and the images in the image library and sort them, and output the sorting results. This includes sub-step D7.

[0055] D7. Use the pedestrian ID-related features output by the model to perform pedestrian image retrieval. That is, in the test phase, calculate the similarity matrix between the query image and the image features in the image library, and calculate the cumulative matching feature curve (CMC) and mean average precision (mAP) based on the pedestrian re-identification evaluation index.

[0056] Figure 2 The model framework involved in the occluded pedestrian re-identification system based on implicit representation decoupling network of the present invention is as follows: Figure 2 As shown in Figure 2, the model framework consists of three parts: (1) occluded sample enhancement (OSA) module; (2) pedestrian feature extraction and semantic decoupling module; and (3) semantic preference guided contrast feature learning module.

[0057] Among them, the occluded sample enhancement (OSA) module is used to process data to enhance the diversity of occluded samples in each batch of training data. Before the training starts, we first start with the training set x train Select the occluders from the set x abstacle (right Figure 2 Acquisition phase). During the training phase, each training batch From x train P pedestrians with different IDs are selected from the occluders, M image samples are selected for each pedestrian, and k occluders are randomly selected from the occluder set. 1 , ..., o k ]∈x abstacle For each training batch, the label is y i Image Sample Use the selected k occluders to synthesize the occlusion enhancement data [x i,1 , ..., x i, k] and shares the same label y as the original image i, the occlusion enhancement data and the original data are used as the network input of the current batch (corresponding to Figure 2 random synthesis stage).

[0058] (2) Pedestrian feature extraction and semantic decoupling module, which is used to:

[0059] (2.1) First, the preprocessed image is input into the convolutional neural network to extract the compact pedestrian global feature f = CNN(x)∈R C×H×W , and then flatten the compact pedestrian global features into a one-dimensional sequence g∈R d×HW Assisted by learnable position encoding, the flattened features of the pedestrian image are overlaid with position encoding and input into the Transformer encoder.

[0060] (2.2) Under the guidance of learnable semantic object queries, the decoder of Transformer decouples the input pedestrian features and obtains pedestrian ID-related features and ID-independent features. Specifically, in order to generate features of semantic components, a set of semantically preferred object queries are defined, which is a set of learnable input embeddings of the decoding layer. Specifically, there are N q -1 human semantic object query (also known as ID-related object query) and 1 occlusion semantic query. Semantic preference object query is added to the input of the attention layer of the decoder and guides the decoder to decouple the input pedestrian image features into the features of the corresponding semantic components Where N q -1 The features of the semantic parts related to the human body are spliced ​​into the ID related features And an ID-independent feature generated using occlusion semantic query guidance

[0061] In order to enable Transformer to decouple the features of different semantic components without external supervision, an object query decorrelation constraint is proposed to make object queries orthogonal to each other and encourage object queries to have different semantic preferences. The calculation formula for object query decorrelation constraint loss is as follows:

[0062]

[0063] where abs(·) represents the absolute value function, <·,·> represents the inner product, ||·|| represents the modulus length, and α is the penalty factor for the decorrelation constraint loss.

[0064] (3) The semantic preference guided contrast feature learning module is used to impose opposite discriminative constraints on pedestrian ID-related features and ID-irrelevant features, separate occlusions and background noise from pedestrian features, and suppress the interference of occlusions on pedestrian matching.

[0065] The specific operation is: for a given pedestrian image x n , using the occluded sample enhancement (OSA) method proposed in step D1, we can construct x n The comparison triples, including x n itself as an anchor, a pedestrian image with the same ID but different occlusions as a positive sample, and a pedestrian image with a different ID but the same obstacle as a negative sample; for the anchor image x n ID-related features f n , f n+ / f n- Respectively represent and x n The ID-related features of pedestrian images with the same / different IDs are discriminated by using the contrast triplet loss; for the anchor image x n ID-independent features Respectively represent and x n The ID-independent features of pedestrian images with different / same occluders are subjected to opposite discriminative constraints using the proposed reverse contrast triplet loss.

[0066] The contrast triplet loss for ID-related feature constraints is expressed as:

[0067]

[0068] where f n is the anchor image x n ID-related features, f n+ / f n- Respectively represent and x n ID-related features of pedestrian images with the same / different IDs, is a function for calculating feature distance, and δ is a boundary parameter.

[0069] In addition, a reverse contrast triplet loss is proposed for ID-independent feature constraints, so that ID-independent features focus on occlusion and noise. n , the positive and negative samples of the reverse contrast triplet loss are opposite to those of the triplet loss. The positive samples are pedestrian images with different IDs but the same obstacles, and the negative samples are pedestrian images with the same ID but different occluders. The reverse contrast triplet loss for ID-independent feature constraints is expressed as:

[0070]

[0071] in is the anchor image x n The ID has nothing to do with the feature, Respectively represent and xn ID-independent features of pedestrian images with different / same occluders.

[0072] The present invention designs an implicit representation learning network based on Transformer, which does not require strict alignment of human body parts and any additional supervision information to solve the problem of pedestrian re-identification with occlusion. Transformer is a deep neural network based on the "encoder-decoder" architecture and uses a self-attention mechanism. It shows good performance in natural language processing tasks and some recent computer vision tasks. Compared with the traditional convolutional neural network (CNN), Transformer has better performance in semantic feature extraction and long-distance feature capture. The present invention extends Transformer to the study of occluded pedestrian re-identification. First, CNN is used to extract compact local information from the image of the person, and then Transformer is used to perform global reasoning to obtain the features of the target pedestrian for similarity calculation. The occluded pedestrian re-identification method based on the implicit representation decoupling network of the present invention can be used for deep learning-based pedestrian re-identification methods with occlusion in intelligent video surveillance, intelligent security, etc. The method of the present invention uses the Transformer architecture to automatically decouple the features of pedestrian parts with different semantics by performing global reasoning on the local features of the occluded pedestrian image, and uses these features to measure the similarity of two pedestrian images. A contrastive feature learning technique (CFL) is also included to better separate the occlusion features and the features of the target pedestrian.

[0073] The above description is the best embodiment according to the concept and working principle of the invention. The above embodiment should not be understood as limiting the protection scope of the present claims, and other implementations and combinations of implementations according to the concept of the present invention belong to the protection scope of the present invention.

[0074] References:

[0075] [1] N.Carion, F.Massa, G.Synnaeve, N.Usunier, A.Kirillov, and S.Zagoruyko, “End-to-end object detection with transformers,” in Proceedings of the European Conference on Computer Vision (ECCV), 2020, pp.213–229.

Claims

1. An occluded person re-identification method based on implicit representation decoupling network, It is characterized in that The following steps are involved: S1. Perform pedestrian image input, occlusion sample enhancement, and pedestrian image preprocessing; S2. Pedestrian feature extraction and decoupling: A convolutional neural network is used to extract compact pedestrian global features from pedestrian images, and a deep self-attention transformer network is used to decouple the input pedestrian features under the guidance of semantic preference object queries to obtain pedestrian ID-related features and ID-independent features, including the following sub-steps: D3. Input the preprocessed image into the convolutional neural network to extract compact pedestrian global features, and then flatten the compact pedestrian global features into a one-dimensional sequence and supplement it with a learnable position encoding, and input it into the encoder and decoder of the deep self-attention transformer network; D4. The decoder of the deep self-attention transformer network decouples the input pedestrian features under the guidance of the learnable semantic object query to obtain the pedestrian ID-related features and ID-irrelevant features; S3. Contrastive feature learning: performing opposite discriminative constraints on the pedestrian ID-related features and the ID-irrelevant features, separating the occlusions and background noise from the pedestrian features, and suppressing the interference of occlusions on pedestrian matching, including the following sub-steps: D5. Using the semantic preference contrast feature learning method, the pedestrian ID-related features and ID-irrelevant features are subjected to opposite discriminative constraints, occlusions and background noise are separated from pedestrian features, and the interference of occlusion on pedestrian matching is suppressed; D6. During the training process of the model, the extracted pedestrian ID-related features are constrained using cross entropy loss and triple contrast loss, and the pedestrian ID-irrelevant features are constrained using reverse triple contrast loss; as well as S4. Pedestrian image retrieval: use the pedestrian ID related features to calculate the similarity matrix between the query image and the images in the image library and sort them, and output the sorting results.

2. According to claim 1, the occluded pedestrian re-identification method based on implicit representation decoupling network, It is characterized in that Step S1 includes the following sub-steps: D1. Sampling and synthesis of occlusion data: Select a part of occlusion objects from the training set and construct an occlusion object set; as well as D2. Preprocess the image data input to the network, wherein the preprocessing includes scale normalization and random horizontal flipping, random cropping, and random erasing.

3. According to claim 2, the occluded pedestrian re-identification method based on implicit representation decoupling network, It is characterized in that In sub-step D1, during the training phase, the occlusion set is used to perform random occlusion data augmentation on each batch of training data, and the occlusion augmented data and the original data are used together as the network input of the current batch.

4. According to claim 1, the occluded pedestrian re-identification method based on implicit representation decoupling network, It is characterized in that In step S4, pedestrian ID related features output by the model are used to perform pedestrian image retrieval. In the test phase, the similarity matrix between the queried image and the image features in the image library is calculated, and the cumulative matching feature curve and average accuracy are calculated based on the pedestrian re-identification evaluation index.

5. An occluded pedestrian re-identification system based on an implicit representation decoupling network, used to implement the occluded pedestrian re-identification method based on an implicit representation decoupling network as claimed in any one of claims 1 to 4, It is characterized in that It includes an occlusion sample enhancement module, a pedestrian feature extraction and semantic decoupling module, and a semantic preference guided contrast feature learning module. The occlusion sample enhancement module is used to process data to enhance the diversity of occlusion samples in each batch of training data; A pedestrian feature extraction and semantic decoupling module is used to first input the preprocessed image into a convolutional neural network to extract compact pedestrian global features, then flatten the compact pedestrian global features into a one-dimensional sequence and supplement them with learnable position encoding, and then input the flattened features of the pedestrian image and the position encoding into a deep self-attention transformer network, and then the decoder of the deep self-attention transformer network decouples the input pedestrian features under the guidance of a learnable semantic object query to obtain pedestrian ID-related features and ID-irrelevant features; The semantic preference guided contrast feature learning module is used to perform opposite discriminative constraints on the pedestrian ID-related features and the ID-irrelevant features, separate the occlusions and background noise from the pedestrian features, and suppress the interference of occlusions on pedestrian matching.

Citation Information

Patent Citations

  • Shielding downlink pedestrian re-identification model training method and device and shielding downlink pedestrian re-identification method and device

    CN113095263A

  • Local feature alignment pedestrian re-identification method based on deep learning

    CN113221625A