Occluded Pedestrian Re-identification Method, System and Medium Based on Multi-level Feature Refinement

Through multi-level feature refinement convolutional neural network, the problem of insufficient recognition accuracy in occlusion pedestrian recognition is solved. Multi-level feature refinement and multi-branch architecture are adopted to enhance occlusion robustness and improve pedestrian recognition accuracy in occlusion scenarios.

CN115937897BActive Publication Date: 2025-07-29SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211576299.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-09
Publication Date
2025-07-29
Estimated Expiration
2042-12-09

AI Technical Summary

Technical Problem

The existing occlusion pedestrian re-identification methods have insufficient recognition accuracy in occlusion scenarios, making it difficult to effectively utilize the features of occlusion pedestrian images, and methods that rely on pose estimation or attention mechanisms lack robustness.

Method used

A multi-level feature refinement convolutional neural network is designed, and a backbone network is built through ResNet-50 and non-local attention modules are combined with a multi-level multi-branch architecture and multiple attention mechanisms to achieve primary, re-refinement and final refinement of pedestrian characteristics, and finally obtain robust pedestrian search features that block.

Benefits of technology

The accuracy of pedestrian re-identification in occlusion scenarios is improved, and the pedestrian retrieval of occlusion is realized, which enhances the attention to non-occlusion human body parts, while retaining the model's robustness to view changes and internal changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115937897B_ABST
    Figure CN115937897B_ABST
Patent Text Reader

Abstract

The present invention discloses an occluded pedestrian re-identification method, system and storage medium based on multi-level feature refinement. The method includes: acquiring a pedestrian image in an occluded scene and performing preprocessing; using ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network for realizing primary refinement of pedestrian features; constructing a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network for realizing re-refinement and final refinement of features of occluded pedestrians; passing the features output by the two-level multi-branch architecture through a basic operation module in the multi-level feature refinement convolutional neural network to obtain final global features, local features and supplementary features of pedestrians; splicing the obtained features as the final pedestrian features for cross-camera retrieval and matching of pedestrians, so as to realize occluded pedestrian re-identification. The present invention adopts a multi-level feature refinement mechanism, effectively improving the accuracy of occluded pedestrian re-identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to an occluded pedestrian re-identification method, system, terminal device and storage medium based on multi-level feature refinement. Background Art

[0002] With the rapid growth of social productivity, the rapid development of the national economy and the continuous enhancement of national strength, the demand for security protection, event recording and alarm systems is increasing day by day in various fields. As one of the important means of automatic security and alarm, video surveillance has been widely applied in traffic management, public security, military facilities and other aspects, playing a huge role in the stability and progress of society. Although surveillance cameras have been widely installed in public places such as intersections, shopping malls, schools, and stations, their main function is still to record videos and take pictures, and in most cases, actual monitoring work still requires a lot of manual operations.

[0003] Pedestrian re-identification (ReID) aims to solve the problem of retrieving and matching pedestrians with the same identity across cameras by computer, and this method is widely used in multi-camera tracking, video surveillance and criminal search. However, most ReID methods assume that the pedestrians in the matching are full-body standing pedestrian images, which ignores the common occlusion scenarios in the real world. In some real scenarios, such as hospitals, railway stations, shopping malls and other crowded scenes, pedestrians are often partially occluded by other pedestrians or static objects. Due to the lack of partial body information and the difficulty of recognition caused by occlusion, general ReID methods are difficult to achieve satisfactory performance in the occluded pedestrian re-identification task. Therefore, pedestrian re-identification in occlusion scenarios is a challenging research topic.

[0004] The main idea of solving the occlusion problem in ReID is to focus on the non-occluded human body parts from the background noise. Currently, the most advanced occluded pedestrian re-identification methods are basically methods based on key points or attention mechanisms. Key point-based methods usually use pose estimation as additional auxiliary information to make the model focus on the non-occluded parts of the pedestrians. Nevertheless, these methods rely heavily on existing pose estimation models, and there is still room for improvement in their retrieval performance. Attention-based methods usually use the attention mechanism to focus on the non-occluded human body parts in the occluded pedestrian images. Existing attention-based methods lack the full utilization of the results of the feature extraction layer. They focus on obtaining discriminative local features through the attention module, while losing the global features that are robust to subtle view changes and internal changes. Summary of the Invention

[0005] To address the deficiencies of the above-mentioned existing technologies, the present invention provides an occluded pedestrian re-identification method, system, terminal device, and storage medium based on multi-level feature refinement. By designing an end-to-end multi-level feature refinement convolutional neural network and adopting a multi-level feature refinement mechanism, the method realizes the extraction and multi-level refinement of pedestrian image features, and finally obtains pedestrian retrieval features that are robust to occlusion, effectively improving the accuracy of pedestrian re-identification in occlusion scenarios.

[0006] The first object of the present invention is to provide an occluded pedestrian re-identification method based on multi-level feature refinement.

[0007] The second object of the present invention is to provide an occluded pedestrian re-identification system based on multi-level feature refinement.

[0008] The third object of the present invention is to provide a terminal device.

[0009] The fourth object of the present invention is to provide a storage medium.

[0010] The first object of the present invention can be achieved by adopting the following technical solutions:

[0011] An occluded pedestrian re-identification method based on multi-level feature refinement, the method comprising:

[0012] Obtain a pedestrian image in an occlusion scenario and perform preprocessing;

[0013] Use ResNet-50 and a non-local attention module to construct a backbone network in the multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian image to achieve primary refinement of pedestrian features;

[0014] Construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and sub-branches in the two-level multi-branch architecture perform re-refinement and final refinement of the occluded pedestrian features according to the occluded pedestrian features; Input the features output by all sub-branches into a basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features, and supplementary features;

[0015] Concatenate the pedestrian global features, local features, and supplementary features, and use the concatenated features as the final pedestrian features for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in occlusion scenarios.

[0016] Furthermore, the backbone network is a ResNet-50 backbone network enhanced by a non-local attention module, which only includes the modules before conv5_x; the non-local attention module maps the input feature map with three different convolutional layers and downsamples to obtain Query, Key, and Value respectively, then performs a dot product calculation on Query and Key and passes the result through the Softmax function as the attention map of Value, multiplies Value by the attention map and restores the dimension of the input feature map by a convolutional layer, and finally adds the output and the input feature map residually to obtain the final output of the non-local attention module.

[0017] Furthermore, the main branch in the two-stage multi-branch architecture includes a global branch, a self-attention branch, and a cross-attention branch, where:

[0018] The downsampling operation of the conv5_x block in the global branch is cancelled, and the global branch is used to achieve a robust representation of pedestrian features for view changes and internal changes;

[0019] The downsampling operation of the conv5_x block in the self-attention branch is retained. The self-attention branch includes a self-attention module, and a multi-head self-attention module operation is used in the self-attention module to capture long-range dependencies of features;

[0020] The cross-attention mechanism in the cross-attention branch is used to reduce the difference between the features output by the cross-attention branch and the features output by the self-attention branch, so that the attention regions between the two branches are kept roughly the same; in addition, the unique attention region of the cross-attention branch is regarded as a complementary attention region for the non-occluded human body parts missed by the self-attention branch.

[0021] Furthermore, two sub-branches are connected after each main branch, and the two sub-branches perform global max pooling operation and global average pooling operation respectively to obtain richer feature information.

[0022] Furthermore, using a multi-head self-attention module operation in the self-attention module to capture long-range dependencies of features includes:

[0023] Input the preprocessed pedestrian image into the backbone network, and after downsampling the image features output by the backbone network in the self-attention branch, obtain a feature map;

[0024] Convert the feature map into a two-dimensional sequence, and multiply the two-dimensional sequence by three different weight matrices to obtain key, query, and value respectively;

[0025] Average the obtained key, query, and value along the channel direction into h subsets such as {Q1, Q2,..., Q h}, {K1, K2, ..., K h}, and {V1, V2, ..., V h}, for the j-th head, Q j , K j , V j are fed into the independent j-th attention operation head for independent cross-attention operations to capture long-range dependencies in the feature map space; where h is the number of heads in the multi-head self-attention operation;

[0026] The two-dimensional sequences output under each head are concatenated and aggregated, and then transformed into a feature map as the output of the multi-head self-attention module;

[0027] Through the above operations, the distinguishable local features noticed are obtained from the self-attention branch.

[0028] Furthermore, the structure of the cross-attention branch is the same as that of the self-attention branch, except that the key and value in the multi-head cross-attention operation used therein both come from the self-attention main branch; Q j , K j , V j are fed into the independent j-th attention operation head for independent cross-attention operations to supplement the long-range dependencies in the richer feature map space.

[0029] Furthermore, the basic operation module includes the following operations: First, use a convolutional layer to reduce the dimension of the input features, then use a separate BN layer to normalize the features, calculate the triplet loss using the Euclidean distance before the BN layer, and calculate the ID loss using the cosine distance after the BN layer.

[0030] Furthermore, in the training stage of the multi-level feature refinement convolutional neural network, the preprocessed pedestrian images are subjected to image enhancement, and in order to focus the features learned by the network on the discriminative information related to identity, the cross-entropy loss with label smoothing and the triplet loss with hard sample mining are used to supervise the training process.

[0031] Furthermore, the obtaining of the pedestrian images in the occluded scene and the preprocessing thereof include:

[0032] Use a pedestrian detection algorithm to detect and crop pedestrians with the same identity from multiple cameras to obtain the pedestrian images in the occluded scene;

[0033] Save the pedestrian images in the same size.

[0034] The second object of the present invention can be achieved by adopting the following technical solutions:

[0035] An occluded pedestrian re-identification system based on multi-level feature refinement, the system includes:

[0036] A pedestrian image acquisition module, configured to acquire pedestrian images in an occluded scenario and perform preprocessing;

[0037] A primary pedestrian feature refinement module, configured to use ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network, and the backbone network extracts occluded pedestrian features according to the preprocessed pedestrian images to achieve primary refinement of pedestrian features;

[0038] A pedestrian feature re-refinement module, configured to construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and sub-branches in the two-level multi-branch architecture are based on the occluded pedestrian features to achieve re-refinement and final refinement of the occluded pedestrian features; input the features output by all sub-branches into a basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features, and supplementary features;

[0039] A pedestrian re-identification module, configured to splice the pedestrian global features, local features, and supplementary features, and use the spliced features as the final pedestrian features for cross-camera retrieval and matching of pedestrians, so as to achieve pedestrian re-identification in an occluded scenario.

[0040] The third object of the present invention can be achieved by adopting the following technical solutions:

[0041] A terminal device, including a processor and a memory for storing a program executable by the processor. When the processor executes the program stored in the memory, the above-mentioned occluded pedestrian re-identification method is implemented.

[0042] The fourth object of the present invention can be achieved by adopting the following technical solutions:

[0043] A storage medium stores a program, and when the program is executed by a processor, the above-mentioned occluded pedestrian re-identification method is implemented.

[0044] The present invention has the following beneficial effects compared with the prior art:

[0045] 1. Compared with the existing methods, the multi-level feature refinement convolutional neural network designed by the method provided by the present invention is an end-to-end network and does not rely on any auxiliary models.

[0046] 2. The method provided by the present invention adopts a multi-level feature refinement mechanism through a multi-level feature refinement convolutional neural network, utilizes a multi-branch architecture and various attention mechanisms, realizes more refined attention to the non-occluded human body parts in the pedestrian image under the occluded scenario, and combines local features and their complementary features with global features, while paying attention to the local, retains the robustness of the model to overall view changes and internal changes.

[0047] 3. The method provided by the present invention finally obtains pedestrian retrieval features that are robust to occlusion, effectively improving the accuracy of pedestrian re-identification under the occluded scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0049] Figure 1 It is a flowchart of the method for occluded pedestrian re-identification based on multi-level feature refinement in Embodiment 1 of the present invention.

[0050] Figure 2 It is a structural diagram of the multi-level feature refinement convolutional neural network in Embodiment 1 of the present invention.

[0051] Figure 3 It is a structural diagram of the basic operation module in Embodiment 1 of the present invention.

[0052] Figure 4 It is a graph of the learning rate change during the training phase in Embodiment 1 of the present invention.

[0053] Figure 5 It is a structural block diagram of the occluded pedestrian re-identification system based on multi-level feature refinement in Embodiment 2 of the present invention.

[0054] Figure 6 It is a structural block diagram of the terminal device in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention. It should be understood that the described specific embodiments are only used to explain the present application and are not used to limit the present application.

[0056] Example 1:

[0057] As Figure 1 shown, this embodiment provides an occluded pedestrian re-identification method based on multi-level feature refinement, which mainly includes the following steps:

[0058] S101. Obtain pedestrian images in an occluded scene and perform preprocessing.

[0059] Use a pedestrian detection algorithm to detect and crop pedestrians with the same identity from multiple cameras to obtain pedestrian images in an occluded scene.

[0060] The preprocessing includes changing the size of the pedestrian images, as well as image enhancement operations such as randomly horizontally flipping, randomly cropping, and randomly erasing the pedestrian images during the training phase.

[0061] Specifically, the preprocessing includes changing the size of the pedestrian images to 384×128, randomly horizontally flipping the pedestrian images with a probability of 50% during the training phase, randomly cropping them back to the original size after magnifying and padding by 10 pixels, and randomly erasing a region of the input image with a probability of 50% and other image enhancement operations.

[0062] S102. Use ResNet-50 and a non-local attention module to construct the backbone network in the multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian images and realizes the primary refinement of pedestrian features.

[0063] As Figure 2 shown, the backbone network of the multi-level feature refinement network is a ResNet-50 backbone network enhanced by a non-local attention module. Specifically, ResNet-50 in the backbone network is first pre-trained on the Imagenet dataset. Only the modules before conv5_x are included in the backbone network of this method, and its conv5_x module is processed differently and used in three main branches of the subsequent branches; the non-local attention module linearly maps the input feature map with three different 1×1 convolutional layers and reduces the dimension to obtain Query, Key, and Value respectively. Then, the dot product is calculated between Query and Key, and the result is passed through the Softmax function as the attention map of Value. Multiply Value by this attention map. Thus, the non-local operation is realized. The output of the non-local operation is input into a new 1×1 convolutional layer to restore the dimension of the original input feature map. Finally, this output is added to the original input as a residual to obtain the final output, that is, the output z of the non-local attention module is realized according to the following formula i :

[0064] z i = Wz × φ(x i ) + x i

[0065] Among them, W z is the weight matrix to be learned, and φ(x i ) represents the residual formed by the non-local operation and the input x i .

[0066] S103. Construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and the sub-branches in the two-level multi-branch architecture are based on the occluded pedestrian features to realize the re-refinement and final refinement of the occluded pedestrian features; input the features output by all sub-branches into the basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features and supplementary features.

[0067] As Figure 2 shown, the two-level multi-branch architecture includes a main branch with three branches and six sub-branches formed by connecting two branches behind each main branch. The three main branches include a global branch, a self-attention branch and a cross-attention branch.

[0068] (1) Global branch.

[0069] To obtain a larger feature map, the downsampling operation of the conv5_x block in the global branch is cancelled; to capture richer statistical characteristics, there are two sub-branches behind the global main branch, which perform global max pooling operation (GMP) and global average pooling operation (GAP) respectively; the global branch and its sub-branches learn global features by capturing the global receptive field, and finally are used to realize the robust representation of pedestrian features to view changes and internal changes.

[0070] (2) Self-attention main branch.

[0071] The self-attention main branch contains a self-attention module. The downsampling operation in the conv5_x block is retained in the self-attention branch to reduce the size of the model and the computational amount of the self-attention module; in the self-attention module, the multi-head self-attention module operation is used to capture the long-distance dependence of features to solve the problem of missing partial human features in the occluded pedestrian images. The multi-head mechanism ensures that the network can pay attention to multiple distinguishable human parts, and each head can be regarded as a separate attention operation.

[0072] Specifically, after inputting the image into the backbone network and performing downsampling in this branch, a feature map is obtained where c, h, and w represent the number of channels, height, and width respectively; referring to the design of the ViT network, Layernorm is applied before the multi-head self-attention operation, and a residual connection is applied afterwards; since the input required for the multi-head self-attention operation is a series of one-dimensional sequences, the feature map F is transformed into a two-dimensional sequence, and the spatial dimension and channel dimension of the feature map are exchanged for convenient calculation, and finally Then, the two-dimensional sequence f is multiplied by three different weight matrices W Q , W K , W V to obtain key, query, and value respectively, and the expressions are:

[0073] Q = fW Q , K = fW K , V = fW V ,

[0074] where W K , are both linear mappings, d k is the dimension of the key and query input to the attention module, and d v is the dimension of the value input to the attention module;

[0075] After that, Q, K, and V are each evenly divided into h subsets along the channel direction, such as {Q1, Q2,..., Q h} where h is the number of heads in the multi-head self-attention operation; for the j-th self-attention operation head, Q j , K j , V j are fed into it for independent self-attention operations to capture the long-range dependencies in the feature map space, and the specific calculation is as follows:

[0076]

[0077] where is a scaling factor; finally, the attention values calculated under each head are aggregated and projected into a new feature map, and the calculation is as follows:

[0078] MultiHead(Q, K, V) = Concat(head1, head2..., head h )W O ,

[0079] where is the parameter matrix of the linear projection, and Concat(·) is the concatenation operation of the feature maps output by each operation head;

[0080] Through a series of operations in the above self-attention module, discriminative local features can be obtained from the self-attention main branch; then, the output features of the self-attention module are processed using the same two pooling operations as the global main branch.

[0081] (3) Cross-attention branch.

[0082] The structure of the cross-attention main branch is basically the same as that of the self-attention main branch, except that the key and value in the multi-head cross-attention operation used therein both come from the self-attention main branch; for the j-th head, Q j , K j , V j are fed into an independent cross-attention operation to supplement the long-range dependencies in the richer feature map space, and the calculation is as follows:

[0083]

[0084] where self and cross respectively represent from the self-attention branch and the cross-attention branch. By reducing the difference between the features output by the cross-attention main branch and the features of the self-attention main branch through the cross-attention mechanism, the attention regions between the two branches can be kept roughly consistent; in addition, the unique attention region of the cross-attention main branch can be regarded as a complementary attention region for the non-occluded human body parts missed by the self-attention main branch; in the sub-branch of this main branch, the same pooling operations as those of other main branches are used;

[0085] As Figure 3 shown, the basic operation module includes three parts of operations. First, a 1×1 convolutional layer is used to reduce the dimension of the features. The two 2048-dimensional features of the global sub-branch are respectively reduced to 512 dimensions, and the features of other sub-branches are respectively reduced to 256 dimensions; then, since there is no consistent objective for the ID loss and the triplet loss in the embedding space, a separate BN layer is used to normalize the features. The Euclidean distance is used to calculate the triplet loss before the BN layer, and the cosine distance is used to calculate the ID loss after the BN layer.

[0086] S104. Concatenate the features of the above-obtained pedestrians as the final pedestrian features for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in an occlusion scenario.

[0087] Specifically, feature concatenation is performed in the test retrieval stage. In the training stage, in order to focus the features learned by the network on the discriminative information related to identity, cross-entropy loss with label smoothing and triplet loss with hard sample mining are used to supervise the training process. The cross-entropy loss aims to guide the network to directly learn the mapping between samples and specific identities, where label smoothing is set to reduce the confidence of the training set labels to improve the generalization ability of the network. The triplet loss aims to achieve identity clustering by minimizing the feature distance between positive samples while maximizing the feature distance between negative samples. The purpose of using hard samples to train the network is also to improve the generalization ability of the network. The cross-entropy loss with label smoothing can be expressed as:

[0088]

[0089]

[0090] where N represents the total number of classes, y is the true ID label, p i is the ID prediction logarithm of the i-th class, and ε is a hyperparameter used to reduce the confidence of the training set labels and improve the generalization ability of the model. In this example, it is 0.1. The triplet loss with hard sample mining can be expressed as:

[0091]

[0092] where α is a threshold parameter set manually, [z] + represents max(z, 0), P represents the number of identities sampled in a batch, K is the number of pedestrian images for each identity. The triplet loss with hard sample mining calculates the Euclidean distance of each sample in the batch in the feature space, and then selects the positive sample f a farthest from the anchor sample in the feature space and the negative sample f p nearest to the anchor sample f n to calculate the triplet loss. Finally, the total loss function is as follows:

[0093] L = λL ID + βL tri

[0094] where λ and β are regularization parameters used to balance different losses, and the specific weights are 1.5 and 2 respectively. The losses are calculated separately for the features of each sub-branch after dimensionality reduction, and the average loss of each branch is used as the training loss. The curve of the learning rate changing with the training cycle during the training process is as Figure 4 shown, including the warm-up strategy and the cosine annealing strategy. [[ID=3〕]

[0095] In this example, α is taken as 0.3, P is taken as 16, and K is taken as 4.

[0096] In the test stage, the features after the BN layer of all sub-branches are concatenated to obtain 2048-dimensional features for the final pedestrian retrieval to achieve pedestrian re-identification.

[0097] It should be noted that although the method operations of the above embodiments are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. On the contrary, the described steps can be changed in the execution order. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.

[0098] Embodiment 2:

[0099] As Figure 5 shown, this embodiment provides an occluded pedestrian re-identification system based on multi-level feature refinement. The system includes a pedestrian image acquisition module 501, a primary pedestrian feature refinement module 502, a secondary pedestrian feature refinement module 503, and a pedestrian re-identification module 504, where:

[0100] The pedestrian image acquisition module 501 is configured to acquire a pedestrian image in an occluded scene and perform preprocessing;

[0101] The primary pedestrian feature refinement module 502 is configured to use ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian image to achieve primary refinement of the pedestrian features;

[0102] The secondary pedestrian feature refinement module 503 is configured to construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and sub-branches in the two-level multi-branch architecture are based on the occluded pedestrian features to achieve secondary refinement and final refinement of the occluded pedestrian features; input the features output by all sub-branches into a basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features, and supplementary features;

[0103] The pedestrian re-identification module 504 is configured to concatenate the pedestrian global features, local features, and supplementary features, and use the concatenated features as the final pedestrian features for cross-camera retrieval and matching of pedestrians to achieve occluded pedestrian re-identification.

[0104] For the specific implementation of each module in this embodiment, reference can be made to Embodiment 1 above, which will not be elaborated here one by one. It should be noted that the system provided in this embodiment is only illustrated by the above division of each functional module. In practical applications, the above functions can be allocated to different functional modules as needed, that is, the internal structure can be divided into different functional modules to complete all or part of the functions described above.

[0105] Embodiment 3:

[0106] This embodiment provides a terminal device, which can be a computer. As Figure 6 shown, it includes a processor 602, a memory, an input device 603, a display 604, and a network interface 605 connected through a system bus 601. The processor is used to provide computing and control capabilities. The memory includes a non-volatile storage medium 606 and an internal memory 607. The non-volatile storage medium 606 stores an operating system, a computer program, and a database. The internal memory 607 provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the processor 602 executes the computer program stored in the memory, the occluded pedestrian re-identification method of Embodiment 1 above is implemented as follows:

[0107] Obtain a pedestrian image in an occluded scene and perform preprocessing;

[0108] Use ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian image to achieve primary refinement of pedestrian features;

[0109] Construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and sub-branches in the two-level multi-branch architecture refine the occluded pedestrian features again and finally according to the occluded pedestrian features; Input the features output by all sub-branches into a basic operation module in the multi-level feature refinement convolutional neural network to obtain the final global pedestrian feature, local feature, and supplementary feature;

[0110] Concatenate the global pedestrian feature, local feature, and supplementary feature, and use the concatenated feature as the final pedestrian feature for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in an occluded scene.

[0111] Embodiment 4:

[0112] This embodiment provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the occluded pedestrian re-identification method of Embodiment 1 above is implemented as follows:

[0113] Obtain a pedestrian image in an occluded scene and perform preprocessing;

[0114] Use ResNet-50 and a non-local attention module to construct the backbone network in a multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian image to achieve primary refinement of pedestrian features;

[0115] Construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and sub-branches in the two-level multi-branch architecture refine the occluded pedestrian features again and finally based on the occluded pedestrian features; input the features output by all sub-branches into the basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features, and supplementary features;

[0116] Concatenate the pedestrian global features, local features, and supplementary features, and use the concatenated features as the final pedestrian features for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in an occluded scene.

[0117] It should be noted that the computer-readable storage medium in this embodiment can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0118] In summary, the method provided by the present invention includes: obtaining a pedestrian image in an occluded scene and performing preprocessing; using ResNet-50 and a non-local attention module to construct the backbone network of a convolutional neural network for extracting occluded pedestrian features and achieving primary refinement of pedestrian features; constructing a two-level multi-branch architecture, where the main branch and sub-branches in the two-level multi-branch structure refine the pedestrian features again and finally respectively, and obtaining the final pedestrian global features, local features, and supplementary features by passing the features output by all sub-branches through the basic operation module; finally, concatenating the above-obtained pedestrian features as the final pedestrian features for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in an occluded scene. The present invention designs the structure of a convolutional neural network, adopts a multi-level feature refinement mechanism to achieve the extraction and multi-level refinement of pedestrian image features, and finally obtains pedestrian retrieval features that are robust to occlusion, effectively improving the accuracy of pedestrian re-identification in an occluded scene.

[0119] As described above, it is only a preferred embodiment of the present invention patent. However, the protection scope of the present invention patent is not limited thereto. Any person skilled in the art within the scope disclosed by the present invention patent, making equivalent substitutions or changes based on the technical solution and inventive concept of the present invention patent, shall fall within the protection scope of the present invention patent.

Claims

1. An occluded pedestrian re-identification method based on multi-level feature refinement, characterized in that, The method includes: Obtaining a pedestrian image in an occlusion scenario and performing preprocessing; Using ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian image to achieve primary refinement of pedestrian features. The backbone network is a ResNet-50 backbone network enhanced by a non-local attention module, and only includes modules before conv5_x. The non-local attention module maps the input feature map with three different convolutional layers and reduces the dimension to obtain Query, Key, and Value respectively. Then, a dot product calculation is performed on Query and Key, and the result is passed through the Softmax function as the attention map of Value. Multiply Value by the attention map and restore the dimension of the input feature map by a convolutional layer. Finally, add the output and the input feature map residually to obtain the final output of the non-local attention module; Constructing a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and the sub-branches in the two-level multi-branch architecture refine and finally refine the occluded pedestrian features according to the occluded pedestrian features. Input the features output by all sub-branches into the basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features, and supplementary features. The main branch in the two-level multi-branch architecture includes a global branch, a self-attention branch, and a cross-attention branch. Among them, the downsampling operation of the conv5_x block in the global branch is cancelled, and the global branch is used to achieve a robust representation of pedestrian features for view changes and internal changes. The downsampling operation of the conv5_x block in the self-attention branch is retained. The self-attention branch includes a self-attention module, and a multi-head self-attention module operation is used in the self-attention module to capture the long-distance dependence of features. The cross-attention mechanism in the cross-attention branch is used to reduce the difference between the features output by the cross-attention branch and the features output by the self-attention branch, so that the attention regions between the two branches are kept roughly the same. In addition, the unique attention region of the cross-attention branch is regarded as a supplementary attention region for the non-occluded human body parts missed by the self-attention branch; Concatenate the pedestrian global features, local features, and supplementary features, and use the concatenated features as the final pedestrian features for pedestrian cross-camera retrieval and matching to achieve pedestrian re-identification in an occlusion scenario.

2. The occluded pedestrian re-identification method according to claim 1, wherein After each main branch, two sub-branches are connected, and the two sub-branches perform global max pooling operation and global average pooling operation respectively to obtain richer feature information.

3. The occluded person re-identification method according to claim 2, characterized in that The use of a multi-head self-attention module operation in the self-attention module to capture the long-distance dependence of features includes: Input the preprocessed pedestrian image into the backbone network, and after downsampling the image features output by the backbone network in the self-attention branch, obtain a feature map; Convert the feature map into a two-dimensional sequence, and multiply the two-dimensional sequence by three different weight matrices to obtain key, query, and value respectively; The obtained key, query, and value are respectively averaged into h subsets along the channel direction, such as { Q 1. Q 2 ,…, Q h}, { K 1. K 2 ,…, K h}, and { V 1. V 2 ,…, V h}. For the j th head, Q j , K j , V j are fed into the independent j th attention operation head for independent cross-attention operations to capture long-range dependencies in the feature map space; where h is the number of heads in the multi-head self-attention operation; Concatenate and summarize the two-dimensional sequences output under each head, and convert them into feature maps as the output of the multi-head self-attention module; Through the above operations, obtain the discriminative local features noticed from the self-attention branch.

4. The occluded person re-identification method according to claim 3, wherein The structure of the cross-attention branch is the same as that of the self-attention branch, except that the keys and values in the multi-head cross-attention operation used therein both come from the self-attention main branch; Q j and K j and V j are sent to an independent j th attention operation head for independent cross-attention operation to supplement the long-range dependencies in the richer feature map space.

5. The occluded person re-identification method according to claim 1, wherein The basic operation module includes the following operations: First, use a convolutional layer to reduce the dimension of the input features, then use a separate BN layer to normalize the features, calculate the triplet loss using the Euclidean distance before the BN layer, and calculate the ID loss using the cosine distance after the BN layer.

6. The occluded person re-identification method according to claim 1, wherein In the training stage of the multi-level feature refinement convolutional neural network, perform image enhancement on the preprocessed pedestrian images, and in order to focus the features learned by the network on the discriminative information related to identity, use the cross-entropy loss with label smoothing and the triplet loss with hard sample mining to supervise the training process.

7. The occluded pedestrian re-identification method according to any one of claims 1-6, characterized in that The acquisition of pedestrian images in an occlusion scenario and preprocessing thereof include: Use a pedestrian detection algorithm to detect and crop pedestrians with the same identity from multiple cameras to obtain pedestrian images in an occlusion scenario; Save the pedestrian images to the same size.

8. An occluded pedestrian re-identification system based on multi-level feature refinement, characterized in that The system includes: A pedestrian image acquisition module for acquiring and preprocessing pedestrian images in an occlusion scenario; A pedestrian feature primary refinement module for using ResNet-50 and a non-local attention module to construct a backbone network in a multi-level feature refinement convolutional neural network. The backbone network extracts occluded pedestrian features from the preprocessed pedestrian images to achieve primary refinement of pedestrian features; the backbone network is a ResNet-50 backbone network enhanced by a non-local attention module, and only includes the modules before conv5_x; the non-local attention module maps the input feature map with three different convolutional layers and reduces the dimension to obtain Query, Key, and Value respectively, then performs a dot product calculation on Query and Key and passes the result through the Softmax function to be used as the attention map of Value, multiplies Value by the attention map and restores the dimension of the input feature map by a convolutional layer, and finally adds the output and the input feature map as a residual to obtain the final output of the non-local attention module; The pedestrian feature refinement module is used to construct a two-level multi-branch architecture in the multi-level feature refinement convolutional neural network. The main branch and the sub-branches in the two-level multi-branch architecture are based on the occluded pedestrian features to realize the re-refinement and final refinement of the occluded pedestrian features; the features output by all sub-branches are input into the basic operation module in the multi-level feature refinement convolutional neural network to obtain the final pedestrian global features, local features and supplementary features; the main branch in the two-level multi-branch architecture includes a global branch, a self-attention branch and a cross-attention branch. Among them, the downsampling operation of the conv5_x block in the global branch is cancelled, and the global branch is used to realize the robust representation of pedestrian features to view changes and internal changes; the downsampling operation of the conv5_x block in the self-attention branch is retained. The self-attention branch includes a self-attention module, and the multi-head self-attention module operation is used in the self-attention module to capture the long-distance dependence of features; the cross-attention mechanism in the cross-attention branch is used to reduce the difference between the features output by the cross-attention branch and the features output by the self-attention branch, so that the attention regions between the two branches are kept roughly the same; in addition, the unique attention region of the cross-attention branch is regarded as a supplementary attention region for the non-occluded human body parts missed by the self-attention branch; The pedestrian re-identification module is used to splice the pedestrian global features, local features and supplementary features, and use the spliced features as the final pedestrian features for pedestrian cross-camera retrieval and matching to realize pedestrian re-identification in occluded scenarios.

Citation Information

Patent Citations

  • Pedestrian re-identification method based on second-order mixed attention

    CN112733590A

  • Multi-scale pedestrian re-identification method based on multi-granularity depth feature fusion

    CN112818931A