A target re-identification method, system, device and medium
The swimmer network encoder and dual-mode decoder optimize re-ID models for NPUs, addressing the challenge of speed and accuracy on mobile devices by reducing computational load and model size, enhancing feature extraction efficiency.
Patent Information
- Application Number
- CN202211165576.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-09-23
AI Technical Summary
The existing target re-identification technology has poor forward speed in NPU on-board. The encoder and decoder have problems with large calculations and many operations that do not support, making it difficult to achieve efficient operation on mobile devices.
The naive block re-identification model is adopted, the encoder uses a goggle network, and the decoder uses a dual-mode network. Through the reasonable combination of 3x3 convolution and channel segmentation, the feature extraction process is optimized, the calculation amount is reduced, and the on-board acceleration operator is used for acceleration.
While ensuring the accuracy of target re-identification, the forward speed of NPU on-board is improved, efficient operation on mobile devices is achieved, and the feature representation ability is stronger than that of existing models.
Smart Images

Figure CN115471817B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image retrieval, relates to the field of the combination of pedestrian re-identification and mobile terminals, and particularly relates to a target re-identification method, system, device and medium. Background Art
[0002] Object re-identification aims to search for the same target under different cameras. This technology is mainly applied to multi-object tracking and plays an important algorithmic role in criminal investigation tasks and autonomous driving tasks. In recent years, in the field of re-identification, research on feature representation has been highly concerned by scholars. With the development of deep learning technology, the combination of re-ID models and it has become increasingly close. Therefore, a large number of neural networks are included, and their structures are becoming more and more complex. However, autonomous driving is a system that requires machines to make decisions in a very short time. As a subtask in autonomous driving, re-ID will undoubtedly also be required to achieve real-time performance and accuracy in in-vehicle systems with limited computing resources.
[0003] Currently, in the daily algorithm development of researchers in the field of artificial intelligence, the experimental equipment used mainly consists of high-performance computers equipped with GPUs, while NPUs are often integrated into mobile artificial intelligence board hardware (such as GX8010, MLU100, RK3399Pro, etc.). Specifically, NPU (Neural Processing Units) is a hardware chip specifically designed for quickly implementing neural networks and is currently widely used in mobile edge computing scenarios.
[0004] Compared with GPUs, NPUs have the advantages of high efficiency and low power consumption when running neural network models. For the same neural network model, the frame rate on NPUs can usually reach several times that on GPUs. At the same time, NPUs also have a relatively simpler operator combination than GPUs to ensure their excellent speed, and have smaller memory and computing power to maintain an ideal power consumption level. This causes some limiting conditions when deploying models to NPUs. For example, more complex operations may not be supported by NPUs, and a large number of model parameters may cause the model to be unable to be deployed on NPUs. In summary, for models with NPU on-board as the final implementation method, these limiting factors need to be considered in the algorithm design stage.
[0005] In the field of re-ID, there is little research on algorithm optimization based on the characteristics of NPU when designing models. Therefore, it is difficult for existing re-ID algorithms to achieve both accuracy and onboard running speed on NPU. Structurally, the re-ID model can be divided into two parts: encoder and decoder. The encoder and decoder currently used in the re-ID field have some problems when deployed on NPU, including:
[0006] (1) The encoder is mainly responsible for mapping shallow image texture information to deep semantic features. The most commonly used encoder in existing re-ID models is ResNet50, which is not a lightweight method. The huge number of parameters and computational complexity make it unsuitable for deployment on mobile devices. In current lightweight encoders, due to the extensive use of depthwise separable convolutions, memory access consumption is too high, which slows down the onboard speed. Therefore, although these existing lightweight methods have fewer parameters and computational complexity, they do not achieve the purpose of acceleration through lightweighting in the actual hardware onboard process.
[0007] (2) For the decoder, it is usually responsible for deeper information mining and conversion of the feature map generated by the encoder. Although existing decoders can achieve or even exceed human retrieval accuracy in re-ID tasks, most of them contain complex structures (such as attention mechanism operations) or operations that are not supported by the NPU (such as graph convolution, spatial transformation network, etc.), making it difficult for the model to run at an ideal speed on the NPU. Summary of the invention
[0008] The purpose of the present invention is to provide a target re-identification method, system, device and medium to solve one or more of the above-mentioned technical problems. The method provided by the present invention can solve the problem that the encoder and decoder in the existing target re-identification technology have unsatisfactory forward speed in the NPU onboard; the technical solution provided by the present invention has a good forward speed in the NPU onboard while ensuring the target re-identification accuracy performance.
[0009] In order to achieve the above object, the present invention adopts the following technical solutions:
[0010] A first aspect of the present invention provides a method for target re-identification, comprising the following steps:
[0011] Input the target image for re-identification and the target image in the candidate library into the pre-trained naive block re-identification model for feature extraction to obtain the target image feature vector and the candidate library target image feature vector;
[0012] Calculate the similarity and sort the target image feature vector and the candidate library target image feature vector to obtain the target re-identification result;
[0013] Among them, the naive block re-identification model includes an encoder and a decoder;
[0014] The encoder adopts the first goggle network or the second goggle network. The overall structures of the first goggle network and the second goggle network are as follows. The goggle module is used to convert shallow picture texture information into deep semantic information; a in the table represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 2. For the first goggle network, a is 7, and for the second goggle network, a is 3; b represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 3. For the first goggle network, b is 13, and for the second goggle network, b is 6; c represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 4. For both the first goggle network and the second goggle network, c is 2.
[0015] The decoder adopts a dual-mode network; the dual-mode network includes:
[0016] A replication unit, which is used to input the feature map output from the decoder and replicate it, and output three replicated feature maps.
[0017] A splitting unit, which is used to input the third feature map output from the replication unit and perform horizontal splitting processing, and output two bisected feature maps after splitting.
[0018] A global maximum pooling unit, which is used to input the first feature map output from the replication unit and the two bisected feature maps output from the splitting unit, and perform global maximum pooling processing, and output the feature vectors after pooling for each feature map.
[0019] A global average pooling unit, which is used to input the first and second feature maps output from the replication unit, perform global average pooling processing, and output the feature vectors after pooling for each feature map.
[0020] A 1x1 convolution unit, which is used to input the pooled feature vectors output from the global maximum pooling unit and the global average pooling unit and perform convolution dimensionality reduction processing, and output the feature vectors after dimensionality reduction for each feature vector.
[0021] A further improvement of the method of the present invention lies in that in the goggle network, in stage 4 and other goggle modules except the first goggle module in stage 2 and stage 3, no spatial downsampling is performed; when the number of channels of the feature map input to the goggle module is N, the processing steps inside the goggle module without accompanying spatial downsampling for the feature map include:
[0022] Step 1: The input feature map is evenly divided into two groups in the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 4 of the original number of channels, the convolution stride is 1, and the number of feature map channels in each group becomes N / 8;
[0023] Step 2: Concatenate the feature maps with N / 8 channels from the upper and lower branches in the channel dimension to form a feature map with N / 4 channels; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches;
[0024] Step 3: Evenly divide the feature map with N / 4 channels obtained in Step 2 into two parts in the channel dimension to obtain two groups of features with N / 8 channels; perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels on each convolution in the two branches to obtain two groups of feature maps with N / 2 channels;
[0025] Step 4: Concatenate the feature maps of the two branches in the channel dimension to form a feature map with N channels;
[0026] Step 5: Add the original feature map input in Step 1 to the feature map obtained in Step 4 to obtain the final feature map.
[0027] A further improvement of the method of the present invention is that in the swimming goggles network, the first swimming goggles module in each of Stage 2 and Stage 3 is used for spatial downsampling; when the number of channels of the feature map input to the swimming goggles module is N, the processing steps of the feature map inside the swimming goggles module accompanied by spatial downsampling include:
[0028] (1) Evenly divide the input feature map into two groups in the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 2 of the original number of channels, the convolution stride is 2, and the number of feature map channels in each group becomes N / 4;
[0029] (2) Concatenate the feature maps with N / 4 channels from the upper and lower branches in the channel dimension to form a feature map with N / 2 channels; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches;
[0030] (3) Evenly divide the feature map with N / 2 channels obtained in step (2) into two parts in the channel dimension to obtain two groups of features with N / 4 channels; perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels on each convolution in the two branches to obtain two groups of feature maps with N channels;
[0031] (4) Concatenate the feature maps of the two branches in the channel dimension to form a feature map with 2N channels;
[0032] (5) In the residual branch, an average pooling layer with a stride of 2 is used for spatial downsampling, and then a 1x1 convolution is used to adjust the number of channels to 2N to obtain an adjusted feature map;
[0033] (6) Add the feature maps obtained in steps (4) and (5) to obtain the final feature map.
[0034] A further improvement of the method of the present invention is that in the step of calculating the similarity based on the target picture feature vector and the candidate library target picture feature vector and sorting to obtain the target re-identification result, the metric used for similarity calculation is the Euclidean distance.
[0035] A target re-identification system provided in the second aspect of the present invention includes:
[0036] A feature vector acquisition module, configured to input the target picture for re-identification and the candidate library target picture into a pre-trained naive block re-identification model for feature extraction to obtain a target picture feature vector and a candidate library target picture feature vector;
[0037] A target re-identification result acquisition module, configured to calculate the similarity based on the target picture feature vector and the candidate library target picture feature vector and sort to obtain the target re-identification result;
[0038] Wherein, the naive block re-identification model includes an encoder and a decoder;
[0039] The encoder adopts the first goggle network or the second goggle network. The overall structures of the first goggle network and the second goggle network are The goggle module is used to convert the shallow picture texture information into deep semantic information; a in the table represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 2. For the first goggle network, a is 7, and for the second goggle network, a is 3; b represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 3. For the first goggle network, b is 13, and for the second goggle network, b is 6; c represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 4. In both the first goggle network and the second goggle network, c is 2;
[0040] The decoder adopts a dual-mode network; the dual-mode network includes:
[0041] A replication unit, configured to input the feature map output from the decoder and replicate it, and output three replicated feature maps;
[0042] A splitting unit, configured to input the third feature map output by the replication unit and perform horizontal splitting processing, and output two bisected feature maps after splitting;
[0043] A global max pooling unit, configured to input the first feature map output by the replication unit and the two bisected feature maps output by the splitting unit, and perform global max pooling processing, and output the feature vectors after pooling for each feature map;
[0044] A global average pooling unit, configured to input the first feature map and the second feature map output by the replication unit, perform global average pooling processing, and output the feature vectors after pooling for each feature map;
[0045] A 1x1 convolution unit, configured to input the pooled feature vectors output by the global max pooling unit and the global average pooling unit and perform convolution dimensionality reduction processing, and output the feature vectors after dimensionality reduction for each feature vector.
[0046] A further improvement of the system of the present invention lies in that in the goggle network, in stage 4 and other goggle modules in stages 2 and 3 except for the first goggle module, no spatial downsampling is performed; when the number of channels of the feature map input to the goggle module is N, the processing steps of the feature map inside the goggle module without accompanying spatial downsampling include:
[0047] Step 1, divide the input feature map into 2 groups on average from the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 4 of the original number of channels, the convolution stride is 1, and the number of channels of the feature map in each group becomes N / 8;
[0048] Step 2, splice the feature maps with N / 8 channels in the upper and lower branches into a feature map with N / 4 channels from the channel dimension; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches;
[0049] Step 3, divide the feature map with N / 4 channels obtained in Step 2 into two parts on average from the channel dimension to obtain two groups of features with N / 8 channels; perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels on each convolution in the two branches to obtain two groups of feature maps with N / 2 channels;
[0050] Step 4, splice the feature maps of the two branches into a feature map with N channels from the channel dimension;
[0051] Step 5, add the original feature map input in Step 1 to the feature map obtained in Step 4 to obtain the final feature map.
[0052] A further improvement of the system of the present invention lies in that, in the goggle network, the first goggle module in each of stage 2 and stage 3 is used for spatial downsampling; when the number of channels of the feature map input to the goggle module is N, the processing steps of the feature map inside the goggle module accompanied by spatial downsampling include:
[0053] (1) Divide the input feature map into two groups on average from the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 2 of the original number of channels, the convolution stride is 2, and the number of channels of the feature map in each group becomes N / 4;
[0054] (2) Concatenate the feature maps with N / 4 channels in the upper and lower branches from the channel dimension to form a feature map with N / 2 channels; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches;
[0055] (3) Divide the feature map with N / 2 channels obtained in step (2) into two parts on average from the channel dimension to obtain two groups of features with N / 4 channels; perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels on each convolution in the two branches to obtain two groups of feature maps with N channels;
[0056] (4) Concatenate the feature maps of the two branches from the channel dimension to form a feature map with 2N channels;
[0057] (5) In the residual branch, use an average pooling layer with a stride of 2 for spatial downsampling, and then adjust the number of channels to 2N through a 1x1 convolution to obtain an adjusted feature map;
[0058] (6) Add the feature maps obtained in step (4) and step (5) to obtain the final feature map.
[0059] A further improvement of the system of the present invention lies in that, in the step of calculating the similarity based on the target picture feature vector and the candidate library target picture feature vector and sorting to obtain the target re-identification result, the metric method used for similarity calculation is the Euclidean distance.
[0060] An electronic device provided in the third aspect of the present invention includes:
[0061] At least one processor; and,
[0062] A memory communicatively connected to the at least one processor; wherein,
[0063] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the target re-identification method as described in any one of the above of the present invention.
[0064] A computer-readable storage medium provided by the fourth aspect of the present invention stores a computer program, and when the computer program is executed by a processor, the target re-identification method described above in any one of the present invention is implemented.
[0065] Compared with the prior art, the present invention has the following beneficial effects:
[0066] The method provided by the present invention can solve the problem that the forward speed of the encoder and decoder in the existing target re-identification technology is not ideal in the NPU board; the technical solution provided by the present invention has a good forward speed in the NPU board while ensuring the accuracy performance of target re-identification. Specifically, for the mobile hardware deployment scenario, the present invention designs a re-ID algorithm suitable for NPU deployment; in the target re-identification method provided by the present invention, a pre-trained naive block re-identification model is used for identification and the result is obtained; the encoder of the naive block re-identification model uses a goggle network, and the decoder uses a dual-mode block network; the goggle network, through the reasonable combination of 3x3 convolution and channel splitting, while being able to fully utilize the on-board acceleration operator to accelerate the convolution operation, controls the amount of calculation and the model size within an ideal range, and at the same time can provide excellent feature representation ability for the system; the dual-mode block network can benefit from its design concept of dual-mode feature extraction, so as to obtain a feature representation ability stronger than the existing model with a relatively simple network structure. Description of the Drawings
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art; obviously, the drawings in the following description are some embodiments of the present invention, and for those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0068] Figure 1 It is a schematic flowchart of a target re-identification method provided by an embodiment of the present invention;
[0069] Figure 2 It is a schematic flowchart of the operation of the goggle module on the feature map proposed by an embodiment of the present invention;
[0070] Figure 3 It is a schematic flowchart of the operation of the dual-mode network on the feature map proposed by an embodiment of the present invention;
[0071] Figure 4 It is a schematic diagram of a target re-identification system provided by an embodiment of the present invention. Detailed Embodiments
[0072] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0073] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0074] The present invention will be further described in detail below in conjunction with the accompanying drawings:
[0075] Please refer to Figure 1 , a target re-identification method provided by an embodiment of the present invention includes the following steps:
[0076] Step 1, input the target image for re-identification and the candidate library target images into a pre-trained naive block re-identification model for feature extraction to obtain the target image feature vector and the candidate library target image feature vector;
[0077] Step 2, calculate the similarity based on the target image feature vector and the candidate library target image feature vector and sort to obtain the target re-identification result.
[0078] In the embodiment of the present invention, the naive block re-identification model includes an encoder and a decoder.
[0079] The encoder adopts the first goggle network or the second goggle network, and the overall structures of the first goggle network and the second goggle network are shown in Table 1.
[0080] Table 1. Overall structures of the first goggle network and the second goggle network
[0081]
[0082] In a further specific and explanatory embodiment of the present invention, the goggle module is used to convert shallow picture texture information into deep semantic information; in the table, a represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 2. For the first goggle network, a is 7, and for the second goggle network, a is 3; b represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 3. For the first goggle network, b is 13, and for the second goggle network, b is 6; c represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 4. For both the first goggle network and the second goggle network, c is 2.
[0083] In a further specific and explanatory embodiment of the present invention, stages 1 to 4 of the goggle network are all composed of several basic modules, namely "goggle modules". The processing process of the goggle module for features is as Figure 2 shown. Assume that the number of channels of the feature map input to the goggle module is N;
[0084] The operations experienced by the feature map inside the goggle module without accompanying spatial downsampling are as follows:
[0085] S1.4.1: The input feature map is evenly divided into 2 groups from the channel dimension, and the number of feature channels in each group is N / 2;
[0086] S1.4.2: For each group of feature maps, perform a 3x3 convolution once. The number of convolution kernels is 1 / 4 of the original number of channels, the convolution stride is 1, and the number of channels of the feature map in each group becomes N / 8;
[0087] S1.4.3: Concatenate the feature maps with N / 8 channels in the upper and lower branches from the channel dimension to form a feature map with N / 4 channels;
[0088] S1.4.4: Perform a 3x3 convolution with the number of convolution kernels equal to the number of channels for each convolution in the two branches;
[0089] S1.4.5: Evenly divide the feature map with N / 4 channels obtained in S1.4.4 into two parts from the channel dimension to obtain two groups of features with N / 8 channels again;
[0090] S1.4.6: Perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels for each convolution in the two branches to obtain two groups of feature maps with N / 2 channels;
[0091] S1.4.7: Concatenate the feature maps of the two branches from the channel dimension to form a feature map with N channels;
[0092] S1.4.8: Finally, add the features of the residual branch (before S1.4.1) to the existing features (after S1.4.7) to obtain the final feature map.
[0093] In an embodiment of the present invention, for the goggle module accompanying spatial downsampling, the operations experienced by the feature map inside the goggle module are as follows:
[0094] S1.4.1: The input feature map is evenly divided into 2 groups from the channel dimension, and the number of feature channels in each group is N / 2;
[0095] S1.4.2: For each group of feature maps, perform a 3x3 convolution once. The number of convolution kernels is 1 / 2 of the original number of channels, and the stride of the convolution is 2. In this way, the number of channels of the feature maps in each group becomes N / 4, and at the same time, 2-fold spatial downsampling is performed;
[0096] S1.4.3: Concatenate the feature maps with N / 4 channels in the upper and lower branches from the channel dimension to form a feature map with N / 2 channels;
[0097] S1.4.4: Perform a 3x3 convolution with the number of convolution kernels equal to the number of channels for each convolution in the two branches;
[0098] S1.4.5: Evenly divide the feature map with N / 2 channels obtained in S1.4.4 into two parts from the channel dimension to obtain two groups of features with N / 4 channels again;
[0099] S1.4.6: Perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels for each convolution in the two branches to obtain two groups of feature maps with N channels;
[0100] S1.4.7: Concatenate the feature maps of the two branches from the channel dimension to form a feature map with 2N channels;
[0101] S1.4.8: In the residual branch, use an additional average pooling layer with a stride (referring to the stride in the convolution operation) of 2 for spatial downsampling, followed by a 1x1 convolution to adjust to 2N channels;
[0102] S1.4.9: Finally, add the features of the residual branch (after S1.4.8) to the original features (after S1.4.7).
[0103] In an embodiment of the present invention. The decoder adopts a dual-mode network; the dual-mode network includes:
[0104] A replication unit for inputting the feature map output from the goggle network and replicating the feature map, and outputting three replicated feature maps;
[0105] A splitting unit for inputting the third feature map output from the replication unit and performing horizontal splitting processing, and outputting two bisected feature maps after splitting;
[0106] A global maximum pooling unit, which is used to input the first feature map copied by the input copy unit and the two bisected feature maps obtained by the splitting unit, perform global maximum pooling processing, and output the feature vectors after pooling for each feature map;
[0107] A global average pooling unit, which is used to input the first feature map copied by the copy unit and the second feature map copied by the copy unit, perform global average pooling processing, and output the feature vectors after pooling for each feature map;
[0108] A 1x1 convolution unit, which is used to input the five feature vectors obtained after pooling and perform convolution dimensionality reduction processing, and output the feature vectors after dimensionality reduction for each feature vector.
[0109] Please refer to Figure 3 , the feature map in the embodiment of the present invention specifically undergoes the following process inside the dual-mode network:
[0110] S1.5.1: Copy the feature map twice to obtain a total of 3 identical feature maps;
[0111] S1.5.2: Perform global maximum pooling on the first feature map, and use 1x1 convolution to reduce the dimensionality of the obtained feature vector with 1024 channels to make its number of channels 256;
[0112] S1.5.3: Perform global average pooling on the first feature map, and use 1x1 convolution to reduce the dimensionality of the obtained feature vector with 1024 channels to make its number of channels 256;
[0113] S1.5.4: Perform global average pooling on the second feature map, and use 1x1 convolution to reduce the dimensionality of the obtained feature vector with 1024 channels to make its number of channels 256;
[0114] S1.5.5: Divide the third feature map into two equal parts in the horizontal direction, perform global maximum pooling on the two split feature maps respectively, and use 1x1 convolution to reduce the dimensionality of the obtained feature vector with 1024 channels to make its number of channels 256.
[0115] In the embodiment of the present invention, the metric used for similarity calculation is the Euclidean distance:
[0116] Eu(feat1, feat2) = feat1 · feat2 / |feat1||feat2|;
[0117] In the formula: feat1 represents the first feature vector participating in the Euclidean distance calculation, and feat2 represents the other feature vector participating in the Euclidean distance calculation.
[0118] The technical solution provided by the embodiments of the present invention, through the reasonable combination of 3x3 convolution and channel splitting, while being able to make full use of the on-board acceleration operator to accelerate the convolution operation, controls the computational complexity and model size within an ideal range, and at the same time can provide excellent feature representation ability for the system; the core component of the goggle network, the goggle module, can endow the model with powerful feature representation ability while enabling the model to run at high speed on the NPU; compared with existing similar models, the dual-mode network benefits from its design concept of dual-mode feature extraction and can obtain stronger feature representation ability than existing models with a relatively simple network structure. The embodiments of the present invention explain the naive block re-identification model. The block model is a type of method in the re-identification method that divides the feature map to obtain an improvement in feature learning ability and achieve performance gain, while the naive block model refers to the method that only performs the first-layer division of the feature map without setting the second-layer or third-layer division of the feature map.
[0119] An efficient naive block re-identification method for mobile devices provided by the embodiments of the present invention specifically includes the following steps:
[0120] S1: Divide the training set of the Market-1501 dataset into several batches with 32 images in each batch, and sequentially input them into the efficient naive block model for mobile devices;
[0121] S2: For each batch, uniformly adjust the images in this batch to a resolution of 384*128;
[0122] S3: Perform data augmentation on this batch using horizontal random flipping with a probability of 50% and random erasing with a probability of 40% and an area ratio of 40%;
[0123] S4: Input the images in this batch into the encoder goggle network;
[0124] S5: Input the feature map output by the goggle network into the decoder dual-mode network;
[0125] S6: Sequentially splice the feature vectors output by the dual-mode network as the feature description vector of the image;
[0126] S7: Use the loss function to train the entire model;
[0127] S8: Use the trained model to extract features from the query set and candidate set images to obtain feature description vectors;
[0128] S9: Calculate the similarity between the query set and candidate set images using the feature description vectors, and output the query results after sorting.
[0129] In the goggle network in S4 of the embodiments of the present invention, the processing process of its core module for features is as Figure 2As shown, assuming that the number of channels of the feature map input to the goggle module is N, the operations experienced by the feature map inside the goggle module without accompanying spatial downsampling are as follows:
[0130] S4.1: The input feature map is evenly divided into 2 groups from the channel dimension, and the number of feature channels in each group is N / 2;
[0131] S4.2: For each group of feature maps, a 3x3 convolution is performed once, and the number of convolution kernels is 1 / 4 of the original number of channels. In this way, the number of channels of the feature maps in each group becomes N / 8;
[0132] S4.3: The feature maps with N / 8 channels in the upper and lower branches are concatenated from the channel dimension to form a feature map with N / 4 channels;
[0133] S4.4: For each convolution in the two branches, a 3x3 convolution with the number of convolution kernels equal to the number of channels is performed;
[0134] S4.5: The feature map with N / 4 channels obtained in S4.4 is evenly divided into two parts from the channel dimension to obtain two groups of features with N / 8 channels again;
[0135] S4.6: For each convolution in the two branches, a 3x3 convolution with the number of convolution kernels 4 times the number of channels is performed to obtain two groups of feature maps with N / 2 channels;
[0136] S4.7: The feature maps of the two branches are concatenated from the channel dimension to form a feature map with N channels;
[0137] S4.8: Finally, the features of the residual branch (before S4.1) are added to the existing features (after S4.7).
[0138] For the goggle module with accompanying spatial downsampling, the differences in its internal operations from S4.1 - S4.8 above are as follows:
[0139] (1) In the residual branch, an additional average pooling layer with a stride (referring to the stride in the convolution operation) of 2 is used for spatial downsampling, followed by a 1x1 convolution to adjust to the target channel dimension;
[0140] (2) In the main path, the stride of the first 3x3 convolution is adjusted from 1 to 2, and the number of convolution kernels is adjusted from N / 4 to N / 2 to achieve the purpose of spatial downsampling and channel upsampling at the same time.
[0141] The overall structure of the goggle network in S4 is shown in Table 1. After a pedestrian picture with a size of 384 * 128 passes through this network, a feature map with a size of 24 * 8 * 1024 is obtained.
[0142] The goggle network in S4 needs to be pre-trained on ImageNet before S7, with a batch size of 512. The stochastic gradient descent algorithm with a momentum parameter of 0.9 is used, and the model is optimized with a weight decay of 0.0001. The Dropout ratio before the classification layer is 0.2. In the first 5 epochs, the warm-up strategy is used to gradually change the learning rate from 0.001 to 0.1, and in the remaining 295 epochs, the cosine annealing learning rate decay strategy is used to gradually change the learning rate from 0.1 to 0.
[0143] The dual-mode network in S5 of the embodiment of the present invention has a schematic structural diagram as Figure 3 shown, and the feature map undergoes the following process inside:
[0144] S5.1: Duplicate the feature map twice to obtain a total of 3 identical feature maps;
[0145] S5.2: Perform global max pooling on the first feature map, and use a 1x1 convolution to reduce the dimension of the obtained feature vector so that its number of channels is 256;
[0146] S5.3: Perform global average pooling on the first feature map, and use a 1x1 convolution to reduce the dimension of the obtained feature vector so that its number of channels is 256;
[0147] S5.4: Perform global average pooling on the second feature map, and use a 1x1 convolution to reduce the dimension of the obtained feature vector so that its number of channels is 256;
[0148] S5.5: Divide the third feature map into two equal parts horizontally, perform global max pooling on the two divided feature maps respectively, and use a 1x1 convolution to reduce the dimension of the obtained feature vector so that its number of channels is 256.
[0149] The loss function training process described in step S7 of the embodiment of the present invention includes the following steps:
[0150] S7.1: For a batch of training sets currently participating in training, assume that this batch contains pedestrians, and each pedestrian has sample images; for an image where i ∈ {1, 2, …, 64}, l is the pedestrian identity corresponding to this image, the corresponding feature description vector is then all with f = i in this batch are positive samples all with e = l are negative samples For each in this batch, use the hard sample criterion to find its positive and negative samples, and the objective function for hard sample mining should satisfy:
[0151]
[0152] Among them, Eu(·) is the calculation of Euclidean distance (cosine): Eu(feat1, feat2) = feat1 · feat2 / |feat1||feat2|;
[0153] S7.2: Add a fully connected layer after the feature descriptor and the output is v = [v1,..., v I , where I is the number of pedestrian identity categories in the training dataset. Then the pedestrian identity prediction objective function should satisfy:
[0154]
[0155] Among them, Let y l be the true pedestrian identity of this picture. Then q(y l ) = 1, and when n ≠ y l , q(n) = 0;
[0156] S7.3: Optimize the overall objective function Calculate the overall objective function for all training set pictures in batches, train the network weights, and repeat the training set 75 times;
[0157] S7.4: Use the stochastic gradient descent algorithm with a momentum parameter of 0.9 for model optimization. In the first 5 iterations, use the warm-up strategy to gradually change the learning rate from 0.001 to 0.1, and in the remaining 70 cycles, use the cosine learning rate decay strategy to gradually change the learning rate from 0.1 to 0.
[0158] After the embodiments of the present invention complete S9, CMC and mAP evaluations are used, and the forward speed of the model on the CPU, GPU, and NPU of the development board is tested on Rockchip RK3399Pro. The obtained results are compared with those of current other re-identification methods as shown in Tables 2 and 3. Among them, the method using ResNet as the encoder is a non-lightweight method, and AutoREID, CDNet, OSNet, and DSLNet are lightweight re-identification methods.
[0159] Table 2. Comparison of the method of the embodiments of the present invention with current other advanced methods on the pedestrian dataset
[0160]
[0161] The comparison datasets in Table 2 are the pedestrian re-identification datasets Market-1501, DukeMTMC-reID, and MSMT17. R-1 and mAP are re-identification accuracy metrics; Params is the parameter quantity metric; FLOPs is the computational volume metric; FPS is the forward frame rate of the model on a certain hardware (CPU, GPU, or NPU), which is the on-board speed metric of the model.
[0162] Table 3. Comparison of the method of the embodiment of the present invention with other current advanced methods on the vehicle dataset
[0163]
[0164] The comparison datasets in Table 3 are the vehicle re-identification datasets VeRi-776 and VehicleID. R-1 and mAP are re-identification accuracy metrics; Params is the parameter quantity metric; FLOPs is the computational volume metric; FPS is the forward frame rate of the model on a certain hardware (CPU, GPU, or NPU), which is the on-board speed metric of the model.
[0165] Based on Table 2, it can be seen from the comparison table of the pedestrian dataset that the method proposed by the present invention has a significant lead in the on-board speed when the mAP and Rank-1 metrics are superior compared to the current state-of-the-art lightweight methods OSNet and DSLNet. Among them, the frame rate of the dual-mode network combined with the second goggle network on the NPU is more than 7 times that of OSNet; compared with the non-lightweight re-identification method using ResNet as the encoder, when the mAP and Rank-1 metrics of the dual-mode network combined with the first goggle network reach the same level as such methods, the parameter quantity is 24.6% of such methods, the computational volume is 17.8% of it, the CPU frame rate is 525% of it, and the GPU frame rate is 488% of it. Moreover, ResNet50 cannot be successfully deployed on the NPU due to its excessive parameter quantity, while the dual-mode network combined with the first goggle network can reach a rate of more than 200 frames per second on the NPU.
[0166] Based on Table 3, it can be seen from the comparison table of the vehicle dataset that the dual-mode network combined with the first goggle network has reached the same level as other methods using ResNet50 as the encoder in terms of mAP and Rank-1 metrics, while the parameter quantity is 24.6% of such methods, the computational volume is 17.8% of it, the CPU frame rate is 551% of it, and the GPU frame rate is 578% of it. ResNet50 cannot be successfully deployed on the NPU due to its excessive parameter quantity, while the dual-mode network combined with the first goggle network can reach a rate of more than 180 frames per second on the NPU.
[0167] In summary, in view of the problem that the forward speed of the current person re-identification method on the NPU is slow, the present invention proposes an efficient naive block-based re-identification method for mobile devices, including: pre-training the encoder goggle network using ImageNet; further training the encoder goggle network and the decoder dual-mode network using person images; inputting the person image into the goggle network to obtain a feature map; inputting the feature map output by the goggle network into the dual-mode network to obtain a feature description vector; and using a loss function to guide the training of the two models. The re-identification method proposed by the present invention, through reasonable design of the encoder and decoder, while making full use of the on-board acceleration operator to accelerate the convolution operation, controls the computational amount and model size within an ideal range, and at the same time can provide excellent feature representation ability for the system. The re-identification method provided by the embodiments of the present invention can achieve the same level of re-identification accuracy with fewer parameters and computational amounts compared with the existing non-lightweight re-identification methods; compared with the lightweight re-identification methods, the method provided by the present disclosure examples not only has higher re-identification accuracy, but also is significantly ahead in terms of the on-board speeds of the CPU, GPU, and NPU.
[0168] The following is an apparatus embodiment of the present invention, which can be used to execute the method embodiment of the present invention. For details not disclosed in the apparatus embodiment, please refer to the method embodiment of the present invention.
[0169] Please refer to Figure 4 , in another embodiment of the present invention, a target re-identification system is provided, including:
[0170] A feature vector acquisition module, configured to input a target image for re-identification and a candidate library target image into a pre-trained naive block-based re-identification model for feature extraction, to obtain a target image feature vector and a candidate library target image feature vector;
[0171] A target re-identification result acquisition module, configured to calculate similarities and sort based on the target image feature vector and the candidate library target image feature vector, to obtain a target re-identification result;
[0172] Wherein, the naive block-based re-identification model includes an encoder and a decoder; the encoder adopts a first goggle network or a second goggle network, and the decoder adopts a dual-mode network.
[0173] In another embodiment of the present invention, a computer device is provided. The computer device includes a processor and a memory. The memory is used to store a computer program, and the computer program includes program instructions. The processor is used to execute the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in the computer storage medium to implement the corresponding method flow or corresponding function. The processor described in the embodiments of the present invention can be used for the operations of the target re-identification method.
[0174] In another embodiment of the present invention, a storage medium is also provided, specifically a computer-readable storage medium (Memory). The computer-readable storage medium is a memory device in the computer device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and, of course, the extended storage medium supported by the computer device. The computer-readable storage medium provides a storage space, and the operating system of the terminal is stored in this storage space. Moreover, one or more instructions suitable for being loaded and executed by the processor are stored in this storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. One or more instructions stored in the computer-readable storage medium can be loaded and executed by the processor to implement the corresponding steps of the target re-identification method in the above embodiments.
[0175] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0176] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0177] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0178] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Therefore, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: still can modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A target re-identification method, characterized in that, It includes the following steps: Input the target image for re-identification and the candidate library target images into a pre-trained naive block re-identification model for feature extraction to obtain the target image feature vector and the candidate library target image feature vectors; Calculate the similarity based on the target image feature vector and the candidate library target image feature vectors and sort them to obtain the target re-identification result; Among them, the naive block re-identification model includes an encoder and a decoder; The encoder adopts the first goggle network or the second goggle network. The overall structures of the first goggle network and the second goggle network are as follows: The goggle module is used to convert the shallow image texture information into deep semantic information; in the table, a represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 2. For the first goggle network, a is 7, and for the second goggle network, a is 3; b represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 3. For the first goggle network, b is 13, and for the second goggle network, b is 6; c represents the number of times the goggle module is stacked and repeated after the first goggle module in stage 4. For both the first goggle network and the second goggle network, c is 2; The decoder adopts a dual-mode network; the dual-mode network includes: A replication unit, which is used to input the feature map output from the decoder and replicate it, and output three replicated feature maps; A splitting unit, which is used to input the third feature map output by the replication unit and perform horizontal splitting processing, and output two bisected feature maps after splitting; A global max pooling unit, which is used to input the first feature map output by the replication unit and the two bisected feature maps output by the splitting unit, and perform global max pooling processing, and output the feature vectors after pooling for each feature map; A global average pooling unit, which is used to input the first and second feature maps output by the replication unit, perform global average pooling processing, and output the feature vectors after pooling for each feature map; A 1x1 convolution unit, which is used to input the pooled feature vectors output by the global max pooling unit and the global average pooling unit and perform convolution dimensionality reduction processing, and output the feature vectors after dimensionality reduction for each feature vector.
2. The object re-identification method according to claim 1, wherein In the goggle network, in stage 4 and in stage 2 and stage 3, except for the first goggle module, other goggle modules do not perform spatial downsampling; when the number of channels of the feature map input to the goggle module is N, the processing steps inside the goggle module without accompanying spatial downsampling for the feature map include: Step 1, evenly divide the input feature map into 2 groups from the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 4 of the original number of channels, the convolution stride is 1, and the number of channels of the feature map in each group becomes N / 8; Step 2, splice the feature maps with the number of channels N / 8 in the upper and lower branches from the channel dimension into a feature map with the number of channels N / 4; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches; Step 3: Average the feature map with N / 4 channels obtained in Step 2 into two parts along the channel dimension to obtain two sets of features with N / 8 channels; perform a 3x3 convolution with the number of convolution kernels being 4 times the number of channels on each convolution in the two branches to obtain two sets of feature maps with N / 2 channels; Step 4: Concatenate the feature maps of the two branches along the channel dimension to form a feature map with N channels; Step 5: Add the original feature map input in Step 1 to the feature map obtained in Step 4 to obtain the final feature map.
3. A target re-identification method according to claim 1, characterized in that, In the goggle network, the first goggle module in each of Stage 2 and Stage 3 is used for spatial downsampling; when the number of channels of the feature map input to the goggle module is N, the processing steps of the feature map inside the goggle module accompanied by spatial downsampling include: (1) Average the input feature map into 2 groups along the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each obtained group of feature maps, the number of convolution kernels is 1 / 2 of the original number of channels, the stride of the convolution is 2, and the number of channels of the feature map in each group becomes N / 4; (2) Concatenate the feature maps with N / 4 channels in the upper and lower branches along the channel dimension to form a feature map with N / 2 channels; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels on each convolution in the two branches; (3) Average the feature map with N / 2 channels obtained in step (2) into two parts along the channel dimension to obtain two sets of features with N / 4 channels; perform a 3x3 convolution with the number of convolution kernels being 4 times the number of channels on each convolution in the two branches to obtain two sets of feature maps with N channels; (4) Concatenate the feature maps of the two branches along the channel dimension to form a feature map with 2N channels; (5) In the residual branch, use an average pooling layer with a stride of 2 for spatial downsampling, and then adjust the number of channels to 2N through a 1x1 convolution to obtain the adjusted feature map; (6) Add the feature maps obtained in step (4) and step (5) to obtain the final feature map.
4. The object re-identification method according to claim 1, characterized in that In the step of calculating the similarity based on the target picture feature vector and the candidate library target picture feature vector and sorting to obtain the target re-identification result, the metric used for similarity calculation is the Euclidean distance.
5. A target re-identification system, characterized in that, Including: A feature vector acquisition module, configured to input the target picture for re-identification and the candidate library target picture into a pre-trained naive block re-identification model for feature extraction to obtain the target picture feature vector and the candidate library target picture feature vector; A target re-identification result acquisition module, configured to calculate the similarity based on the target picture feature vector and the candidate library target picture feature vector and sort to obtain the target re-identification result; Among them, the naive block re-identification model includes an encoder and a decoder; The encoder adopts the first goggle network or the second goggle network, and the overall structures of the first goggle network and the second goggle network are The goggle module is used to convert shallow picture texture information into deep semantic information; in the table, a represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 2, where a is 7 for the first goggle network and a is 3 for the second goggle network; b represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 3, where b is 13 for the first goggle network and b is 6 for the second goggle network; c represents the number of times the goggle modules are stacked and repeated after the first goggle module in stage 4, and c is 2 for both the first goggle network and the second goggle network. The decoder adopts a dual-mode network; the dual-mode network includes: A replication unit, which is used to input the feature map output from the decoder, replicate it, and output three replicated feature maps. A splitting unit, which is used to input the third feature map output by the replication unit and perform horizontal splitting processing, and output two bisected feature maps after splitting. A global max pooling unit, which is used to input the first feature map output by the replication unit and the two bisected feature maps output by the splitting unit, and perform global max pooling processing to output the feature vectors after pooling for each feature map. A global average pooling unit, which is used to input the first and second feature maps output by the replication unit, perform global average pooling processing, and output the feature vectors after pooling for each feature map. A 1x1 convolution unit, which is used to input the pooled feature vectors output by the global max pooling unit and the global average pooling unit, perform convolution dimensionality reduction processing, and output the feature vectors after dimensionality reduction for each feature vector.
6. The object re-identification system according to claim 5, characterized in that, In the goggle network, in stage 4 and in stage 2 and stage 3, except for the first goggle module, other goggle modules do not perform spatial downsampling; when the number of channels of the feature map input to the goggle module is N, the processing steps of the feature map inside the goggle module without accompanying spatial downsampling include: Step 1: Divide the input feature map into two groups on average from the channel dimension, and the number of feature channels in each group is N / 2; perform a 3x3 convolution on each group of obtained feature maps, the number of convolution kernels is 1 / 4 of the original number of channels, the convolution stride is 1, and the number of channels of the feature map in each group becomes N / 8. Step 2: Concatenate the feature maps with N / 8 channels in the upper and lower branches from the channel dimension to form a feature map with N / 4 channels; perform a 3x3 convolution with the number of convolution kernels equal to the number of channels for each convolution in the two branches. Step 3: Divide the feature map with N / 4 channels obtained in Step 2 into two parts on average from the channel dimension to obtain two groups of features with N / 8 channels; perform a 3x3 convolution with the number of convolution kernels 4 times the number of channels for each convolution in the two branches to obtain two groups of feature maps with N / 2 channels. Step 4: Concatenate the feature maps of the two branches from the channel dimension to form a feature map with N channels. Step 5: Add the original feature map input in Step 1 to the feature map obtained in Step 4 to obtain the final feature map.
7. The object re-identification system according to claim 5, characterized in that In the goggle network, the first goggle module in each of Stage 2 and Stage 3 is used for spatial downsampling. When the number of channels of the feature map input to the goggle module is N, the processing steps of the feature map inside the goggle module accompanying spatial downsampling include: (1) The input feature map is evenly divided into two groups from the channel dimension, and the number of feature channels in each group is N / 2; a 3x3 convolution is performed on each group of obtained feature maps, the number of convolution kernels is 1 / 2 of the original number of channels, the convolution stride is 2, and the number of channels of the feature map in each group becomes N / 4; (2) The feature maps with N / 4 channels in the upper and lower branches are concatenated from the channel dimension to form a feature map with N / 2 channels; a 3x3 convolution with the number of convolution kernels equal to the number of channels is performed on each convolution in the two branches; (3) The feature map with N / 2 channels obtained in step (2) is evenly divided into two parts from the channel dimension to obtain two groups of features with N / 4 channels; a 3x3 convolution with the number of convolution kernels 4 times the number of channels is performed on each convolution in the two branches to obtain two groups of feature maps with N channels; (4) The feature maps of the two branches are concatenated from the channel dimension to form a feature map with 2N channels; (5) In the residual branch, an average pooling layer with a stride of 2 is used for spatial downsampling, and then the number of channels is adjusted to 2N through a 1x1 convolution to obtain an adjusted feature map; (6) The feature maps obtained in step (4) and step (5) are added to obtain the final feature map.
8. An object re-identification system according to claim 5, characterized in that In the step of calculating the similarity based on the target picture feature vector and the candidate library target picture feature vector and sorting to obtain the target re-identification result, the metric used for similarity calculation is the Euclidean distance.
9. An electronic device, characterized in that, Including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the target re-identification method according to any one of claims 1 to 4.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the target re-identification method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Encoder training method and device and storage medium
CN114418069A
Pedestrian re-identification method and system based on comparison features
CN114429648A