Pedestrian re-identification method, system and device and medium
Through the super-score detection model, multiple super-score detections of pedestrian re-identification images are carried out, and combined with multiple network models, the problem of difficult and low efficiency of feature matching in pedestrian re-identification is solved, achieving a more efficient and stable recognition effect.
Patent Information
- Application Number
- CN202510381884.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing pedestrian re-identification method, the distance between the pedestrian target and the camera leads to blurred images after interpolation amplification, which is difficult to match features, and poor recognition efficiency and stability.
The super-segment detection model is used to perform super-segment detection at image level and detection frame level, combining the size detection network, image super-segment network and object detection network, and improving the accuracy and efficiency of feature matching through deep learning.
Effectively reduce the difficulty of pedestrian re-identification feature matching, improve identification efficiency and stability, reduce computing resource requirements, and reduce hardware costs.
Smart Images

Figure CN120356238A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular to a pedestrian re-identification method, system, device and medium. Background Art
[0002] The goal of pedestrian re-identification is to identify and match the identities of the same person from images taken by different cameras and from different perspectives, and its application in the fields of intelligent monitoring, security, intelligent transportation, etc. is becoming increasingly common.
[0003] Currently, traditional pedestrian re-identification methods usually interpolate and magnify the captured images and then extract discriminative features, and then match the extracted features with the known features in the database to achieve pedestrian re-identification. However, in practical applications, the pedestrian target is usually far away from the camera, and this method is likely to cause image blurring in the interpolated and magnified captured images, and the matching difficulty between the extracted features and the known features is relatively large, the matching takes a long time, and the recognition efficiency of pedestrian re-identification is not satisfactory.
[0004] Therefore, the problems existing in the prior art still need to be solved and optimized urgently. Summary of the Invention
[0005] An object of the present invention is to solve at least to a certain extent one of the technical problems existing in the related art.
[0006] To this end, an object of an embodiment of the present invention is to provide a pedestrian re-identification method, system, device and medium, which can effectively reduce the feature matching difficulty of pedestrian re-identification and improve the recognition efficiency of pedestrian re-identification.
[0007] In order to achieve the above technical object, the technical solutions adopted in the embodiments of the present application include:
[0008] In a first aspect, an embodiment of the present application provides a pedestrian re-identification method, including:
[0009] Obtain the target pedestrian feature and the captured image to be detected;
[0010] Input the captured image into a super-resolution detection model for image super-resolution detection to obtain a detection frame image output by the super-resolution detection model, where the detection frame image includes a plurality of pedestrian detection frames, and the pedestrian detection frame is a candidate frame for the pedestrian feature of a pedestrian image with an image size greater than a first size threshold;
[0011] Input the detection frame image into the super-resolution detection model for detection frame super-resolution detection to obtain pedestrian detection features, where the pedestrian detection features are the pedestrian features of the pedestrian detection frames with a detection frame size greater than a second size threshold;
[0012] Match and re-identify the pedestrian detection features according to the target pedestrian features to obtain the pedestrian re-identification result of the captured image.
[0013] In addition, according to the method of the above embodiments of the present application, the following additional technical features may also be included:
[0014] Further, in an embodiment of the present application, the super-resolution detection model includes a size detection network, an image super-resolution network, and a target detection network. Input the first target image into the super-resolution detection model for super-resolution detection to obtain an output target, including:
[0015] Input the first target image into the size detection network for size detection to obtain a second target image and the target size of the second target image;
[0016] Obtain a target threshold corresponding to the target size;
[0017] Compare the target threshold and the target size to obtain a size determination result;
[0018] If the size determination result is that the target size is greater than the target threshold, input the second target image into the target detection network for target detection to obtain the output target; or, if the size determination result is that the target size is less than or equal to the target threshold, input the second target image into the image super-resolution network to obtain a super-resolution image, and input the super-resolution image into the target detection network for target detection to obtain the output target;
[0019] Wherein, the first target image is the captured image or the detection frame image.
[0020] Further, in an embodiment of the present application, inputting the first target image into the size detection network for size detection to obtain a second target image includes:
[0021] Perform multi-level feature extraction on the first target image to obtain a plurality of first intermediate features, and the feature scales of each of the first intermediate features are different;
[0022] Perform upsampling feature fusion on all the first intermediate features to obtain a plurality of second intermediate features;
[0023] Classify and crop the first target image according to all the second intermediate features to obtain the second target image.
[0024] Further, in an embodiment of the present application, the image super-resolution network includes a plurality of cascaded pyramid residual feature extraction modules. Inputting the second target image into the image super-resolution network to obtain a super-resolved image includes:
[0025] Performing shallow feature extraction on the second target image to obtain shallow features;
[0026] Performing deep feature extraction on the shallow features through the plurality of cascaded pyramid residual feature extraction modules to obtain deep features;
[0027] Performing feature fusion and feature pixel rearrangement on the deep features according to the shallow features to obtain the super-resolved image.
[0028] Further, in an embodiment of the present application, performing deep feature extraction on the input features through the pyramid residual feature extraction module to obtain output features includes:
[0029] Performing multi-level feature convolution extraction on the input features to obtain third intermediate features;
[0030] Performing residual feature fusion on the input features according to the third intermediate features to obtain fourth intermediate features;
[0031] Performing pyramid compression attention extraction on the fourth intermediate features to obtain the output features;
[0032] Wherein, the input features are the shallow features or the output features output by the previous pyramid residual feature extraction module.
[0033] Further, in an embodiment of the present application, inputting a third target image into the target detection network for target detection to obtain the output target includes:
[0034] Performing depthwise separable convolution processing on the image features of the third target image to obtain fifth intermediate features;
[0035] Performing attention extraction on the fifth intermediate features to obtain the feature attention of the fifth intermediate features;
[0036] Performing attention fusion on the fifth intermediate features according to the feature attention to obtain sixth intermediate features;
[0037] Performing feature fusion on the image features of the second target image according to the sixth intermediate features to obtain the output target;
[0038] Wherein, the third target image is the second target image or the super-resolved image.
[0039] Further, in an embodiment of the present application, the step of performing matching re-identification on the pedestrian detection features according to the target pedestrian features to obtain the pedestrian re-identification result of the captured image includes:
[0040] Obtain a similarity distance threshold;
[0041] Calculate the distance between the pedestrian detection features according to the target pedestrian features to obtain a feature distance;
[0042] Compare the feature distance with the similarity distance threshold to obtain a threshold comparison result;
[0043] If the threshold comparison result is that the feature distance is less than the threshold comparison result, then determine the pedestrian detection features as the pedestrian re-identification result of the captured image.
[0044] In a second aspect, an embodiment of the present application provides a pedestrian re-identification system, including:
[0045] A first processing unit, configured to obtain target pedestrian features and a captured image to be detected;
[0046] A second processing unit, configured to input the captured image into a super-resolution detection model for image super-resolution detection to obtain a detection frame image output by the super-resolution detection model, where the detection frame image includes a plurality of pedestrian detection frames, and the pedestrian detection frames are candidate frames of pedestrian features of pedestrian images with an image size greater than a first size threshold;
[0047] A third processing unit, configured to input the detection frame image into the super-resolution detection model for detection frame super-resolution detection to obtain pedestrian detection features, where the pedestrian detection features are pedestrian features of pedestrian detection frames with a detection frame size greater than a second size threshold;
[0048] A fourth processing unit, configured to perform matching re-identification on the pedestrian detection features according to the target pedestrian features to obtain the pedestrian re-identification result of the captured image.
[0049] In a third aspect, an embodiment of the present application further provides an electronic device, including:
[0050] At least one processor;
[0051] At least one memory, configured to store at least one program;
[0052] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.
[0053] Fourthly, an embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by a processor. The program executable by the processor, when executed by the processor, is used to implement the above method.
[0054] The advantages and beneficial effects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application:
[0055] A person re-identification method, system, device and medium disclosed by an embodiment of the present application. The method includes: obtaining a target pedestrian feature and a captured image to be detected; inputting the captured image into a super-resolution detection model for image super-resolution detection to obtain a detection box image output by the super-resolution detection model, where the detection box image includes a plurality of pedestrian detection boxes, and the pedestrian detection box is a candidate box for the pedestrian feature of a pedestrian image with an image size greater than a first size threshold; inputting the detection box image into the super-resolution detection model for detection box super-resolution detection to obtain a pedestrian detection feature, where the pedestrian detection feature is the pedestrian feature of a pedestrian detection box with a detection box size greater than a second size threshold; and performing matching re-identification on the pedestrian detection feature according to the target pedestrian feature to obtain a person re-identification result of the captured image. By performing multiple super-resolution detections on the input image through the super-resolution detection model, specifically, performing image-level super-resolution detection on the captured image through the super-resolution detection model, and performing detection box-level super-resolution detection on the detection box image through the super-resolution detection model, the method can effectively reduce the feature matching difficulty of person re-identification and improve the identification efficiency of person re-identification. Description of the Drawings
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the related technical solution drawings in the embodiments of the present application or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly expressing some embodiments of the technical solutions in the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0057] Figure 1 It is a schematic flowchart of a person re-identification method provided by an embodiment of the present application;
[0058] Figure 2 It is a schematic network structure diagram of a size detection network provided by an embodiment of the present application;
[0059] Figure 3 It is a schematic structure diagram of a pyramid residual feature extraction module provided by an embodiment of the present application;
[0060] Figure 4Schematic structural framework diagram of a person re-identification system provided by an embodiment of the present application;
[0061] Figure 5 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Specific embodiments
[0062] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as limiting the present application. For the step numbers in the following embodiments, they are only set for the convenience of explanation and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0064] Currently, traditional person re-identification methods usually interpolate and magnify the captured image and then extract discriminative features, and then match the extracted features with the known features in the database to achieve person re-identification. However, in actual applications, the person target is usually at a relatively long distance from the camera. This method is likely to cause the interpolated and magnified captured image to be blurred, and the matching difficulty between the extracted features and the known features is relatively large, the matching takes a long time, and the recognition efficiency of person re-identification is not satisfactory; in addition, due to the relatively large matching difficulty between the extracted features and the known features, the matching results obtained by this method often deviate from the actual situation to a certain extent, resulting in poor stability and effect of person re-identification.
[0065] In view of this, the embodiments of the present invention provide a person re-identification method, system, device and medium. Among them, this method performs multiple super-resolution detections on the input image through a super-resolution detection model. Specifically, it performs super-resolution detection at the image level on the captured image through the super-resolution detection model, and performs super-resolution detection at the detection box level on the detection box image through the super-resolution detection model, which can effectively reduce the feature matching difficulty of person re-identification and improve the recognition efficiency of person re-identification.
[0066] In addition, the method performs pedestrian re-identification on the captured images through a super-resolution detection model. Specifically, the method performs pedestrian re-identification on the captured images by combining a size detection network, an image super-resolution network and a target detection network. It can utilize the generalization ability of deep learning to effectively improve the stability and effect of pedestrian re-identification. At the same time, the method performs pedestrian re-identification on the captured images through a super-resolution detection model. The super-resolution detection model can learn the complex nonlinear mapping between low-resolution images and high-resolution images in the pedestrian re-identification task, which is not only conducive to improving the flexibility and application scope of the pedestrian re-identification task, but also conducive to more fully capturing the feature information in the captured images, so that the obtained pedestrian detection features have more feature information, thereby effectively improving the stability and effect of pedestrian re-identification.
[0067] In addition, the method is also based on an image super-resolution network of a cascaded pyramid residual feature extraction model, and specifically implements super-resolution reconstruction of the input second target image through pyramid compression attention extraction. Compared with the existing super-resolution reconstruction technology based on the Enhanced Spatial Attention (ESA) mechanism, while ensuring the accuracy of image super-resolution reconstruction, it effectively reduces the computing resources required for super-resolution reconstruction and improves the reconstruction efficiency of super-resolution reconstruction, which is beneficial to improving the recognition efficiency of pedestrian re-identification and reducing the hardware equipment cost required for pedestrian re-identification tasks. In addition, in practical applications, super-resolution reconstruction of each image will cause computational redundancy, while this method only performs super-resolution reconstruction on the image when the resolution size of the pedestrian image or the detection frame image is small, which can further reduce the required computing resources, which is beneficial to improving the recognition efficiency of pedestrian re-identification.
[0068] Reference Figure 1 In an embodiment of the present application, a pedestrian re-identification method includes:
[0069] Step 110: Acquire target pedestrian features and captured images to be detected;
[0070] In an embodiment of the present application, the target pedestrian features may be pedestrian features known in a database, or may be pedestrian detection features extracted and retained in a previous pedestrian re-identification process; the captured image may be an image captured by a camera device (such as a camera), which records image information of several pedestrians.
[0071] Step 120: input the captured image into a super-resolution detection model to perform image super-resolution detection, and obtain a detection frame image output by the super-resolution detection model, wherein the detection frame image includes a plurality of pedestrian detection frames, and the pedestrian detection frames are pedestrian feature candidate frames of pedestrian images whose image size is greater than a first size threshold;
[0072] In the embodiments of the present application, the captured image can be input into the super-resolution detection model, and through this super-resolution detection model, the captured image is super-resolution reconstructed into a pedestrian image with an image size larger than the first size threshold. Then, the super-resolution detection model performs target detection on the pedestrian information of the pedestrian image and outputs it as a detection box image. Specifically, the super-resolution detection model can output multiple detection box images regarding the pedestrian information of the pedestrian image. At this time, each detection box image is an image representation of a pedestrian detection box; alternatively, the super-resolution detection model can output a detection box image of the pedestrian information of the pedestrian image. At this time, this detection box image is a set of image representations of multiple pedestrian detection boxes, which can be specifically obtained by splicing the image representations of multiple pedestrian detection boxes. This application will not elaborate here.
[0073] Step 130: Input the detection box image into the super-resolution detection model for super-resolution detection of the detection box to obtain pedestrian detection features. The pedestrian detection features are the pedestrian features of the pedestrian detection box with a detection box size larger than the second size threshold.
[0074] In the embodiments of the present application, the detection box image can be input into the super-resolution detection model again, and the super-resolution detection model performs super-resolution reconstruction on the detection box image so that the detection box size of the pedestrian detection box in the detection box image is larger than the second size threshold, obtaining the super-resolution reconstructed detection box image; then, the super-resolution detection model performs target detection on the super-resolution reconstructed detection box image and outputs the pedestrian detection features of the pedestrian detection box.
[0075] It can be understood that the first size threshold and the second size threshold in the embodiments of the present application can be set according to the actual situation, and this application does not limit this here.
[0076] Step 140: According to the target pedestrian features, perform matching and re-identification on the pedestrian detection features to obtain the pedestrian re-identification result of the captured image.
[0077] In the embodiments of the present application, feature matching can be performed on the target pedestrian features and the pedestrian detection features. When the feature matching between the target pedestrian features and the pedestrian detection features is successful, the pedestrian detection features that match the target pedestrian features are generated as the pedestrian re-identification result of the captured image; alternatively, when the feature matching between the target pedestrian features and the pedestrian detection features is unsuccessful, step 120 or step 130 can be returned for execution.
[0078] In some embodiments, the super-resolution detection model includes a size detection network, an image super-resolution network, and a target detection network. Inputting the first target image into the super-resolution detection model for super-resolution detection to obtain an output target includes:
[0079] A1. Input the first target image into the size detection network for size detection to obtain a second target image and the target size of the second target image;
[0080] A2. Obtain a target threshold corresponding to the target size;
[0081] A3. Compare the target threshold and the target size to obtain a size determination result;
[0082] A4. If the size determination result is that the target size is greater than the target threshold, input the second target image into the target detection network for target detection to obtain the output target;
[0083] Or, A5. If the size determination result is that the target size is less than or equal to the target threshold, input the second target image into the image super-resolution network to obtain a super-resolved image, and input the super-resolved image into the target detection network for target detection to obtain the output target;
[0084] In the embodiments of the present application, the first target image may be a captured image or a detection frame image. Specifically, if the first target image is a captured image, the corresponding second target image is a pedestrian image, the target size is the image size of the pedestrian image, the target threshold is the first size threshold, and the output target output by the target detection network is a detection frame image.
[0085] Or, if the first target image is a detection frame image, the corresponding second target image may be the original detection frame image or the finely adjusted detection frame image, the target size is the image size of the second target image, the target threshold is the second size threshold, and the output target output by the target detection network is pedestrian detection features.
[0086] It can be understood that if the first target image is a captured image, step A1 can be to use a size detection model to roughly locate the pedestrian image in the captured image and crop the located pedestrian image to obtain a pedestrian image and the image size of the pedestrian image; then, compare the size relationship between the image size of the pedestrian image and the corresponding first size threshold to obtain a size determination result. Specifically, if the size determination result is that the image size of the pedestrian image is greater than the first size threshold, it means that the pedestrian image is a high-resolution image. At this time, the pedestrian image can be directly input into the target detection network, and the target detection network can be used to perform feature bounding detection on the pedestrians in the pedestrian image to obtain a detection box image including several pedestrian detection boxes; or, if the size determination result is that the image size of the pedestrian image is less than or equal to the first size threshold, it means that the pedestrian image is a low-resolution image. At this time, the pedestrian image can be input into an image super-resolution network, and the image super-resolution network can be used to perform super-resolution reconstruction on the pedestrian image to obtain a super-resolution image, and the target detection network can be used to perform target detection on the super-resolution image to obtain a detection box image.
[0087] It should be noted that if the first target image is a detection box image, the detection box image can be input into the size detection model for detection, and the size detection model can be used to finely locate and crop the detection box image of the pedestrian detection box in the detection box image to obtain the original detection box image or the finely adjusted detection box image. Steps A2 to A5 for the detection box image are similar to steps A2 to A5 for the captured image described above and can be simply deduced by analogy.
[0088] It is worth mentioning that when the fine positioning result of the size detection model is consistent with the original detection box image of the pedestrian detection box, the obtained second target image can be the original detection box image; or, when the fine positioning result of the size detection model is inconsistent with the original detection box image of the pedestrian detection box, the pedestrian detection box can be finely cropped according to the fine positioning result to obtain a finely adjusted detection box image.
[0089] Refer to Figure 2 , in some embodiments, inputting the first target image into the size detection network for size detection to obtain a second target image includes:
[0090] B1. Perform multi-level feature extraction on the first target image to obtain a number of first intermediate features, and the feature scales of each first intermediate feature are different;
[0091] B2. Perform upsampling feature fusion on all the first intermediate features to obtain a number of second intermediate features;
[0092] B3. Classify and crop the first target image according to all the second intermediate features to obtain the second target image.
[0093] In an embodiment of the present application, the size detection network may be an improved yolov11 model. In step B1, the first target image may be input into the backbone layer of the size detection network for feature extraction, and the outputs of the first C3K2 module, the second C3K2 module, and the C2PSA module in the backbone layer of the size detection network may be determined as the first intermediate features corresponding to the feature scales.
[0094] It can be understood that in step B2, first, the first intermediate feature output by the C2PSA module may be upsampled and then feature concatenated with the first intermediate feature output by the second C3K2 module to obtain a first upsampled feature. Then, the first upsampled feature may be further upsampled and then feature concatenated with the first intermediate feature output by the first C3K2 module to obtain a second upsampled feature, and the second upsampled feature may be determined as the first second intermediate feature, and the second upsampled feature may be subjected to sparse convolution to obtain a second second intermediate feature. Then, after the second second intermediate feature is subjected to sparse convolution again, it may be feature concatenated with the first intermediate feature output by the C2PSA module to obtain a third intermediate feature.
[0095] It should be noted that in step B3, the second intermediate feature may be classified and detected through the classification head in the size detection network, and the image frame of the first target image may be cropped based on the classification and detection result output by the size detection network to obtain a second target image. In addition, the size detection model in the embodiment of the present application may also be a neural network model, for example, it may be other versions of the yolo model, and the present application does not limit this here.
[0096] In some embodiments, the image super-resolution network includes a plurality of cascaded pyramid residual feature extraction modules. The step of inputting the second target image into the image super-resolution network to obtain a super-resolution image includes:
[0097] C1. Extract shallow features from the second target image to obtain shallow features;
[0098] C2. Extract deep features from the shallow features through the plurality of cascaded pyramid residual feature extraction modules to obtain deep features;
[0099] C3. Perform feature fusion and feature pixel rearrangement on the deep features according to the shallow features to obtain the super-resolution image.
[0100] In the embodiments of the present application, step C1 may be to perform shallow feature extraction on the second target image. Specifically, it may be to implement shallow feature extraction on the second target image through several convolutional layers to obtain the shallow features of the second target image. Step C2 may input the shallow features into a cascaded pyramid residual feature extraction module, and each pyramid residual feature extraction module extracts deep features layer by layer from the shallow features to obtain the deep features of the second target image.
[0101] It can be understood that step C3 may first be to perform feature fusion on the shallow features and the deep features. Specifically, an element-wise addition operation may be performed on the shallow features and the deep features. Then, the fused features are sequentially input into a convolutional layer and a pixel shuffle layer, and the convolutional layer and the pixel shuffle layer map the fused features into a high-resolution image to obtain a super-resolution image.
[0102] Referring to Figure 3 , in some embodiments, deep feature extraction is performed on the input features through the pyramid residual feature extraction module to obtain output features, including:
[0103] D1. Perform multi-level feature convolution extraction on the input features to obtain third intermediate features;
[0104] D2. According to the third intermediate features, perform residual feature fusion on the input features to obtain fourth intermediate features;
[0105] D3. Perform pyramid compression attention extraction on the fourth intermediate features to obtain the output features;
[0106] In the embodiments of the present application, for a certain pyramid residual feature extraction module, if this pyramid residual feature extraction module is the first pyramid residual feature extraction module among all the cascaded pyramid residual feature extraction modules, its input features may be shallow features; or, if this pyramid residual feature extraction module is the second or subsequent pyramid residual feature extraction module among all the cascaded pyramid residual feature extraction modules, its input features are the output features output by the previous pyramid residual feature extraction module. Also, if this pyramid residual feature extraction module is the last pyramid residual feature extraction module among all the cascaded pyramid residual feature extraction modules, its output features are deep features.
[0107] It can be understood that step D1 can be to input the input features into a combination of three cascaded 3×3 convolutional layers and ReLu activation functions for multi-level feature convolution extraction, so as to obtain the extracted features output by the combination of the last 3×3 convolutional layer and ReLu activation function, and determine the extracted features as the third intermediate features. Step D2 can be to perform an element-wise addition operation on the third intermediate features and the input features, so as to achieve the residual feature fusion of the third intermediate features and the input features in the pyramid residual feature extraction module, and obtain the fourth intermediate features.
[0108] It should be noted that step D3 can be to directly input the fourth intermediate features into the Pyramid Squeeze Attention (PSA) module. Specifically, the PSA module performs multi-scale feature extraction and channel attention recalibration on the fourth intermediate features, so as to obtain the output features.
[0109] In some embodiments, inputting the third target image into the target detection network for target detection to obtain the output target includes:
[0110] E1. Performing depthwise separable convolution processing on the image features of the third target image to obtain fifth intermediate features;
[0111] E2. Extracting attention for the fifth intermediate features to obtain the feature attention of the fifth intermediate features;
[0112] E3. According to the feature attention, performing attention fusion on the fifth intermediate features to obtain sixth intermediate features;
[0113] E4. According to the sixth intermediate features, performing feature fusion on the image features of the second target image to obtain the output target;
[0114] In the embodiments of the present application, the image features of the third target image can be sequentially input into a 1×1 convolutional layer, a Batch normalization layer, and a hard-sigmoid activation function layer, as well as a 3×3 depthwise separable convolutional layer, a Batch Norm normalization layer, and a hard-sigmoid activation function layer, so as to achieve depthwise separable convolution processing of the image features and obtain fifth intermediate features. Step E2 can be to sequentially input the fifth intermediate features into an average pooling layer, a fully connected layer, and a ReLu activation function layer, as well as a fully connected layer and a hard-sigmoid activation function layer, so as to obtain an attention map corresponding to the fifth intermediate features, and determine the attention map as the feature attention of the fifth intermediate features.
[0115] It can be understood that step E3 can first perform an element-wise multiplication operation on the feature attention and the fifth intermediate feature to obtain the multiplied fifth intermediate feature, and then sequentially input the multiplied fifth intermediate feature into a 1×1 convolutional layer, a Batch Norm normalization layer, and a hard-sigmoid activation function layer for processing to obtain the sixth intermediate feature.
[0116] It should be noted that the feature fusion in step E4 can be achieved by a skip connection to perform an element-wise addition operation on the image feature of the second target image and the sixth intermediate feature, thereby realizing the feature fusion of the image feature to obtain the output target.
[0117] In some embodiments, matching and re-identifying the pedestrian detection feature according to the target pedestrian feature to obtain the pedestrian re-identification result of the captured image includes:
[0118] F1. Obtain a similarity distance threshold;
[0119] F2. Calculate the feature distance between the target pedestrian feature and the pedestrian detection feature according to the target pedestrian feature to obtain the feature distance;
[0120] F3. Compare the feature distance and the similarity distance threshold to obtain a threshold comparison result;
[0121] F4. If the threshold comparison result is that the feature distance is less than the threshold comparison result, then determine the pedestrian detection feature as the pedestrian re-identification result of the captured image.
[0122] In the embodiments of the present application, the specific value of the similarity distance threshold can be flexibly set according to the actual situation, such as any one of 0.1, 0.2, 0.4, etc. Step F2 can be to calculate the feature distance between the target pedestrian feature and the pedestrian detection feature, and the specific type of the feature distance can be Euclidean distance, cosine distance, etc. Step F3 can be to compare the size relationship between the feature distance and the similarity distance threshold to obtain the threshold comparison result.
[0123] It can be understood that if the threshold comparison result is that the feature distance is less than the threshold comparison result, then the pedestrian re-identification result of the captured image can be generated according to the pedestrian detection feature; or, if the threshold comparison result is that the feature distance is greater than or equal to the threshold comparison result, then steps 120 or 130 can be returned for execution.
[0124] Next, a pedestrian re-identification system proposed according to an embodiment of the present application will be described in detail with reference to the accompanying drawings.
[0125] Refer to Figure 4 , a pedestrian re-identification system proposed in the embodiments of the present application includes:
[0126] The first processing unit 101 is configured to obtain target pedestrian features and a captured image to be detected;
[0127] The second processing unit 102 is configured to input the captured image into a super-resolution detection model for image super-resolution detection, and obtain a detected box image output by the super-resolution detection model. The detected box image includes a plurality of pedestrian detection boxes, and the pedestrian detection box is a candidate box for pedestrian features of a pedestrian image with an image size greater than a first size threshold;
[0128] The third processing unit 103 is configured to input the detected box image into the super-resolution detection model for detected box super-resolution detection, and obtain pedestrian detection features. The pedestrian detection features are pedestrian features of a pedestrian detection box with a detected box size greater than a second size threshold;
[0129] The fourth processing unit 104 is configured to perform matching and re-identification on the pedestrian detection features according to the target pedestrian features, and obtain a pedestrian re-identification result of the captured image.
[0130] Referring to Figure 5 , an embodiment of the present application further provides an electronic device, including:
[0131] At least one processor 201;
[0132] At least one memory 202, configured to store at least one program;
[0133] When the at least one program is executed by the at least one processor 201, the at least one processor 201 implements the above method embodiment.
[0134] Similarly, it can be understood that the content in the above method embodiment is applicable to this device embodiment. The functions specifically implemented by this device embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method embodiment.
[0135] An embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor 201 is stored. The program executable by the processor 201 is used to implement the above method embodiment when executed by the processor 201.
[0136] Similarly, the content in the above method embodiment is applicable to this computer-readable storage medium embodiment. The functions specifically implemented by this computer-readable storage medium embodiment are the same as those of the above method embodiment, and the beneficial effects achieved are also the same as those of the above method embodiment.
[0137] In some alternative embodiments, the functions / operations recited in the block diagrams may not occur in the order presented in the operational illustrations. For example, depending on the functions / operations involved, two blocks shown in succession may actually be executed substantially concurrently or the blocks can sometimes be executed in the reverse order. Additionally, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more thorough understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered and in which sub-operations described as part of a larger operation are performed independently.
[0138] Furthermore, although the present application has been described in the context of functional modules, it should be understood that one or more of the functions and / or features may be integrated in a single physical device and / or software module unless otherwise stated to the contrary, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for an understanding of the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skill of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.
[0139] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, or a part of such technical solution, may be embodied in the form of a software product stored in a storage medium, including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method according to the embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a portable hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0140] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence of executable instructions for implementing logical functions, which can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device.
[0141] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0142] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0143] In the above description of this specification, the descriptions referring to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.
[0144] Although embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application. The scope of the present application is defined by the claims and their equivalents.
[0145] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.
Claims
1. A pedestrian re-identification method, characterized in that, Including: Obtain target pedestrian features and a captured image to be detected; Input the captured image into a super-resolution detection model for image super-resolution detection to obtain a detection box image output by the super-resolution detection model. The detection box image includes a plurality of pedestrian detection boxes, and the pedestrian detection box is a candidate box for pedestrian features of a pedestrian image with an image size greater than a first size threshold; Input the detection box image into the super-resolution detection model for detection box super-resolution detection to obtain pedestrian detection features, where the pedestrian detection features are the pedestrian features of the pedestrian detection box with a detection box size greater than a second size threshold; According to the target pedestrian features, perform matching and re-identification on the pedestrian detection features to obtain a pedestrian re-identification result of the captured image.
2. The method according to claim 1, wherein The super-resolution detection model includes a size detection network, an image super-resolution network, and a target detection network. Inputting a first target image into the super-resolution detection model for super-resolution detection to obtain an output target includes: Input the first target image into the size detection network for size detection to obtain a second target image and the target size of the second target image; Obtain a target threshold corresponding to the target size; Compare the target threshold and the target size to obtain a size determination result; If the size determination result is that the target size is greater than the target threshold, input the second target image into the target detection network for target detection to obtain the output target; or, if the size determination result is that the target size is less than or equal to the target threshold, input the second target image into the image super-resolution network to obtain a super-resolution image, and input the super-resolution image into the target detection network for target detection to obtain the output target; Wherein, the first target image is the captured image or the detection box image.
3. The method according to claim 2, characterized in that, Input the first target image into the size detection network for size detection to obtain a second target image, including: Perform multi-level feature extraction on the first target image to obtain a plurality of first intermediate features, and the feature scales of each first intermediate feature are different; Perform upsampling feature fusion on all the first intermediate features to obtain a plurality of second intermediate features; According to all the second intermediate features, perform classification and cropping on the first target image to obtain the second target image.
4. The method according to claim 2, wherein The image super-resolution network includes a plurality of cascaded pyramid residual feature extraction modules. Inputting the second target image into the image super-resolution network to obtain a super-resolution image includes: Perform shallow feature extraction on the second target image to obtain shallow features; Perform deep feature extraction on the shallow features through the plurality of cascaded pyramid residual feature extraction modules to obtain deep features; According to the shallow features, perform feature fusion and feature pixel rearrangement on the deep features to obtain the super-resolution image.
5. The method according to claim 4, wherein Performing deep feature extraction on the input feature through the pyramid residual feature extraction module to obtain an output feature, including: Perform multi-level feature convolution extraction on the input feature to obtain third intermediate features; According to the third intermediate feature, perform residual feature fusion on the input feature to obtain a fourth intermediate feature; Perform pyramid compression attention extraction on the fourth intermediate feature to obtain the output feature; Wherein, the input feature is the shallow feature or the output feature output by the previous pyramid residual feature extraction module.
6. The method according to claim 2, characterized in that Input the third target image into the target detection network for target detection to obtain the output target, including: Perform depthwise separable convolution processing on the image feature of the third target image to obtain a fifth intermediate feature; Perform attention extraction on the fifth intermediate feature to obtain the feature attention of the fifth intermediate feature; According to the feature attention, perform attention fusion on the fifth intermediate feature to obtain a sixth intermediate feature; According to the sixth intermediate feature, perform feature fusion on the image feature of the second target image to obtain the output target; Wherein, the third target image is the second target image or the super-resolution image.
7. The method according to claim 1, wherein The method of performing matching re-identification on the pedestrian detection feature according to the target pedestrian feature to obtain the pedestrian re-identification result of the captured image includes: Obtain a similarity distance threshold; Calculate the distance of the pedestrian detection feature according to the target pedestrian feature to obtain a feature distance; Compare the feature distance with the similarity distance threshold to obtain a threshold comparison result; If the threshold comparison result is that the feature distance is less than the threshold comparison result, determine the pedestrian detection feature as the pedestrian re-identification result of the captured image.
8. A pedestrian re-identification system, characterized in that, Includes: A first processing unit, configured to obtain a target pedestrian feature and a captured image to be detected; A second processing unit, configured to input the captured image into a super-resolution detection model for image super-resolution detection to obtain a detection box image output by the super-resolution detection model, where the detection box image includes a plurality of pedestrian detection boxes, and the pedestrian detection box is a candidate box for the pedestrian feature of a pedestrian image with an image size greater than a first size threshold; A third processing unit, configured to input the detection box image into the super-resolution detection model for detection box super-resolution detection to obtain a pedestrian detection feature, where the pedestrian detection feature is the pedestrian feature of a pedestrian detection box with a detection box size greater than a second size threshold; A fourth processing unit, configured to perform matching re-identification on the pedestrian detection feature according to the target pedestrian feature to obtain the pedestrian re-identification result of the captured image.
9. An electronic device, characterized in that, Includes: At least one processor; At least one memory, configured to store at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1-7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that, The program executable by the processor, when executed by the processor, is used to implement the method according to any one of claims 1-7.