Aligned pedestrian re-identification system and method based on boundary progressive sample optimization and DDT model

Through the boundary progressive sample optimization and the aligned pedestrian re-identification system of the DDT model, the problem of accurate matching of traditional pedestrian re-identification technology under different cameras is solved, and accurate positioning and monitoring of pedestrians is realized, which is suitable for pedestrian recognition in complex environments.

CN120299062APending Publication Date: 2025-07-11SHANDONG JIANZHU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510356301.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Traditional pedestrian re-identification technology is difficult to achieve accurate identity matching under different cameras, and is greatly affected by environmental changes and has low efficiency.

Method used

The aligned pedestrian re-identification system based on boundary progressive sample optimization and DDT model is adopted, including image preprocessing module, target feature measurement model training module and cross-camera search module, to achieve accurate pedestrian positioning through feature extraction, boundary progressive sample optimization and multi-scale feature alignment technologies.

Benefits of technology

It improves the accuracy and efficiency of pedestrian identification, and can accurately locate target pedestrians in complex environments, and is suitable for monitoring and tracking tasks in actual scenarios such as campuses and streets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299062A_ABST
    Figure CN120299062A_ABST
Patent Text Reader

Abstract

The invention provides an aligned pedestrian re-identification system and method based on boundary progressive sample optimization and a DDT model, and the system comprises an image preprocessing module which is used for carrying out the feature extraction of an existing pedestrian data set based on the DDT model, and obtaining a feature map; the target feature measurement model training module is used for calculating the feature distance of the feature map, performing boundary progressive sample optimization on the existing pedestrian data set based on the feature distance, and training a target feature measurement model; and the cross-camera retrieval module is used for performing target pedestrian recognition and positioning on pedestrian images collected by the monitoring camera based on the trained target feature measurement model in combination with a multi-scale feature alignment technology, obtaining a recognition and positioning result and completing aligned pedestrian re-recognition based on boundary progressive sample optimization and a DDT model. According to the technical scheme of the invention, pedestrians with different backgrounds, different angles and different postures shot by a camera are accurately classified and positioned, and a basis is provided for personnel monitoring and target tracking tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision and deep learning, and particularly relates to an aligned pedestrian re-identification system and method based on boundary progressive sample optimization and DDT model. Background Art

[0002] Pedestrian re-identification (Re-ID) is a key computer vision task, mainly used to identify and match the identities of the same person under different cameras. This technology has wide applications in fields such as security monitoring, public safety, and intelligent transportation.

[0003] Due to the complexity of the urban environment, the behaviors and activities of pedestrians in public areas are affected by factors such as different lighting, camera angles, and occlusions, which makes it difficult for traditional visual recognition methods to achieve accurate identity matching. There may be different shooting conditions and perspectives between different cameras, which poses a great challenge to pedestrian re-identification. Traditional pedestrian recognition technologies rely on manual feature extraction and matching, with low efficiency and being easily affected by environmental changes.

[0004] With the progress of deep learning technology, especially the application of convolutional neural networks (CNNs) and other deep models, the performance of pedestrian re-identification has been significantly improved. Modern pedestrian re-identification systems can automatically extract feature representations of pedestrians by training on large-scale datasets, thereby improving the matching accuracy under different cameras. In addition, pedestrian re-identification technology also combines methods such as image retrieval and metric learning to improve the recognition effect in complex environments. Summary of the Invention

[0005] Aiming at the above technical problems, the present invention proposes an aligned pedestrian re-identification system and method based on boundary progressive sample optimization and DDT model, which can accurately classify and locate pedestrians with different backgrounds, angles, and postures captured by cameras, providing a basis for the implementation of personnel monitoring and target tracking tasks.

[0006] An aligned pedestrian re-identification system based on boundary progressive sample optimization and DDT model, comprising:

[0007] An image preprocessing module, configured to extract features from an existing pedestrian dataset based on the DDT model to obtain a feature map;

[0008] A target feature metric model training module, configured to calculate the feature distance of the feature map and optimize the existing pedestrian dataset based on the feature distance to train a target feature metric model;

[0009] The cross-camera retrieval module is used to perform target pedestrian recognition and positioning on pedestrian images collected by surveillance cameras based on a trained target feature metric model combined with multi-scale feature alignment technology, obtain the recognition and positioning results, and complete the aligned pedestrian re-identification based on boundary progressive sample optimization and the DDT model.

[0010] Preferably, the image preprocessing module includes:

[0011] The image preprocessing unit is used to unify the size of the existing pedestrian dataset and preprocess the unified-size existing pedestrian dataset by using random horizontal flipping, normalization, and variance and standardization illumination compensation to obtain the preprocessed existing pedestrian dataset;

[0012] The feature extraction unit is used to extract features from the preprocessed existing pedestrian dataset based on the DDT model to obtain the feature map.

[0013] Preferably, the feature extraction unit includes:

[0014] The dynamic convolution decomposition sub-unit is used to replace the 7×7 large convolution kernel of the first convolution layer in the input stage of the ResNet-50 basic network with three cascaded 3×3 small convolution kernels, process the preprocessed existing pedestrian dataset, and obtain the first feature map;

[0015] The two-way stride transpose sub-unit is used to perform two-way stride transpose on the strides of the last two convolution layers in the bottleneck layer of the ResNet-50 basic network layer, process the first feature map, and obtain the second feature map;

[0016] The multi-scale delayed downsampling sub-unit is used to add a pooling layer with a preset size and stride before the convolution layer on the second downsampling layer path of the fourth layer of the ResNet-50 basic network, process the second feature map, and obtain the third feature map;

[0017] The output sub-unit is used to pool the third feature map based on the global average pooling layer and output the final feature map through the fully connected layer.

[0018] Preferably, the target feature metric model training module includes:

[0019] The feature calculation unit is used to calculate the global feature of the feature map based on global pooling and calculate the local feature of the feature map based on horizontal maximum pooling;

[0020] The distance calculation unit is used to calculate the global feature distance and the local feature distance respectively based on the global feature and the local feature;

[0021] A boundary progressive sample optimization unit, which is used to perform boundary progressive sample optimization on the preprocessed existing pedestrian dataset based on the global feature distance and the local feature distance to obtain optimized samples;

[0022] A loss calculation unit, which is used to calculate the global loss and the local loss of the optimized samples based on the dynamic curriculum triplet loss algorithm, and complete the construction and training of the target feature metric model.

[0023] Preferably, the process of obtaining the optimized samples includes:

[0024] Based on the sample distance matrix and the positive and negative sample masks of the preprocessed existing pedestrian dataset, obtain the positive and negative sample distances, and define the hard sample boundary; wherein, the positive and negative sample distances are the distances between the samples in the sample distance matrix and the negative samples and the positive samples;

[0025] When the distance between the selected pedestrian image and the target pedestrian image is less than the positive sample distance but the actually selected pedestrian image is a negative sample, perform boundary progressive sample optimization;

[0026] Take the intersection of the hard sample boundary and the boundary progressive sample optimization result, update the positive and negative sample masks, and weight the distance between the hard samples that meet the updated positive and negative sample masks and the boundary progressive sample optimization result to obtain the optimized samples.

[0027] The present invention also provides an aligned pedestrian re-identification method based on boundary progressive sample optimization and the DDT model. Applying the system, the method includes:

[0028] Extract features from the existing pedestrian dataset based on the DDT model to obtain a feature map;

[0029] Calculate the feature distance of the feature map and perform boundary progressive sample optimization on the existing pedestrian dataset based on the feature distance to train the target feature metric model;

[0030] Based on the trained target feature metric model and combined with the multi-scale feature alignment technology, perform target pedestrian recognition and positioning on the pedestrian images collected by the surveillance camera to obtain the recognition and positioning result, and complete the aligned pedestrian re-identification based on boundary progressive sample optimization and the DDT model.

[0031] Preferably, the method of obtaining the feature map includes:

[0032] Unify the size of the existing pedestrian dataset, and perform preprocessing on the existing pedestrian dataset with unified size by using random horizontal flipping, normalization, and variance and standardization illumination compensation to obtain the preprocessed existing pedestrian dataset;

[0033] Feature extraction is performed on the preprocessed existing pedestrian dataset based on the DDT model to obtain the feature map.

[0034] Preferably, the method for feature extraction of the preprocessed existing pedestrian dataset based on the DDT model includes:

[0035] Replace the 7×7 large convolution kernel of the first convolution layer in the input stage of the ResNet-50 basic network with three cascaded 3×3 small convolution kernels, process the preprocessed existing pedestrian dataset, and obtain the first feature map;

[0036] Perform two-way stride transposition on the strides of the last two convolution layers in the bottleneck layer of the ResNet-50 basic network, process the first feature map, and obtain the second feature map;

[0037] Add a pooling layer with a preset size and stride before the convolution layer in the second downsampling layer path of the fourth layer of the ResNet-50 basic network, process the second feature map, and obtain the third feature map;

[0038] Pool the third feature map based on the global average pooling layer, and output the final feature map through the fully connected layer.

[0039] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0040] 1. Realize the pedestrian recognition and classification function in the surveillance camera and complete the pedestrian precise positioning task;

[0041] 2. By using the boundary progressive sample optimization method, the model's ability to identify difficult-to-distinguish samples is improved;

[0042] 3. By using the DDT model, the feature extraction amount and calculation efficiency are greatly improved, and at the same time, the number of parameters is effectively reduced. With a more concise network architecture, better classification accuracy is achieved, and the target pedestrian can be more accurately located in the huge pedestrian image data;

[0043] 4. Apply the pedestrian precise positioning method to actual scenarios such as campuses and streets, and perform precise classification and positioning on pedestrians with different backgrounds, different angles, and different postures captured by the camera, providing a basis for the development of personnel monitoring and target tracking tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the present invention, the following briefly introduces the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 Schematic diagram of the aligned pedestrian re-identification system based on boundary progressive sample optimization and DDT model according to an embodiment of the present invention;

[0046] Figure 2 Schematic diagram of the ResNet-50 basic network according to an embodiment of the present invention;

[0047] Figure 3 Schematic diagram of dynamic convolution decomposition according to an embodiment of the present invention;

[0048] Figure 4 Frame of the image preprocessing module according to an embodiment of the present invention;

[0049] Figure 5 Feature distance calculation framework according to an embodiment of the present invention;

[0050] Figure 6 Effect diagram of boundary progressive sample optimization according to an embodiment of the present invention. Specific implementation manners

[0051] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners.

[0053] Embodiment 1

[0054] As Figure 1 shown, an aligned pedestrian re-identification system based on boundary progressive sample optimization and DDT model includes: an image preprocessing module, a target feature metric model training module, and a cross-camera retrieval module.

[0055] The image preprocessing module is used to extract features from the existing pedestrian dataset based on the DDT model to obtain a feature map. The role of this module is to extract the feature map from the existing dataset. Using the DDT model can greatly increase the scope of feature extraction and save computing resources, thereby solving the problem of the quality of feature map extraction. The DDT (deferred, decompose and transpose) model (the Chinese meaning is "deferred, decompose and transpose", hereinafter referred to as the DDT model).

[0056] Furthermore, the implementation manner lies in that the image preprocessing module includes:

[0057] An image preprocessing unit, which is used to unify the size of the existing pedestrian dataset, and preprocess the unified-size existing pedestrian dataset by adopting random horizontal flipping, normalization, and variance and standardized illumination compensation to obtain a preprocessed existing pedestrian dataset; specifically, the input image size is unified to 224×224 pixels, and random horizontal flipping, normalization (mean [0.485, 0.456, 0.406], variance [0.229, 0.224, 0.225]) and standardized illumination compensation are adopted. The number of input images is N, and the output size in the input stage will be halved to 112×112.

[0058] A feature extraction unit, which is used to extract features from the preprocessed existing pedestrian dataset based on the DDT model to obtain a feature map.

[0059] A further implementation manner is that the feature extraction unit includes:

[0060] A dynamic convolution decomposition subunit, which is used to replace the 7×7 large convolution kernel of the first convolutional layer in the input stage of the ResNet-50 basic network (as Figure 2 shown) with three cascaded 3×3 small convolution kernels, greatly reducing the computational amount while deepening the network and keeping the output size unchanged at 56×56 (as Figure 3 ), process the preprocessed existing pedestrian dataset to obtain a first feature map.

[0061] A bidirectional stride transpose subunit, which is used to perform bidirectional stride transpose on the strides of the last two convolutional layers in the bottleneck layer of the ResNet-50 basic network layer, process the first feature map to obtain a second feature map; specifically, the bidirectional stride transpose is to change the strides of the last two convolutional layers in the bottleneck layer of the ResNet-50 basic network layer (stage 4) from the original (2, 1) to (1, 2), increasing the sampling area, and the output is 56×28.

[0062] A multi-scale postponed downsampling subunit, which is used to add a pooling layer with a preset size and stride before the convolutional layer of path two in the downsampling layer of the fourth layer of the ResNet-50 basic network, process the second feature map to obtain a third feature map; specifically, add an average pooling layer with a size of 2×2 and a stride of 2 before the convolutional layer of path two in the downsampling layer of layer 4 (stage 4) of the ResNet-50 basic network to increase the scanning range and greatly improve the accuracy of feature extraction.

[0063] An output subunit, which is used to pool the third feature map based on the global average pooling layer and output the final feature map through a fully connected layer. Specifically, change the output size of the global average pooling layer in the output stage of the ResNet-50 basic network to (1, 1), and change the input and output sizes of the fully connected layer to (2048, N).

[0064] Specifically, the passed layers 1, 2, and 3 will maintain their sizes until the first downsampling operation in layer 4 halves the size again to 28×14; finally, the global average pooling layer in the output stage pools the 28×14 feature map into 1×1, with the number of input channels being 2048 (default for the resnet-50 network). After passing through the fully connected layer, the final output is (N, 2048, 1, 1), which is the extracted feature map, where N is the number of pedestrians, 2048 is the number of channels, and (1, 1) is the size of the feature map. Then, the processed feature map is input into the target feature metric model training module. The process is as Figure 4 shown.

[0065] The target feature metric model training module is used to calculate the feature distance of the feature map and optimize the boundary progressive samples of the existing pedestrian dataset based on the feature distance, and train the target feature metric model. The function of this module is to pass the feature map through two paths to obtain the global feature and the local feature respectively. These features are in the form of vectors. Then, the Euclidean distances between the global feature vectors and local feature vectors of different images are calculated respectively. The two parts of the distances are calculated using the dynamic curriculum triplet loss to obtain the most dissimilar samples under the same identity and the most similar samples under different identities, and the loss is calculated to measure the similarity between images. In this process, in order to improve the discrimination ability of the model, that is, the discrimination distance ability of the dynamic curriculum triplet loss, a method of boundary progressive sample optimization is designed to assist in the calculation. This model realizes the conversion of similarity into the gap of distance in a way of from front to back, two-way parallel and then combined. The process is as Figure 5 .

[0066] A further implementation method is that the target feature metric model training module includes:

[0067] The feature calculation unit is used to calculate the global feature of the feature map based on global pooling; calculate the local feature of the feature map based on horizontal max pooling; specifically, local / global feature calculation. Given the size of the input feature map as (N, 2048, 1, 1), after batch normalization and RELU activation, the size remains unchanged as (N, 2048, 1, 1). Then, after horizontal max pooling, input_size[3] is 1, min(input_size[3], 64) is 1, and the pooling kernel size is (1, 1), so there is Output i,j = max(Input i,j , Input j+k-1 ), where k is the size of the pooling kernel in the horizontal direction. For the pooling window area starting from the j-th column of the i-th row in the input feature map, and then passing through the convolutional layer Finally, the output local feature is (N, 128, 8); meanwhile, the input feature map passes through global average pooling where H is the height of the feature map, W is the width of the feature map, C is the number of channels, and Input i,j,c is the pixel value of the c-th channel at the i-th row and j-th column in the input feature map, and the global feature is (N, 2048).

[0068] A distance calculation unit for calculating the global feature distance and the local feature distance respectively based on the global feature and the local feature; specifically, the global distance is defined as calculating the Euclidean distance between two samples (global feature vectors), and the global feature is input where x i and x j are the feature vectors of two samples to obtain a global distance of (32, 32); for the local distance, first calculate the Euclidean distance matrix (dist_mat) between the batch samples (batch local feature vectors), where D ij represents the distance between the feature vector of the i-th sample in the i-th batch and the feature vector of the j-th sample in the j-th batch, and then perform normalization processing calculation using the hyperbolic tangent function, to obtain a local distance of (N, N).

[0069] A boundary progressive sample optimization unit for performing boundary progressive sample optimization on the preprocessed existing pedestrian dataset (i.e., the feature vector of the image) based on the global feature distance and the local feature distance to obtain optimized samples. Existing methods (such as hard sample mining) directly select the most difficult samples, but ignore the dynamic characteristics of the boundary samples, resulting in the model being prone to falling into local optima in the later stage of training. The boundary progressive sample optimization proposed by the present invention screens out more difficult-to-distinguish samples by defining a progressive boundary region. As Figure 6 shown.

[0070] A further implementation manner is that the process of obtaining the optimized samples includes:[[]]

[0071] Based on the sample distance matrix and the positive and negative sample masks of the preprocessed existing pedestrian dataset, obtain the positive and negative sample distances and define the hard sample boundary; wherein, the positive and negative sample distances are the distances between the sample and the negative sample and the positive sample in the sample distance matrix; specifically, first define the hard sample boundary, and according to the distance matrix of formula (2), respectively obtain the distance (dist_an) between the sample and the negative sample and the distance (dist_ap) between the sample and the positive sample in the distance matrix. The formulas are as follows:[[]]

[0072] dist_ap = max(dist_mat[is_pos], dim = 1) (1)

[0073] dist_an = min(dist_mat[is_neg], dim = 1) (2)

[0074] Among them, is_pos and is_neg are positive and negative sample masks. At the same time, a threshold (margin) is defined. When dist_an < margin, it is a hard negative sample, and when dist_ap > margin, it is a hard positive sample. dim = 1 means performing an aggregation operation in the second dimension (row direction).

[0075] When the distance between the selected pedestrian image and the target pedestrian image is less than the positive sample distance but the actually selected pedestrian image is a negative sample, boundary progressive sample optimization is performed. Specifically, considering special cases, if the distance between the selected image and the target image is less than dist_ap but the actually selected image is a negative sample, that is, the case where the two images are extremely similar, the hard sample boundary is no longer reliable. Therefore, for boundary progressive sample optimization (BPO), it is defined as:

[0076] BPO = {(d ap , d an ) ∈ R 2 ∣ d ap < d an ∧ d an - d ap < margin} (3)

[0077] Take the intersection of the hard sample boundary and the boundary progressive sample optimization result, update the positive and negative sample masks, and weight the distance between the hard samples that meet the updated positive and negative sample masks and the boundary progressive sample optimization result to obtain the optimized samples. Specifically, the distance between the hard samples that meet the updated mask and BPO is weighted by 1.5 times to increase the contribution to the loss function, and the final distances dist_an and dist_ap of the positive and negative samples are output.

[0078] The loss calculation unit is used to calculate the global loss and local loss of the optimized samples based on the dynamic curriculum triplet loss algorithm, and complete the construction and training of the target feature metric model.

[0079] The main defects of the traditional triplet loss are: the preset boundary value cannot adapt to different training stages and sample difficulty distributions; treating simple / hard samples equally, resulting in unstable model convergence; not considering the learning objective differences in the initial and later stages of training. Therefore, a three-level dynamic adjustment mechanism, namely the dynamic curriculum triplet loss, is proposed. The technical solution is as follows:

[0080] First is the stage-aware boundary adjustment, designing a hyperbolic tangent decay function to achieve dynamic boundary contraction:

[0081] α(t) = δ_max - (δ_max - δ_min) * tanh(5 * t / T) (4)

[0082] Among them, δ_max = 1.5 (initial maximum boundary) prevents the model from falling into local optima prematurely and provides sufficient exploration range for the initialization of the feature space. δ_min = 0.4 (final minimum boundary) enhances the sharpness of the decision boundary and the ability to capture subtle differences. t is the current training step, and T is the total number of training steps. The purpose of using the hyperbolic tangent function is to maintain the characteristics of rapid contraction in the initial stage and gentle transition in the later stage compared with linear attenuation. Through dynamic adjustment, the gradient contribution degree of the boundary optimization samples changes adaptively with the training process.

[0083] Secondly, there is the curriculum learning sampling strategy. To better utilize the samples obtained by optimizing the boundary progressive samples and other samples, it is necessary to optimize the learning efficiency of model training, that is, let the model learn gradually for different samples, so as to achieve the effect of improving the learning efficiency. In the initial stage of training, simple samples (70%) are emphasized to quickly establish a stable topological structure of the feature space; in the middle stage of training, balanced sampling is carried out to prevent the model from overfitting to samples of a specific difficulty. In the later stage of training, boundary progressive optimization samples (80%) are used to specifically enhance the robustness of the model to occlusion and pose changes. The dynamic curriculum triplet loss function is defined as:

[0084]

[0085] Among them, β represents the boundary parameter in the dynamic curriculum triplet loss, which is used to control the minimum interval between positive samples and negative samples. t is the current training round, and T is the total number of rounds; the simple sample set S easy is defined as d pos < 0.5 and d neg > 1.0 samples; the boundary progressive optimization sample set S BPO is dynamically screened by equation (3).

[0086] Finally, for the loss evaluation during the training process, the gradient is adaptively weighted, and a confidence-based loss weight is introduced:

[0087] weight = 1 + log(1 + exp(10 * (d_an_i - d_ap_i))) (6)

[0088] Gradient suppression is performed on potential negative samples (d_an < d_ap), and gradient enhancement is performed on high-set boundary progressive optimization samples (d_an >> d_ap), that is, the occurrence of negative samples is reduced, and at the same time, more attention is paid to boundary progressive optimization samples.

[0089] The cross-camera retrieval module is used to perform target pedestrian recognition and positioning on pedestrian images collected by surveillance cameras based on a trained target feature metric model combined with multi-scale feature alignment technology, obtain the recognition and positioning results, and complete the aligned pedestrian re-identification based on boundary progressive sample optimization and the DDT model.

[0090] The function of this module is to accurately classify pedestrians in surveillance cameras, using a multi-granularity feature fusion and positioning mechanism. Through the collection of pedestrian image data taken on-site and input into a high-precision aligned pedestrian re-identification model (target feature metric model) established by boundary progressive sample optimization and deep network improvement. First, extract the feature map and take each piece of the feature map as an input vector; second, feed the input vector to the above-mentioned trained target positioning model to accurately position the pedestrian; finally, input the output vector into the classifier, and different pedestrians are classified by the classifier to achieve the purpose of accurate positioning. Specifically, it includes the following technical solutions:

[0091] Firstly, it is the multi-scale feature alignment technology. On the 28×14 resolution feature map output by the DDT model, use spatial pyramid pooling (SPP) to construct multi-scale context information: design three-level pyramid pooling kernels (1×1, 3×3, 6×6), maintain the feature map resolution through dilated convolution, and at the same time use an attention-guided feature fusion formula:

[0092]

[0093] Among them, α i is the weight of the attention channel.

[0094] Secondly, it is to construct a human body part heat map based on local features: use a deformable convolutional network (DCN) to adapt to different postures and design a part correlation matrix:

[0095]

[0096] Among them, f i 、f j respectively represent the feature vectors of different parts.

[0097] Finally, it is to introduce a temporal memory module into the camera video stream data: design a gated recurrent unit (GRU) to store historical features and establish a spatio-temporal similarity metric function:

[0098]

[0099] Among them, λ is an adaptive weight parameter and T is the time window length.

[0100] Embodiment 2

[0101] The present invention also provides an aligned pedestrian re-identification method based on boundary progressive sample optimization and DDT model, and an application system. The method includes:

[0102] Extract features from the existing pedestrian dataset based on the DDT model to obtain a feature map;

[0103] Calculate the feature distance of the feature map and perform boundary progressive sample optimization on the existing pedestrian dataset based on the feature distance to train the target feature metric model;

[0104] Based on the trained target feature metric model, combine the multi-scale feature alignment technology to identify and locate the target pedestrian in the pedestrian image collected by the surveillance camera, obtain the recognition and location result, and complete the aligned pedestrian re-identification based on boundary progressive sample optimization and DDT model.

[0105] A further implementation manner is that the method for obtaining the feature map includes:

[0106] Unify the size of the existing pedestrian dataset, and perform preprocessing on the unified-size existing pedestrian dataset by using random horizontal flipping, normalization, and variance and standardization illumination compensation to obtain the preprocessed existing pedestrian dataset;

[0107] Extract features from the preprocessed existing pedestrian dataset based on the DDT model to obtain a feature map.

[0108] A further implementation manner is that the method for extracting features from the preprocessed existing pedestrian dataset based on the DDT model includes:

[0109] Replace the 7×7 large convolution kernel in the first convolution layer of the input stage of the ResNet-50 basic network with three cascaded 3×3 small convolution kernels, process the preprocessed existing pedestrian dataset to obtain a first feature map;

[0110] Transpose the stride of the last two convolution layers in the bottleneck layer of the ResNet-50 basic network in both directions, process the first feature map to obtain a second feature map;

[0111] Add a pooling layer with a preset size and stride before the convolution layer in the second downsampling layer path of the fourth layer of the ResNet-50 basic network, process the second feature map to obtain a third feature map;

[0112] Pool the third feature map based on the global average pooling layer and output the final feature map through the fully connected layer.

[0113] The embodiments described above are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An alignment pedestrian re-identification system based on boundary progressive sample optimization and DDT model, characterized in that Including: An image preprocessing module, configured to extract features from an existing pedestrian dataset based on the DDT model to obtain a feature map; A target feature metric model training module, configured to calculate the feature distance of the feature map and optimize the boundary progressive samples of the existing pedestrian dataset based on the feature distance, and train a target feature metric model; A cross-camera retrieval module, configured to perform target pedestrian recognition and positioning on pedestrian images collected by a surveillance camera based on the trained target feature metric model in combination with a multi-scale feature alignment technique, obtain a recognition and positioning result, and complete aligned pedestrian re-identification based on boundary progressive sample optimization and the DDT model.

2. The system according to claim 1, characterized in that, The image preprocessing module includes: An image preprocessing unit, configured to unify the size of the existing pedestrian dataset, and perform preprocessing on the unified-size existing pedestrian dataset by using random horizontal flipping, normalization, and variance and standardization light compensation to obtain a preprocessed existing pedestrian dataset; A feature extraction unit, configured to extract features from the preprocessed existing pedestrian dataset based on the DDT model to obtain the feature map.

3. The system according to claim 2, wherein The feature extraction unit includes: A dynamic convolution decomposition sub-unit, configured to replace the 7×7 large convolution kernel of the first convolution layer in the input stage of the ResNet-50 basic network with three cascaded 3×3 small convolution kernels, process the preprocessed existing pedestrian dataset, and obtain a first feature map; A bidirectional stride transpose sub-unit, configured to perform bidirectional stride transpose on the strides of the last two convolution layers in the bottleneck layer of the ResNet-50 basic network layer, process the first feature map, and obtain a second feature map; A multi-scale deferred downsampling sub-unit, configured to add a pooling layer with a preset size and stride before the convolution layer of the second path of the downsampling layer in the fourth layer of the ResNet-50 basic network, process the second feature map, and obtain a third feature map; An output sub-unit, configured to pool the third feature map based on a global average pooling layer and output the final feature map through a fully connected layer.

4. The system according to claim 2, wherein The target feature metric model training module includes: A feature calculation unit, configured to calculate the global feature of the feature map based on global pooling; calculate the local feature of the feature map based on horizontal maximum pooling; A distance calculation unit, configured to calculate a global feature distance and a local feature distance respectively based on the global feature and the local feature; A boundary progressive sample optimization unit, configured to optimize the boundary progressive samples of the preprocessed existing pedestrian dataset based on the global feature distance and the local feature distance to obtain optimized samples; A loss calculation unit, configured to calculate the global loss and the local loss of the optimized samples based on a dynamic curriculum triplet loss algorithm, and complete the construction and training of the target feature metric model.

5. The system according to claim 4, wherein The process of obtaining the optimized samples includes: Obtaining positive and negative sample distances based on the sample distance matrix and positive and negative sample masks of the preprocessed existing pedestrian dataset, and defining a hard sample boundary; wherein, the positive and negative sample distances are the distances between the samples and the negative samples and the positive samples in the sample distance matrix. When the distance between the selected pedestrian image and the target pedestrian image is less than the positive sample distance but the actually selected pedestrian image is a negative sample, boundary progressive sample optimization is performed; Take the intersection of the hard sample boundary and the boundary progressive sample optimization result, update the positive and negative sample masks, and weight the distance between the hard samples that meet the updated positive and negative sample masks and the boundary progressive sample optimization result to obtain the optimized samples.

6. A pedestrian re-identification method for alignment based on boundary progressive sample optimization and DDT model, applying the system according to any one of claims 1-5, characterized in that, The method includes: Extract features from the existing pedestrian dataset based on the DDT model to obtain a feature map; Calculate the feature distance of the feature map and perform boundary progressive sample optimization on the existing pedestrian dataset based on the feature distance to train the target feature metric model; Based on the trained target feature metric model and combined with the multi-scale feature alignment technology, perform target pedestrian recognition and localization on the pedestrian images collected by the surveillance camera to obtain the recognition and localization result, and complete the aligned pedestrian re-identification based on boundary progressive sample optimization and the DDT model.

7. The method according to claim 6, wherein The method for obtaining the feature map includes: Unify the size of the existing pedestrian dataset, and perform preprocessing on the unified-size existing pedestrian dataset by using random horizontal flipping, normalization, and variance and standardization illumination compensation to obtain the preprocessed existing pedestrian dataset; Extract features from the preprocessed existing pedestrian dataset based on the DDT model to obtain the feature map.

8. The method according to claim 7, characterized in that, The method for extracting features from the preprocessed existing pedestrian dataset based on the DDT model includes: Replace the 7×7 large convolutional kernel in the first convolutional layer of the input stage of the ResNet-50 basic network with three cascaded 3×3 small convolutional kernels, process the preprocessed existing pedestrian dataset to obtain a first feature map; Transpose the strides of the last two convolutional layers in the bottleneck layer of the ResNet-50 basic network layer in a two-way manner, process the first feature map to obtain a second feature map; Add a pooling layer with a preset size and stride before the convolutional layer in the second downsampling layer path of the fourth layer of the ResNet-50 basic network, process the second feature map to obtain a third feature map; Pool the third feature map based on the global average pooling layer and output the final feature map through the fully connected layer.

Citation Information

Cited By

  • Normalized structure reconstruction and feature sharpness perception online adaptation method and related equipment

    CN121189425A

  • Online Adaptive Methods and Related Equipment for Normalized Structure Reconstruction and Feature Sharpness Perception

    CN121189425B