Visible-infrared pedestrian re-identification method based on background weakening and joint loss
Patent Information
- Application Number
- CN202610861662.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-15
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]本发明的目的是为了解决现有可见光红外跨模态行人重识别技术在复杂场景下由于背景干扰引发的特征判别性不足和跨模态对齐偏差问题,提出了一种基于背景削弱和联合损失的可见光红外行人重识别方法
(1)本发明设计了背景削弱模块,通过自适应前景分割、背景结构破坏、纹理抑制与边缘保护技术,在输入层面主动抑制背景噪声,引导模型聚焦行人前景的身份特征,解决了背景干扰导致的特征学习偏差问题。
Smart Images

Figure CN122676571A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence computer vision technology, specifically relating to the design of a visible light infrared pedestrian re-identification method based on background attenuation and joint loss. Background Technology
[0002] With the acceleration of urbanization and the rapid development of intelligent security systems in my country, Person Re-Identification (Re-ID) technology has become one of the core supporting technologies for intelligent video surveillance systems. It aims to achieve cross-scene retrieval and identity association of specific pedestrians in non-overlapping multi-camera networks using computer vision methods. Traditional Re-ID technology has achieved mature applications in well-lit daytime environments, but in low-light or no-visible-light environments such as nighttime, tunnels, and backlit areas, the light-sensing capability of visible-light cameras is severely limited, leading to a sharp decline in system performance or even system failure, failing to meet the core requirement of all-weather monitoring.
[0003] Visible-Infrared Person Re-Identification (VI-ReID) combines the complementary advantages of visible light and infrared imaging. Visible light images provide rich color and texture details, while infrared images are unaffected by ambient light and can clearly present pedestrian outlines and structural information in low-light or even no-light environments, making it a core technology for solving the problem of continuous day and night surveillance. Current research in this field mainly focuses on mitigating the modal gap, using techniques such as representation learning, metric learning, and modal transformation to reduce the modal differences between visible light and infrared images, achieving some technological progress. However, in real-world complex surveillance scenarios, pedestrian foregrounds are often highly mixed with background elements such as buildings, vegetation, and vehicles. Background information presents completely different visual representations and feature distributions under visible light and infrared modalities, which can easily lead to the network over-focusing on background-irrelevant information during feature learning, weakening the extraction of discriminative features for pedestrian identity, and ultimately causing cross-modal feature alignment bias, limiting the robustness and generalization ability of the model in complex real-world scenarios. Current mainstream research focuses primarily on mitigating modal differences, and systematic suppression schemes for background interference still have significant shortcomings: First, there is a lack of dedicated modules for actively suppressing background noise at the input level, making it difficult to effectively guide the model to focus on the identity features of pedestrians in the foreground; second, the decoupling of shared and private features across modalities often adopts a fixed ratio partitioning strategy, which cannot adaptively achieve precise decoupling based on data distribution, making it difficult to provide refined guidance for cross-modal feature alignment; third, the loss function design fails to fully incorporate the feature characteristics after background weakening, making it impossible to achieve synergistic optimization of feature alignment and discriminative enhancement, resulting in the model's retrieval performance in complex background scenarios failing to meet actual security needs. Summary of the Invention
[0004] The purpose of this invention is to solve the problems of insufficient feature discriminativeness and cross-modal alignment deviation caused by background interference in complex scenes in existing visible light infrared cross-modal pedestrian re-identification technology. A visible light infrared pedestrian re-identification method based on background attenuation and joint loss is proposed.
[0005] The technical solution of this invention is: a visible light infrared pedestrian re-identification method based on background attenuation and joint loss, comprising the following steps: S1. Adjust the size of the input visible light pedestrian image and infrared pedestrian image to a preset size, and construct training batch data according to the preset sampling strategy.
[0006] S2. Construct a dual-path parallel feature extraction structure based on the basic feature extraction path of the original image and the enhanced feature path based on the background weakening module, and input the training batch data into the dual-path parallel feature extraction structure to extract multi-channel feature maps.
[0007] S3. Construct a weight-sharing backbone network based on the ResNet50 architecture, and input the multi-channel feature map into the shared feature extraction layer in the backbone network for deep feature extraction to obtain the global feature representation of pedestrians.
[0008] S4. Based on the global feature representation of pedestrians, a common dimension mask is generated by using frequency domain analysis and dynamic thresholding mechanism through the dynamic common dimension partitioning module.
[0009] S5. Construct a multi-objective joint loss function and complete feature constraints based on a shared dimension mask to realize an end-to-end dual-path parallel feature extraction structure and backbone network parameter optimization, thereby obtaining a visible light infrared pedestrian re-identification model to achieve visible light infrared pedestrian re-identification.
[0010] Furthermore, the preset sampling strategy in step S1 is as follows: for each batch, four pedestrians with different identities are selected in both the visible light mode and the infrared mode, and four images are sampled for each identity. At the same time, data augmentation operations such as random cropping, random horizontal flipping, and random erasure are performed on the training images to obtain the training batch data.
[0011] Furthermore, the background weakening module in step S2 includes an adaptive foreground segmentation submodule, a background structure destruction submodule, a texture suppression submodule, and an edge protection mechanism connected in sequence.
[0012] Furthermore, the data processing method for the adaptive foreground segmentation submodule is as follows: A1. The feature map F is obtained by extracting features from the input image X using a block encoder.
[0013] A2. Based on the feature map F, an initial logical value is generated through a mask prediction head, and the initial logical value is processed by the Sigmoid function to obtain the probability map S.
[0014] A3. Obtain the binarized mask B by hard-thresholding the probability map S using a learnable threshold τ.
[0015] A4. Perform morphological erosion on the binarized mask B to smooth the foreground region and remove isolated noise points, and upsample to the original input image size to obtain the foreground mask M.
[0016] Furthermore, the data processing method for the background structure destruction submodule is as follows: B1. Separate the foreground region from the input image X based on the foreground mask M. With background area ,in This indicates element-wise multiplication.
[0017] B2. Apply a learnable Gaussian blur to the entire input image X to obtain the blurred image. .
[0018] B3. To blurry image Structured noise is added to the background region to obtain an image with corrupted background regions. .
[0019] B4. Foreground area Image with background area destroyed Merge the images to obtain a complete image after the background structure has been destroyed. : .
[0020] Furthermore, the data processing method for the texture suppression submodule is as follows: C1. Extracting the complete image after background structure destruction using a small convolutional network. The texture feature map T.
[0021] C2. Based on texture features, calculate the texture energy map E and the high-texture region mask H: in This represents the number of channels in the texture feature map T. Indicates the first Feature map of each channel This represents the mean value of all pixels in the texture energy map E. Represents the mean function, This indicates an indicator function that takes the value 1 if the condition is true and 0 otherwise.
[0022] C3. Calculate the suppression factor S based on the high-texture region mask H: C4. By applying an inhibition factor S to the original pixels at high-texture locations in the background region to reduce texture interference, a texture-suppressed image is obtained: in This represents a texture-suppressed image.
[0023] Furthermore, the data processing method for the edge protection mechanism is as follows: D1. Predict the edge probability map from the feature map F through the edge prediction branch and upsample it to the original input image size.
[0024] D2. Based on the edge probability map, the foreground edge region of the texture-suppressed image is contrast-enhanced, and the output is an enhanced image with weakened background.
[0025] Furthermore, step S4 includes the following sub-steps: S41. Perform discrete Fourier transform on the visible light and infrared depth features of each batch in the global feature representation of pedestrians, transform them to the frequency domain and decompose them to obtain the amplitude spectrum. Calculate the average frequency importance weight vector of each batch of samples based on the amplitude spectrum to complete the frequency domain quantization of dimensional characteristics.
[0026] S42. Based on cross-modal distribution consistency and feature discriminative strength, calculate the potential score for each feature dimension.
[0027] S43. The Otsu method of maximum inter-class variance is used to automatically learn the optimal splitting threshold from the potential score distribution of the feature dimensions. The feature dimensions are divided into two categories: high-potential shared dimensions and low-potential private dimensions, so that the variance between the two categories is maximized. Dimensions with scores higher than the optimal splitting threshold are selected to form the initial set of shared dimensions.
[0028] S44. Set the shared dimension ratio range, perform ratio verification and adjustment on the initial shared dimension set, and output the binarized shared dimension mask.
[0029] Furthermore, the formula for the multi-objective joint loss function in step S5 is: in Represents the joint loss function for multiple objectives. Represents identity cross-entropy loss, Indicates the loss of the triplet. This represents the frequency-aware decoupling loss. This represents the cross-modal feature comparison regularization loss.
[0030] Furthermore, frequency-aware decoupling loss The calculation formula is: in This represents the shared representation loss across modalities. This represents the intramodal characterization enhancement loss.
[0031] Cross-modal shared representation loss Based on shared-dimensional masks, common subspace features of bimodal features are extracted. Cross-modal common feature alignment is achieved by constraining the diagonalization of the feature cross-correlation matrix, and intramodal representation enhancement loss is applied. By constraining the diagonalization of the correlation matrix between the original features and the background-weakened features, the consistency and discriminativeness of features within a modality are enhanced.
[0032] Cross-modal feature contrast regularization loss Using cosine distance as the metric, four sets of cross-modal feature pairs are constructed to form a comparison constraint. By minimizing the distance between cross-modal feature pairs with the same identity and maximizing the distance between cross-modal feature pairs with different identities, intra-class aggregation and inter-class separation are achieved. At the same time, a difficult negative sample screening mechanism is introduced to improve the model's ability to distinguish easily confused samples.
[0033] The beneficial effects of this invention are: (1) The present invention designs a background reduction module, which actively suppresses background noise at the input level through adaptive foreground segmentation, background structure destruction, texture suppression and edge protection techniques, guides the model to focus on the identity features of the pedestrian foreground, and solves the problem of feature learning bias caused by background interference.
[0034] (2) This invention proposes a dynamic common dimension partitioning module, which realizes adaptive decoupling of modal common features and private features based on frequency domain analysis, providing refined dimension guidance for cross-modal feature alignment and breaking through the limitations of fixed dimension partitioning.
[0035] (3) This invention constructs a joint loss function consisting of frequency-aware decoupling loss and cross-modal feature comparison regularization loss, and combines identity loss and triplet loss to achieve end-to-end optimization. At the same time, a hard negative sample mining strategy is introduced to achieve synergistic optimization of cross-modal feature alignment and discriminative enhancement.
[0036] (4) The present invention outperforms mainstream advanced methods on the two major public datasets RegDB and SYSU-MM01, and has stronger robustness and generalization ability in complex real monitoring scenarios. Attached Figure Description
[0037] Figure 1The diagram shown is a flowchart of the visible light infrared pedestrian re-identification method based on background attenuation and joint loss provided by an embodiment of the present invention.
[0038] Figure 2 The diagram shown is an overall architecture block diagram of the visible light infrared pedestrian re-identification method based on background attenuation and joint loss provided in an embodiment of the present invention.
[0039] Figure 3 The diagram shown is a structural block diagram of the background weakening module provided in an embodiment of the present invention.
[0040] Figure 4 The diagram shown is a structural block diagram of the adaptive foreground segmentation submodule provided in an embodiment of the present invention.
[0041] Figure 5 The diagram shown is a structural block diagram of the background structure destruction submodule provided in an embodiment of the present invention.
[0042] Figure 6 The diagram shown is a structural block diagram of the texture suppression submodule provided in an embodiment of the present invention.
[0043] Figure 7 The diagram shows the algorithm flowchart for the dynamic common dimension partitioning module provided in an embodiment of the present invention. Detailed Implementation
[0044] Exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be understood that the embodiments shown and described in the drawings are merely exemplary and are intended to illustrate the principles and spirit of the invention, and are not intended to limit the scope of the invention.
[0045] This invention provides a visible-infrared pedestrian re-identification method based on background attenuation and joint loss, such as... Figure 1 and Figure 2 As shown, the steps S1 to S5 are included: S1. Adjust the size of the input visible light pedestrian image and infrared pedestrian image to a preset size, and construct training batch data according to the preset sampling strategy.
[0046] In this embodiment of the invention, the preset size is 3×288×144, and the preset sampling strategy is as follows: for each batch, four pedestrians with different identities are selected in both the visible light mode and the infrared mode, and four images are sampled for each identity. At the same time, data augmentation operations such as random cropping, random horizontal flipping, and random erasure are performed on the training images to obtain training batch data.
[0047] S2. Construct a dual-path parallel feature extraction structure based on the basic feature extraction path of the original image and the enhanced feature path based on the background reduction module (SSM). Input the training batch data into the dual-path parallel feature extraction structure to extract multi-channel feature maps.
[0048] like Figure 3 As shown, the background reduction module includes an adaptive foreground segmentation submodule, a background structure destruction submodule, a texture suppression submodule, and an edge protection mechanism connected in sequence.
[0049] like Figure 4 As shown, the core objective of the adaptive foreground segmentation submodule is to generate a foreground mask M to accurately distinguish between pedestrian and background regions. Its data processing method is as follows: A1. The feature map F is obtained by extracting features from the input image X using a block encoder.
[0050] In this embodiment of the invention, the block encoder includes two convolutional layers. The first convolutional layer uses p×p convolutions with stride p to achieve small block embedding, followed by batch normalization and ReLU activation. The second convolutional layer is a 3×3 convolution, which further extracts local context information and ultimately divides the spatial dimensions into non-overlapping blocks. A small piece, in which and These represent the height and width of the input image, respectively.
[0051] A2. Based on the feature map F, an initial logical value is generated through a mask prediction head, and the initial logical value is processed by the Sigmoid function to obtain the probability map S.
[0052] A3. Obtain the binarized mask B by hard-thresholding the probability map S using a learnable threshold τ.
[0053] In this embodiment of the invention, the learnable threshold τ is dynamically predicted by a threshold learner based on global features.
[0054] A4. Perform morphological erosion on the binarized mask B to smooth the foreground region and remove isolated noise points, and upsample to the original input image size to obtain the foreground mask. .
[0055] like Figure 5 As shown, the core of the background structure destruction submodule is to eliminate the structured spatial information in the background, which is achieved by adding noise to Gaussian blur. The data processing method is as follows: B1. Separate the foreground region from the input image X based on the foreground mask M. With background area ,in This indicates element-wise multiplication.
[0056] In this embodiment of the invention, the foreground region It will be retained without any modification.
[0057] B2. Apply a learnable Gaussian blur to the entire input image X to obtain the blurred image. .
[0058] In this embodiment of the invention, the fuzzy kernel size k=7, and the standard deviation σ is a learnable parameter constrained between [0.5, 5.0].
[0059] B3. To blurry image Structured noise is added to the background region to obtain an image with corrupted background regions. .
[0060] In this embodiment of the invention, the noise includes low-frequency noise and high-frequency noise, which are combined in a certain proportion.
[0061] B4. Foreground area Image with background area destroyed Merge the images to obtain a complete image after the background structure has been destroyed. : .
[0062] like Figure 6 As shown, the data processing method of the texture suppression submodule is as follows: C1. Extracting the complete image after background structure destruction using a small convolutional network. The texture feature map T.
[0063] C2. Based on texture features, calculate the texture energy map E and the high-texture region mask H: in This represents the number of channels in the texture feature map T. Indicates the first Feature map of each channel This represents the mean value of all pixels in the texture energy map E. Represents the mean function, This indicates an indicator function that takes the value 1 if the condition is true and 0 otherwise. Pixels with a value of 1 in the high-texture region mask H correspond to high-texture background regions and will be suppressed in subsequent step C4.
[0064] C3. Calculate the suppression factor S based on the high-texture region mask H: C4. By applying an inhibition factor S to the original pixels at high-texture locations in the background region to reduce texture interference, a texture-suppressed image is obtained: in This represents a texture-suppressed image.
[0065] The data processing method for edge protection mechanisms is as follows: D1. Predict the edge probability map from the feature map F through the edge prediction branch and upsample it to the original input image size.
[0066] D2. Based on the edge probability map, the foreground edge region of the texture-suppressed image is contrast-enhanced, and the output is an enhanced image with weakened background.
[0067] In this embodiment of the invention, the basic feature extraction path directly performs shallow feature mapping on the input original image, fully preserving the global semantic information and original feature distribution of the image. The input image of the enhanced feature path is first processed by a background reduction module, effectively suppressing the interference of background noise on identity feature learning and enhancing the expressive ability of pedestrian local discriminative features. The visible light and infrared modal images from the training batch data are respectively input into these two paths, and the output four types of features are fused through multi-path feature fusion by element-wise addition to obtain a multi-channel feature map.
[0068] S3. Construct a weight-sharing backbone network based on the ResNet50 architecture, and input the multi-channel feature map into the shared feature extraction layer in the backbone network for deep feature extraction to obtain the global feature representation of pedestrians.
[0069] In this embodiment of the invention, the shared feature extraction layer adopts a hierarchical convolutional neural network structure with fully shared weights to perform unified deep feature extraction on the fusion features of visible light and infrared modes, and outputs a 2048-dimensional global feature representation of pedestrians.
[0070] S4. Based on the global feature representation of pedestrians, a common dimension mask is generated by using frequency domain analysis and dynamic thresholding mechanism through the dynamic common dimension partitioning module.
[0071] In this embodiment of the invention, the core of the dynamic shared dimension partitioning module is to adaptively identify shared feature dimensions with high discriminative power among modalities, thereby achieving precise decoupling between shared and private features of modalities. For example... Figure 7 As shown, step S4 includes the following sub-steps S41 to S44: S41. Perform discrete Fourier transform on the visible light and infrared depth features of each batch in the global feature representation of pedestrians, transform them to the frequency domain and decompose them to obtain the amplitude spectrum. Calculate the average frequency importance weight vector of each batch of samples based on the amplitude spectrum to complete the frequency domain quantization of dimensional characteristics.
[0072] In this embodiment of the invention, the amplitude spectrum reflects the energy intensity of the feature at each frequency component.
[0073] S42. Based on cross-modal distribution consistency and feature discriminative strength, calculate the potential score for each feature dimension.
[0074] In this embodiment of the invention, cross-modal distribution consistency is calculated by cosine similarity of visible light and infrared feature amplitude spectra, and feature discriminative power is calculated by arithmetic mean of bimodal amplitude spectra. The final potential score of the feature dimension is the sum of the two, ensuring that the high-scoring dimension simultaneously satisfies the two conditions of strong cross-modal consistency and high self-discrimination power.
[0075] S43. The Otsu method of maximum inter-class variance is used to automatically learn the optimal splitting threshold from the potential score distribution of the feature dimensions. The feature dimensions are divided into two categories: high-potential shared dimensions and low-potential private dimensions, so that the variance between the two categories is maximized. Dimensions with scores higher than the optimal splitting threshold are selected to form the initial set of shared dimensions.
[0076] S44. Set the shared dimension ratio range, perform ratio verification and adjustment on the initial shared dimension set, and output the binarized shared dimension mask.
[0077] like Figure 7 As shown, when the initial shared dimension ratio exceeds a preset range, the boundary value of the preset range is used as the target ratio, and finally, a binarized shared dimension mask is output. In this embodiment of the invention, the lower limit of the shared dimension ratio is set to 0.4 and the upper limit is set to 0.6 to ensure that the partitioning results are within a reasonable and stable range.
[0078] S5. Construct a multi-objective joint loss function and complete feature constraints based on a shared dimension mask to realize an end-to-end dual-path parallel feature extraction structure and backbone network parameter optimization, thereby obtaining a visible light infrared pedestrian re-identification model to achieve visible light infrared pedestrian re-identification.
[0079] In this embodiment of the invention, the formula for the multi-objective joint loss function is: in Represents the joint loss function for multiple objectives. Represents identity cross-entropy loss, Indicates the loss of the triplet. This represents the frequency-aware decoupling loss. This represents the cross-modal feature comparison regularization loss.
[0080] Identity cross-entropy loss Each pedestrian is categorized into a separate class to ensure accurate pedestrian identification.
[0081] Triple loss This is used to reduce the distance between the anchor point and positive samples, while increasing the distance between the anchor point and negative samples, thereby enhancing the intra-class aggregation and inter-class separation capabilities of features.
[0082] Frequency-aware decoupling loss The calculation formula is: in This represents the shared representation loss across modalities. This represents the intramodal characterization enhancement loss.
[0083] Among them, cross-modal shared representation loss Based on common dimension mask The common subspace features of the bimodal model are extracted. By constraining the diagonalization of the feature cross-correlation matrix, the diagonal elements of the matrix approach 1, and the off-diagonal elements approach 0, thus achieving cross-modal common feature alignment. The specific calculation method involves using a common dimension mask. Original visible light characteristics Original infrared features Visible light characteristics after background attenuation Infrared features after background attenuation Perform element-wise multiplication to obtain the corresponding common subspace features. , , , Calculate the normalized cross-correlation matrices for the original feature pairs and the background-reduced feature pairs, respectively: in This represents the normalized cross-correlation matrix of the original feature pairs. This represents the normalized cross-correlation matrix of feature pairs after background weakening, with the superscript indicating the normalized cross-correlation matrix. To represent the transpose of a matrix, Indicates the batch size.
[0084] For the normalized cross-correlation matrix and Perform diagonalization constraints: Finally, the cross-modal shared representation loss is obtained by merging the results. : in This represents the result of diagonalization constraint on the original feature pairs. This represents the diagonalization constraint result of the feature pairs after background weakening. This indicates the preset balance hyperparameters. 、 All are matrix indices.
[0085] Intramodal characterization enhancement loss By constraining the diagonalization of the correlation matrix between the original features and the background-weakened features, the consistency and discriminativeness of intramodal features are enhanced, preventing the collapse of private dimensions during training. The specific calculation method is as follows: Construct four feature pair combinations, namely , , , Calculate the normalized cross-correlation matrix for each combination: in Indicates the first The normalized cross-correlation matrix of each combination, This represents the first feature pair in the combination. This represents the second feature pair in the combination.
[0086] Apply diagonalization constraints to each cross-correlation matrix: in Indicates the first The diagonalization constraint results of each combination This represents the preset equilibrium hyperparameters. The final result is the intra-modal characterization enhancement loss. : Cross-modal feature contrast regularization loss Using cosine distance as the metric, four sets of cross-modal feature pairs are constructed as contrast constraints. The four sets of cross-modal feature pairs are: original visible light feature - original infrared feature. Visible light features after background attenuation - original infrared features Original visible light features - infrared features after background attenuation Visible light features after background attenuation - Infrared features after background attenuation .
[0087] For the Let there be a combination, and let the total loss of all its anchor-negative sample pairs be... The corresponding total logarithm is Then the average loss of this combination for: Cross-modal feature contrast regularization loss The calculation formula is: in To balance the hyperparameters, the default value is 1.0.
[0088] By minimizing the distance between cross-modal feature pairs with the same identity and maximizing the distance between cross-modal feature pairs with different identities, intra-class aggregation and inter-class separation are achieved. At the same time, a difficult negative sample screening mechanism is introduced, which only retains negative samples with a cosine similarity greater than -0.1 with the anchor point, so that the model can focus more on processing difficult negative samples that are hard to distinguish and improve the model's ability to distinguish easily confused samples.
[0089] The technical effects of the present invention will be further explained in detail below through specific experimental examples.
[0090] The experimental examples of this invention were conducted using the PyTorch framework, with an RTX 4060ti 16GB GPU. Model training lasted for 160 epochs. The SGD optimizer was used, with a weight decay of 5 × 10⁻⁶. -4 The momentum term was set to 0.9; a dynamic scheduling strategy was adopted for the learning rate. For the RegDB dataset, the learning rate linearly increased from 0 to 0.1 in the first 30 iterations, and then decreased by 0.1 at the 80th and 120th iterations, respectively. For the SYSU-MM01 dataset, the learning rate linearly increased to 0.1 in the first 20 iterations, and then decreased by 0.1 at the 20th, 80th, and 120th iterations, respectively. The threshold for triplet loss was set to 0.3, and the input images were uniformly scaled to 3×288×144.
[0091] The experimental examples of this invention were verified on two public visible light infrared cross-modal person re-identification datasets, RegDB and SYSU-MM01. Following the standard evaluation protocol, Rank-n (Rank-1, Rank-10, Rank-20), mAP, and mINP were used as core evaluation indicators to comprehensively evaluate the cross-modal retrieval performance of the model.
[0092] Table 1. Ablation experiments using the model in infrared retrieval of visible light on the RegDB dataset (%) Table 2. Model performance (%) on the RegDB dataset. Table 3. Model performance (%) in indoor and outdoor evaluations on the SYSU-MM01 dataset. As shown in Table 1 of the ablation experiment results, compared with the baseline method, the introduction of the background attenuation module (SSM) alone improved the Rank-1 accuracy by 3.6 percentage points and the mAP by 4.0 percentage points, verifying that the background attenuation module can effectively suppress background interference and enhance the discriminativeness of features. Frequency-aware decoupling loss was then introduced on top of the SSM module. Regularization loss compared to cross-modal features Subsequently, the model performance was further improved, proving that the two loss functions can effectively enhance the cross-modal feature alignment effect; when all modules are used together, the model performance reaches the optimal level, with Rank-1 accuracy reaching 97.91% and mAP reaching 91.70%, which fully verifies the synergistic optimization effect of each module.
[0093] Subsequently, the performance of the method of this invention was compared with that of current mainstream advanced methods on the SYSU-MM01 and RegDB datasets. The results are shown in Tables 2 and 3. In the visible light to infrared retrieval task of the RegDB dataset, the Rank-1 accuracy of the method of this invention reached 97.91%, and the mAP reached 91.70%, which is 6.31 percentage points and 7.60 percentage points higher than the second-best MMN method, respectively. In the infrared to visible light retrieval task, the Rank-1 accuracy of the method of this invention reached 94.51%, and the mAP reached 88.32%, which is also significantly better than the existing mainstream methods. On the more challenging SYSU-MM01 dataset, the method of this invention achieves a Rank-1 accuracy of 76.13% and an mAP of 78.70% in indoor scenes, and a Rank-1 accuracy of 70.23% and an mAP of 64.34% in outdoor scenes. Both of these results surpass mainstream advanced methods such as RAT, MCLNet, and PMWGCN, fully demonstrating the robustness and advancement of the method in complex scenes. It can effectively solve the problem of cross-modal feature alignment deviation caused by background interference and achieve high-precision visible light and infrared cross-modal pedestrian re-identification.
[0094] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A visible-infrared pedestrian re-identification method based on background attenuation and joint loss, characterized in that, Includes the following steps: S1. Adjust the size of the input visible light pedestrian image and infrared pedestrian image to a preset size, and construct training batch data according to the preset sampling strategy; S2. Construct a dual-path parallel feature extraction structure based on the basic feature extraction path of the original image and the enhanced feature path based on the background weakening module, and input the training batch data into the dual-path parallel feature extraction structure to extract multi-channel feature maps; S3. Construct a weight-sharing backbone network based on the ResNet50 architecture, and input the multi-channel feature map into the shared feature extraction layer in the backbone network for deep feature extraction to obtain the global feature representation of pedestrians; S4. Based on the global feature representation of pedestrians, a common dimension mask is generated by frequency domain analysis and dynamic thresholding mechanism through a dynamic common dimension partitioning module. S5. Construct a multi-objective joint loss function and complete feature constraints based on a shared dimension mask to realize an end-to-end dual-path parallel feature extraction structure and backbone network parameter optimization, thereby obtaining a visible light infrared pedestrian re-identification model to achieve visible light infrared pedestrian re-identification.
2. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 1, characterized in that, The preset sampling strategy in step S1 is as follows: for each batch, four pedestrians with different identities are selected in both the visible light mode and the infrared mode, and four images are sampled for each identity. At the same time, data augmentation operations such as random cropping, random horizontal flipping, and random erasure are performed on the training images to obtain training batch data.
3. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 1, characterized in that, The background weakening module in step S2 includes an adaptive foreground segmentation submodule, a background structure destruction submodule, a texture suppression submodule, and an edge protection mechanism connected in sequence.
4. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 3, characterized in that, The data processing method of the adaptive foreground segmentation submodule is as follows: A1. The feature map F is obtained by extracting features from the input image X using a block encoder; A2. Based on the feature map F, an initial logical value is generated through a mask prediction head, and the initial logical value is processed by the Sigmoid function to obtain the probability map S; A3. Obtain the binarized mask B by performing hard thresholding on the probability map S using a learnable threshold τ. A4. Perform morphological erosion on the binarized mask B to smooth the foreground region and remove isolated noise points, and upsample to the original input image size to obtain the foreground mask M.
5. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 4, characterized in that, The data processing method for the background structure destruction submodule is as follows: B1. Separate the foreground region from the input image X based on the foreground mask M. With background area ,in This indicates element-wise multiplication; B2. Apply a learnable Gaussian blur to the entire input image X to obtain the blurred image. ; B3. To blurry image Structured noise is added to the background region to obtain an image with corrupted background regions. ; B4. Foreground area Image with background area destroyed Merge the images to obtain a complete image after the background structure has been destroyed. : 。 6. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 5, characterized in that, The data processing method of the texture suppression submodule is as follows: C1. Extracting the complete image after background structure destruction using a small convolutional network. Texture feature map T; C2. Based on texture features, calculate the texture energy map E and the high-texture region mask H: in This represents the number of channels in the texture feature map T. Indicates the first Feature map of each channel This represents the mean value of all pixels in the texture energy map E. Represents the mean function, This indicates an indicator function that takes the value 1 if the condition is true and 0 otherwise. C3. Calculate the suppression factor S based on the high-texture region mask H: C4. By applying an inhibition factor S to the original pixels at high-texture locations in the background region to reduce texture interference, a texture-suppressed image is obtained: in This represents a texture-suppressed image.
7. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 6, characterized in that, The data processing method for the edge protection mechanism is as follows: D1. Predict the edge probability map from the feature map F through the edge prediction branch and upsample it to the original input image size; D2. Based on the edge probability map, the foreground edge region of the texture-suppressed image is contrast-enhanced, and the output is an enhanced image with weakened background.
8. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 1, characterized in that, Step S4 includes the following sub-steps: S41. Perform discrete Fourier transform on each batch of visible light and infrared depth features in the global feature representation of pedestrians, convert to the frequency domain and decompose to obtain the amplitude spectrum. Calculate the average frequency importance weight vector of each batch of samples based on the amplitude spectrum to complete the frequency domain quantization of dimensional characteristics. S42. Based on cross-modal distribution consistency and feature discriminative strength, calculate the potential score for each feature dimension; S43. The Otsu maximum inter-class variance method is used to automatically learn the optimal splitting threshold from the potential score distribution of the feature dimensions. The feature dimensions are divided into two categories: high-potential shared dimensions and low-potential private dimensions, so that the variance between the two categories is maximized. Dimensions with scores higher than the optimal splitting threshold are selected to form the initial set of shared dimensions. S44. Set the shared dimension ratio range, perform ratio verification and adjustment on the initial shared dimension set, and output the binarized shared dimension mask.
9. The visible light infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 1, characterized in that, The formula for the multi-objective joint loss function in step S5 is as follows: in Represents the joint loss function for multiple objectives. Represents identity cross-entropy loss, Indicates the loss of the triplet. This represents the frequency-aware decoupling loss. This represents the cross-modal feature comparison regularization loss.
10. The visible-infrared pedestrian re-identification method based on background attenuation and joint loss according to claim 9, characterized in that, The frequency-aware decoupling loss The calculation formula is: in This represents the shared representation loss across modalities. This represents the intramodal characterization enhancement loss; The cross-modal shared representation loss Based on the extraction of shared subspace features of the bimodal model using a shared-dimensional mask, cross-modal shared feature alignment is achieved by constraining the diagonalization of the feature cross-correlation matrix. The intra-modal representation enhancement loss... By constraining the diagonalization of the correlation matrix between the original features and the background-weakened features, the consistency and discriminativeness of features within a modality are enhanced. The cross-modal feature contrast regularization loss Using cosine distance as the metric, four sets of cross-modal feature pairs are constructed to form a comparison constraint. By minimizing the distance between cross-modal feature pairs with the same identity and maximizing the distance between cross-modal feature pairs with different identities, intra-class aggregation and inter-class separation are achieved. At the same time, a difficult negative sample screening mechanism is introduced to improve the model's ability to distinguish easily confused samples.