Pedestrian re-identification model training method based on adaptive modal perception
By training an adaptive modality-aware person re-identification model, auxiliary modality images are generated and feature alignment loss is used to solve the semantic conflict problem in VI-ReID, achieving efficient person re-identification under all-weather conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-04-14
AI Technical Summary
Existing VI-ReID methods are prone to semantic conflicts when generating auxiliary modalities, which increases the difficulty of model learning and limits the improvement of recognition performance, especially in low-light or complex nighttime scenes where performance degrades.
By training an adaptive modal perception pedestrian re-identification model, an auxiliary modal image is generated using a spatial attention mechanism and an adaptive hierarchical replacement module. The alignment of visible light and infrared image features is achieved by combining symmetric cross-modal consistency loss and cascaded aggregation loss.
It effectively compensates for missing feature information in cross-modal data, reduces modal differences, and improves recognition performance, especially significantly improving the accuracy of pedestrian re-identification under all-weather conditions.
Smart Images

Figure CN121861691A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically to a method for training a pedestrian re-identification model based on adaptive modal perception. Background Technology
[0002] Person re-identification (ReID) is an important task in the field of computer vision, aiming to accurately match different images of the same pedestrian across different cameras and scenes. In recent years, with the widespread application of deep learning technology, person re-identification has made significant progress in matching performance.
[0003] However, traditional single-modal ReID methods primarily rely on visible light cameras, and their performance deteriorates sharply in low-light or complex nighttime scenes due to severe image texture degradation. To address this issue, Visible-Infrared Person Re-Identification (VI-ReID) technology has emerged, which addresses all-weather identification needs by searching between two heterogeneous modalities: visible light and infrared. The VI-ReID task not only needs to overcome intra-class discrepancies in traditional ReID but also needs to bridge the significant modal gap caused by different imaging mechanisms.
[0004] Despite the progress made in VI-ReID technology, it still faces many challenges. Existing VI-ReID methods, when generating auxiliary modalities, such as generating grayscale images from a single source or randomly mixing pixels from two sources, are prone to phenomena such as double contours in the generated images (e.g., visible light texture and infrared thermal boundary simultaneously existing on the human body edge), causing semantic conflicts between modalities. This, in turn, increases the learning difficulty of the model and limits further improvement in recognition performance. Summary of the Invention
[0005] Based on this, the present invention provides a method and device for training a pedestrian re-identification model based on adaptive modal perception, which solves at least one problem in the prior art.
[0006] In a first aspect, the present invention provides a method for training a person re-identification model based on adaptive modal perception, which includes the following steps: Generate auxiliary modal images based on visible light and infrared images; Features are extracted from the visible light image, infrared image, and auxiliary modal image respectively to obtain the features of the visible light image, the infrared image, and the auxiliary modal image; The pedestrian re-identification model is trained using features from visible light images, infrared images, and auxiliary modal images. Update the parameters of the pedestrian re-identification model based on the target loss; The generation of auxiliary modal images based on visible light and infrared images includes: Visible light images and infrared images are divided into... A semantically consistent horizontal stripe, among which It is a positive integer greater than 1; Based on the probability of replacement Horizontal stripes from either a visible light image or an infrared image are selected to generate the auxiliary modal image.
[0007] In some optional embodiments, the replacement probability Calculated by the following formula: in, Indicates the first The probability of replacing a horizontal stripe. Indicates the first Prior factors for the human body parts corresponding to each horizontal stripe. ; Indicates the first Normalized significance weights of horizontal stripes; This represents a non-zero constant.
[0008] In some optional embodiments, the strategy for generating auxiliary modal images is as follows: in, Represents auxiliary modal images. This indicates an image stitching operation; Indicates the infrared image number 1 The image region corresponding to each stripe Represents the visible light image. The image area corresponding to each stripe; Let represent a random variable that is uniformly distributed on the interval [0,1].
[0009] In some alternative embodiments, the first The normalized significance weights of the horizontal stripes are calculated by the following formula: in, This represents the smallest significance weight among all horizontal stripes. This represents the one with the largest significance weight among all horizontal stripes; Indicates the first The significance weight of each stripe is calculated by the following formula: Indicates the height of the horizontal stripes. Indicates the width of the feature map; Indicates the first One horizontal stripe; Spatial attention map exist Location values, spatial attention map Calculated by the following formula: This represents the global max pooling result across the channel dimensions of the input feature tensor. This represents the result of global average pooling across the channel dimensions of the input feature tensor. This indicates a splicing operation. This represents a two-dimensional convolution with a kernel size of 7. This indicates that the Sigmoid function is activated.
[0010] In some optional embodiments, prior factors for human body parts are determined based on the body parts themselves. Preferably, the human body parts are divided into the upper body and the lower body, with the prior factors for the upper body being smaller than those for the lower body.
[0011] In some alternative embodiments, The prior factor is 4, the prior factor for the upper body is 0.3, and the prior factor for the lower body is 0.7.
[0012] In some optional embodiments, the target loss includes symmetric cross-modal consistency loss (SCCL), wherein the symmetric cross-modal consistency loss includes symmetric cross-modal consistency loss for generating auxiliary modal images based on visible light images. Symmetric cross-modal consistency loss based on generating auxiliary modal images from infrared images As shown in the following formula: in, This indicates the number of labeled categories in the training batch. This indicates the number of image instances for each identifier; Represents the L2 norm; Indicates the first The first identity (i.e., the same identity identifier) under the first identity (i.e., the same identity identifier) Visible light image features, Indicates the first The first identity under the th Infrared image features, Indicates the first The first identity under the th Auxiliary image features; This indicates the control parameter.
[0013] In some optional embodiments, the target loss also includes a cascaded aggregation loss (CAL), as shown below: in, The identity index (pedestrian's index number) is... Visible light image feature center, ; Indicates identity index is The infrared image feature center, ; This represents the semantic center of the visible light image obtained through a two-layer fully connected network. ; This represents the semantic center of the infrared image extracted through a two-layer fully connected network. ; This represents a two-layer fully connected network.
[0014] In some alternative embodiments, the target loss also includes identity loss and triplet loss.
[0015] Secondly, the present invention provides a pedestrian re-identification method, which includes the following steps: The visible light image and infrared image to be identified are input into the pedestrian re-identification model, and the pedestrian re-identification model outputs the identification result. The pedestrian re-identification model is obtained by the above-mentioned pedestrian re-identification model training method based on adaptive modal perception.
[0016] Thirdly, the present invention provides a pedestrian re-identification device, comprising: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that, when executed by at least one processor, implement the pedestrian re-identification method described above.
[0017] Fourthly, the present invention provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the above-described pedestrian re-identification method.
[0018] Due to the adoption of the above technical solutions, the embodiments of the present invention have at least the following beneficial effects: Guided by human anatomy, the replacement probability and prior factors of different human body parts are determined, and auxiliary modal images are generated accordingly, which effectively compensates for the missing feature information in cross-modal data. Symmetric cross-modal consistency loss (SSCL) and cascaded aggregation loss (CAL) are used to align visible light and infrared images in the feature space indirectly and directly. SSCL uses an auxiliary mode as an intermediate bridge to enable bidirectional knowledge flow between visible light and infrared modes, ultimately achieving the goal of simultaneous alignment of dual semantics (visible light-auxiliary and infrared-auxiliary) in the feature space. CAL further aligns the features of visible light and infrared images across multiple scales, significantly reducing modal differences and making the establishment of spectral correspondences more direct and efficient. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the network structure of a pedestrian re-identification model in one embodiment of the present invention.
[0020] Figure 2 This is a schematic diagram of the network structure of the AHRM module in one embodiment of the present invention.
[0021] Figure 3 The figure shows the verification results of the AMAF model's ability to mitigate modal differences in one embodiment of the present invention. (a) is the result obtained before model training, and (b) is the result after model training. inter-class represents the inter-class distance, and intra-class represents the intra-class distance. Detailed Implementation
[0022] The following will provide a clear and complete description of the concept and technical effects of the present invention, so as to fully explain the purpose, solution and effects of the present invention.
[0023] To address the issue of poor pedestrian matching performance in low-light or no-light environments using traditional visible light pedestrian re-identification, this invention presents a pedestrian re-identification model training method based on adaptive modal perception. Specifically, a pedestrian re-identification model is trained, and pedestrian re-identification is achieved through this model. This pedestrian re-identification model is also known as an adaptive modal perception fusion (AMAF) model.
[0024] like Figure 1 As shown, when training the pedestrian re-identification model, the visible light image is first processed through a spatial attention mechanism, and then an auxiliary modality image is generated based on the visible light image and the infrared image. Features of the visible light image, the infrared image, and the auxiliary modality image are extracted respectively to obtain the features of the visible light image. Features of infrared images Features of auxiliary modal images ; Using features of visible light images Features of infrared images Features of auxiliary modal images Train the pedestrian re-identification model; and update the parameters of the pedestrian re-identification model based on the target loss.
[0025] Specifically, spatial attention is applied through a lightweight spatial attention module. This module takes the original image as input and dynamically generates an attention weight map of size [1, H, W] using a dual-dimensional channel and spatial attention mechanism. High-response regions precisely correspond to highly discriminative semantic features (such as facial contours and personalized clothing textures). The spatial attention mechanism quantifies the discriminative power of different spatial locations in a visible light image, generating a pixel-level saliency map (spatial attention map) as a prior for subsequent mixing strategies. Specifically, given a visible light image… Convert it into a feature tensor In one embodiment of the present invention, C=3 is used; global max pooling is performed by fusing the channel dimension. Compared with global average pooling Obtain spatial attention map As shown in the following formula: in, Indicates the number of channels. Indicates the height of the feature map, Indicates the width of the feature map; Indicates the first Feature maps of each channel; This represents the output feature map of global max pooling. This represents the output feature map of global average pooling. This indicates a splicing operation. This represents a two-dimensional convolution with a kernel size of 7. M represents the activation of the Sigmoid function; M represents the spatial attention map.
[0026] Subsequently, an auxiliary modal image is generated using the Adaptive Hierarchical Replacement Module (AHRM). Specifically, the visible light image and the infrared image are divided into multiple semantically consistent horizontal stripes. Then, a portion of the horizontal stripes in the visible light image is replaced with corresponding horizontal stripes in the infrared image, or vice versa. Figure 2 As shown, the visible light image and the infrared image are uniformly divided along the height direction into two parts. Horizontal stripes The height of each stripe is . No. Significance weight of each stripe The mean attention value within this region is defined as follows: To avoid the influence of extreme values, the significance weights are further normalized to [0,1], as shown in the following formula: The normalized weight Used to guide the replacement probability of horizontal stripes when generating cross-modal images.
[0027] AHRM divides the image into semantically consistent horizontal stripes and, based on the principle of "high saliency → low replacement probability, low saliency → high replacement probability," achieves fine-grained blending for discriminative region protection and non-discriminative region enhancement. In one embodiment of the invention, K=4 (i.e., 4 horizontal stripes), each stripe... The set of covered pixels is shown in the following formula: This division corresponds to prior knowledge of the human body: and Roughly covering the head and torso, and Cover the thigh and foot (shoe) area. Weight the significance level. with human a priori After combination, the first The probability of replacing a horizontal stripe is: in, This represents the prior factors for human body parts. Prior factors for human body parts are set based on these parts. The upper body region (k=1,2) contains more discriminative features (such as the head and torso), therefore a more conservative setting is used. To reduce the probability of replacement while preserving visible light features; the lower body region (k=3,4) has fewer discriminative features, so a more enhanced [feature] is set. This increases the probability of feature replacement. Represents the minimum noise term, preventing hour This leads to a loss of training diversity. In one embodiment of the invention, In one embodiment of the present invention, Initial settings: Replacement probability It is determined by two parts: prior factors of human body parts. Significance weights (Region importance). This approach leverages prior human knowledge while also considering adaptive adjustments to image content. The final hybrid strategy for the generated auxiliary modalities can be expressed as: in, Represents auxiliary modal images. This indicates an image stitching operation; Indicates the infrared image number 1 The image region corresponding to each stripe Represents the visible light image. The image area corresponding to each stripe; Let represent a random variable uniformly distributed in the interval [0,1], used for decision-making. This means that for each stripe k, the infrared or visible light portion is selected based on the conditions.
[0028] AHRM has the following advantages: First, the generated auxiliary modal pedestrian semantics are complete, natural, and clear; second, the generation process is simple and does not increase the complexity of the model; third, the replacement strategy forms a closed-loop control of protecting important regions and enhancing secondary regions, almost avoiding damage to these regions. AHRM introduces an upper-lower body asymmetric processing mechanism for the first time. For the upper body region (including key identity information such as the face and chest), a low replacement probability is set, prioritizing the preservation of fine-grained texture information from the visible light modality. This is because high-frequency information such as faces and clothing patterns has significant discriminative power in the visible light modality, while the infrared modality often suffers from information loss in this region due to uneven thermal radiation. A high replacement probability is set for the lower body region (low-frequency structures such as legs and shoes), actively introducing contour features from the infrared modality.
[0029] Using the auxiliary modal image generated by AHRM as a key intermediate medium, the auxiliary modality cleverly decomposes the large modal gap between the original visible and infrared regions into two relatively smaller sub-gap: the visible-auxiliary sub-gap and the auxiliary-infrared sub-gap. This gap decomposition strategy has significant advantages, making it easier to accurately capture and effectively extract modal sharing cues in any feature subspace. Smaller sub-gap means tighter correlations between different modalities and smoother information transmission, thus providing more favorable conditions for the model to mine common features between modalities.
[0030] Considering that the generated auxiliary modes still have significant modal differences from the visible and infrared modes, this invention uses symmetric cross-modal consistency loss (SCCL) to achieve visible-infrared channel alignment, thereby indirectly mitigating the modal differences.
[0031] The visible-auxiliary stage symmetric cross-modal consistency loss (SCCL) can be expressed as: Similarly, the infrared-assisted phase symmetric cross-modal consistency loss (SCCL) can be expressed as: in, This indicates the number of labeled categories in the training batch. This indicates the number of image instances for each identifier; Represents the L2 norm; Indicates the first The first identity (i.e., the same identity identifier) under the first identity (i.e., the same identity identifier) Visible light image features, Indicates the first The first identity under the th Infrared image features, Indicates the first The first identity under the th Auxiliary image features; This represents the control parameter, which is responsible for controlling the relative proportion of knowledge transfer between strongly correlated sample pairs and weakly correlated sample pairs. It can be defined as: in, For hyperparameters, and These represent the indices for visible light and auxiliary modal images, respectively.
[0032] The overall SSCL expression can be represented as: in, This represents the balance parameter, used for control. . contributions.
[0033] To further eliminate modal differences, this invention also uses Cascaded Aggregation Loss (CAL) to directly align visible and infrared features. On the one hand, it narrows the feature distance between samples of the same identity at the instance level; on the other hand, it imposes constraints on the global feature centers of visible and infrared light at the distribution level, using the identity center as a metric, thereby achieving a progressive alignment from local to global. CAL can be represented as follows: in, Indicates identity index is Visible light image feature center, ; Indicates identity index is The infrared image feature center, ; This represents the semantic center of the visible light image obtained through a two-layer fully connected network. ; This represents the semantic center of the infrared image extracted through a two-layer fully connected network. ; This represents a two-layer fully connected network used to extract latent semantics, enabling the model to more easily construct cross-modal associations.
[0034] In addition to SSCL loss ( ) and CAL loss ( In addition to identity loss, and batch hard ternary group loss Let's work together to optimize AMAF, using identity loss. And batch hard triplet loss The baseline loss is used to learn the discriminative global features. Therefore, the overall objective loss of AMAF can be expressed as: in, and Represents the global loss term. and The coefficient of relative importance.
[0035] After training the pedestrian re-identification model, features from the query image and the image library are extracted, cosine similarity is calculated, and the re-identification results are displayed in order of cosine similarity. Using a pedestrian re-identification model is an existing technique and will not be elaborated upon here.
[0036] To verify the effectiveness of the pedestrian re-identification model (AMAF model) in this embodiment of the invention, experiments were conducted on the SYSU-MM01 dataset and the RegDB dataset, and a systematic comparison was made with the existing VI-ReID model.
[0037] SYSU-MM01 is a multimodal person re-identification dataset containing data on 491 different identities, totaling 44,745 images, of which 29,033 are visible light images and 15,712 are infrared images. This data was acquired by four visible light cameras and two infrared cameras in mixed indoor and outdoor scenes. The training set covers 395 identities, containing 22,258 visible light images and 11,909 infrared images. In the testing phase, the query set contains 3,803 infrared images of 96 identities, while the image library randomly selects either 301 (single-lens) or 3,010 (multi-lens) visible light images for retrieval. This dataset provides two evaluation modes: a full search mode that examines cross-modal matching capabilities in both indoor and outdoor scenes, and an indoor search mode that focuses on person re-identification performance in indoor environments.
[0038] RegDB is a cross-modal person re-identification dataset containing data on 412 different pedestrians, with 10 visible light images and 10 infrared images collected for each identity, totaling 8,240 images. The dataset uses a dual-camera system to simultaneously acquire visible light and infrared modal data. For data partitioning, a standard segmentation strategy was followed, randomly selecting 206 identities (4,120 images) as the training set, and the remaining 206 identities (4,120 images) for testing. The dataset incorporates a bidirectional cross-modal retrieval task: in the visible light to infrared (VIS→IR) retrieval mode, the system needs to match the queried visible light image from the infrared image library; while in the infrared to visible light (IR→VIS) mode, the retrieval task is performed in reverse, i.e., the infrared image is used to query and match the visible light image library.
[0039] The experiment was implemented using PyTorch on a GeForce RTX 3080ti GPU. For the network architecture, a ResNet50 pre-trained on ImageNet was used as the backbone, with the stride of the last layer set to 1. The training input images were uniformly adjusted to 3×288×144, and data augmentation techniques such as random horizontal flipping, random cropping, random erasing, and color dithering were employed during training. The learning threshold was initially set to 1×10. -2 The first 10 rounds of linear heating to 1×10 -1 The weights were then decayed by 10% at 20 and 60 epochs. Each batch consisted of 64 image samples (4 visible light images + 4 infrared images, covering 8 identities), and the training process lasted for 110 epochs. The optimizer used stochastic gradient descent (SGD) with a momentum of 0.9 and a weight decay of 5 × 10⁻⁶. -4 The hyperparameters are set to... , , , .
[0040] As shown in Table 1, the AMAF model in this invention achieves significant performance using a simple network structure. Specifically, in All-search mode, the AMAF model achieves a Rank-1 of 75.06% and an mAP of 71.40%. In Indoor-Search mode, the Rank-1 and mAP performance of the AMAF model are improved to 81.97% and 84.72%, respectively.
[0041] Table 1. Test results (%) of different models on the SYSU-MMO1 dataset In the table, Rank-k represents the probability of containing a correctly matching image in the top k search results. For example, Rank-5 or Rank-10 represents the proportion of the correct identity contained in the top 5 or top 10 results. mAP represents the mean precision (AP) over all query images, reflecting the model's overall performance on the ranked list.
[0042] As shown in Table 2, for the VIS-IR mode, the AMAF model achieved a Rank-1 score of 91.69% and an mAP of 88.15%. Furthermore, for the IR-VIS mode, the AMAF model achieved a Rank-1 score of 90.16% and an mAP of 86.95%. The comparison results demonstrate that the AMAF model exhibits superior performance compared to other models on the RegDB dataset.
[0043] Table 2. Test results (%) of different models on the RegDB dataset To verify the effectiveness of each component of the AMAF model, ablation experiments were conducted on the SYSU-MM01 dataset in All-search mode. A two-stream base network structure (i.e., the base, which could be a ResNet50 network) was used as the baseline model for the ablation experiments. As shown in Table 3, by comparing the baseline model with the model that incorporates the AHRM module without spatial attention (base+AHRM*), significant improvements in both Rank-1 and mAP metrics can be observed. When the spatial attention mechanism is then introduced into AHRM (base+AHRM), further improvements in Rank-1 and mAP metrics are observed. This fully demonstrates that the auxiliary modal images generated by AHRM can effectively correlate visible light and infrared modalities, providing strong support for cross-modal feature learning. Furthermore, when SSCL loss (base+AHRM+SSCL) or a combination of SSCL and CAL loss components (base+AHRM+SSCL+CAL) is incorporated into the framework, performance is significantly improved again; this strongly demonstrates the effectiveness of SSCL and CAL losses in enhancing intra-identity similarity and inter-identity differences under the guidance of auxiliary representations. Ablation experiments fully demonstrate that each module makes a positive contribution to the overall performance improvement. In fact, the organic combination of the modules in the AMAF model enables it to exhibit performance that is competitive with most visible-infrared person re-identification (VI-ReID) models.
[0044] Table 3 Ablation Experiment Results To further verify the advantages of the AMAF model compared to other auxiliary modality generation modules, a comparative analysis was conducted on the auxiliary modalities generated by Patch-Mixed, Global-Mixed, CutMix, SCSA, and PedMixed. In practice, the AHRM module in the AMAF model was replaced with other auxiliary modality generation modules, and experimental verification was performed on the SYSU-MM01 dataset. As shown in Table 4, the AHRM module outperforms all other auxiliary modality generation modules, fully demonstrating its significant advantages over existing auxiliary modality generation modules.
[0045] Table 4. Test results (%) of VI-ReID models containing different auxiliary modality generation modules The effectiveness of different blending methods for the upper and lower body regions was investigated experimentally, and the results are shown in Table 5. In Table 5, "√" indicates that an auxiliary image was generated by replacing the horizontal stripes in either the upper or lower body region. It can be observed that replacing only the horizontal stripes in the upper body region significantly improves all three metrics. Replacing only the horizontal stripes in the lower body region also shows a significant improvement compared to the domain baseline, but it is not as good as blending only the upper body horizontal stripes. This is because the upper body often has more discriminative features. When both regions are included in the blending strategy, the three metrics show even greater improvement. This indicates that the pedestrian proportion in each region affects the model's performance.
[0046] Table 5. Impact of horizontal stripe blending strategies on model performance for different body parts (%) Further experiments were conducted to explore the impact of the number of horizontal stripes in different human body regions on auxiliary modality generation, and the results are shown in Table 6. The experiments show that there is a clear sweet spot for the number of stripes: the model achieves optimal performance when two stripes are taken for each of the upper and lower body regions; excessively large or small stripes weaken matching accuracy. When the stripes are too narrow, it is difficult for a single region to completely represent the human body structure, making local features susceptible to perturbations by viewpoint and posture, significantly increasing the difficulty of representation. As the randomly spliced region becomes larger, the auxiliary modality tends to be biased towards either visible light or infrared light, thus increasing the modal spacing and hindering knowledge transfer.
[0047] Table 6. Effect of different horizontal stripe numbers on model performance (%) In addition, prior factors for human body parts were explored through experiments. The impact on model performance is shown in Table 7. It can be seen that when the upper body region... Take 0.3 and the lower body area The optimal ratio is 0.7. Compared to other ratios, 0.3 / 0.7 also exhibits the characteristics of the golden ratio mathematically, approaching the ideal distribution ratio of discriminant features; while 0.4 / 0.6 is too balanced, leading to a dilution of discriminant power; and 0.2 / 0.8 is significantly unbalanced.
[0048] Table 7. Effects of prior factors from different body parts on model performance (%) To further verify the AMAF model's ability to mitigate modal differences, the frequency distribution of intra-class and inter-class distances was visualized on the SYSU-MM01 dataset. For example... Figure 3 As shown, the vertical lines mark the means of each distribution, and... This represents the difference between two means. (Comparison) Figure 3 (a) and Figure 3 As shown in (b), the AMAF model successfully amplified the mean interval ( This indicates that the AMAF model is robust to identity recognition under the influence of modal changes.
[0049] The above description is merely a preferred embodiment of the present invention. The present invention is not limited to the above-described embodiments. Any embodiment that achieves the technical effects of the present invention by the same or equivalent means should fall within the protection scope of the present invention. Within the protection scope of the present invention, various modifications and variations can be made to the technical solutions and / or implementation methods.
Claims
1. A method for training a pedestrian re-identification model based on adaptive modal perception, characterized in that, Includes the following steps: Generate auxiliary modal images based on visible light and infrared images; Features are extracted from the visible light image, infrared image, and auxiliary modal image respectively to obtain the features of the visible light image, the infrared image, and the auxiliary modal image; The pedestrian re-identification model is trained using features from visible light images, infrared images, and auxiliary modal images. Update the parameters of the pedestrian re-identification model based on the target loss; The step of generating auxiliary modal images based on visible light and infrared images includes: Visible light images and infrared images are divided into... A semantically consistent horizontal stripe, among which It is a positive integer greater than 1; Based on the probability of replacement Horizontal stripes from either a visible light image or an infrared image are selected to generate the auxiliary modal image.
2. The method according to claim 1, characterized in that, The replacement probability Calculated by the following formula: in, Indicates the first The probability of replacing a horizontal stripe. Indicates the first Prior factors for the human body parts corresponding to each horizontal stripe. ; Indicates the first Normalized significance weights of horizontal stripes; This represents a non-zero constant.
3. The method according to claim 2, characterized in that, The strategy for generating auxiliary modal images is shown in the following equation: in, Represents auxiliary modal images. This indicates an image stitching operation; Indicates the infrared image number 1 The image region corresponding to each stripe Represents the visible light image. The image area corresponding to each stripe; Let represent a random variable that is uniformly distributed on the interval [0,1].
4. The method according to claim 2, characterized in that, The first The normalized significance weights of the horizontal stripes are calculated by the following formula: in, This represents the smallest significance weight among all horizontal stripes. This represents the one with the largest significance weight among all horizontal stripes; Indicates the first The significance weight of each stripe is calculated by the following formula: Indicates the height of the horizontal stripes. Indicates the width of the feature map; Indicates the first One horizontal stripe; Spatial attention map exist Location values, spatial attention map Calculated by the following formula: This represents the global max pooling result across the channel dimensions of the input feature tensor. This represents the result of global average pooling across the channel dimensions of the input feature tensor. This indicates a splicing operation. This represents a two-dimensional convolution with a kernel size of 7. This indicates that the Sigmoid function is activated.
5. The method according to claim 2, characterized in that, Determine the prior factors of human body parts based on the body parts themselves.
6. The method according to claim 1, characterized in that, The target loss includes a symmetric cross-modal consistency loss, wherein the symmetric cross-modal consistency loss includes a symmetric cross-modal consistency loss for generating auxiliary modal images based on visible light images. Symmetric cross-modal consistency loss based on generating auxiliary modal images from infrared images As shown in the following formula: in, This indicates the number of labeled categories in the training batch. This indicates the number of image instances for each identifier; Represents the L2 norm; Indicates the first The first identity under the th Visible light image features, Indicates the first The first identity under the th Infrared image features, Indicates the first The first identity under the th Auxiliary image features; This indicates the control parameter.
7. The method according to claim 1, characterized in that, The target loss also includes a cascaded aggregation loss, as shown in the following formula: in, Indicates identity index is Visible light image feature center, ; Indicates identity index is The infrared image feature center, ; This represents the semantic center of the visible light image obtained through a two-layer fully connected network. ; This represents the semantic center of the infrared image extracted through a two-layer fully connected network. ; This represents a two-layer fully connected network.
8. A pedestrian re-identification method, characterized in that, Includes the following steps: The visible light image and infrared image to be identified are input into the pedestrian re-identification model, and the pedestrian re-identification model outputs the identification result. The pedestrian re-identification model is obtained by the pedestrian re-identification model training method based on adaptive modal perception as described in any one of claims 1-7.
9. A pedestrian re-identification device, characterized in that, include: At least one processor; and memory that is communicatively connected to at least one processor; The memory stores instructions that, when executed by at least one processor, implement the pedestrian re-identification method of claim 8.
10. A computer-readable storage medium, characterized in that, The system stores instructions that, when executed by a processor, implement the pedestrian re-identification method of claim 8.