Visible light-infrared pedestrian re-identification method based on morphological center representation learning
Through a morphological center representation learning method, combined with a dual-stream network, a shape feature learning network and an infrared shape recovery method, the appearance and shape characteristics of pedestrians are extracted and enhanced, and the accuracy and reliability of visible-infrared pedestrian re-identification in the prior art are solved, achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510115845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing visible-infrared pedestrian re-identification technology has low accuracy and reliability when matching visible light and infrared images, making it difficult to meet the complex and changeable actual monitoring needs.
Using a method based on morphological center representation learning, appearance features are extracted through a dual-stream network, and combined with shape feature learning network and infrared shape recovery method, shape features are extracted and corrected, and finally, appearance features are enhanced through cascading two-stage attention networks.
It improves the accuracy and reliability of pedestrian re-identification, overcomes the problem of difficult infrared image shape extraction in the prior art, enhances the ability to distinguish shape features, and meets the needs in complex and changing scenarios.
Smart Images

Figure CN120047893A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pedestrian re-identification, and specifically to a visible light-infrared pedestrian re-identification method based on morphological center representation learning. Background Art
[0002] Although the current pedestrian re-identification technology has made certain progress and plays a crucial role in intelligent monitoring systems, it has also attracted much attention and made rapid progress in recent years. However, most existing pedestrian re-identification methods are limited to processing scenes visible during the day and rely too much on visible appearance features. When matching pedestrian images captured by visible light and infrared cameras, this excessive reliance on visible appearance features will lead to a sharp decline in performance and cannot meet the complex and changing actual monitoring needs. This is mainly because there are significant differences between visible light and infrared spectral images, making traditional visible light feature-based methods unable to effectively cope.
[0003] To solve this problem, the visible light-infrared pedestrian re-identification task has now been introduced, aiming to achieve the retrieval of pedestrians under different spectra (such as visible light and infrared spectra). However, due to the large amount of intra-modal and inter-modal variations between visible light and infrared spectral images, this brings great difficulties to feature extraction and matching.
[0004] To match pedestrians of different morphologies, although two types of methods, namely feature-level modality alignment and image-level modality alignment, have been developed to suppress or reduce modality differences. However, modality-shared cues, such as shape-centered cues, have not been fully explored and utilized, which greatly limits the discriminability of features and thus affects the accuracy and robustness of pedestrian re-identification.
[0005] Moreover, most of the currently widely studied pedestrian re-identification methods rely heavily on color appearance, which makes it difficult for them to play an effective role in special scenarios such as cloth-changing ReID that require changing the dependence on color appearance. Their application scenarios are greatly restricted and cannot meet the requirements in complex and changing scenarios.
[0006] Even though attempts have been made to introduce artificial semantic parsing to obtain refined local features. But most of them mainly focus on human body parts and ignore that there are other important cues besides human body parts, such as personal items, which can also be used as useful cues, resulting in the difficulty of fully exploiting semantic parsing. And in terms of the current technology, semantic parsing is not completely reliable, which is often ignored, leading to a great reduction in the accuracy and reliability of existing pedestrian re-identification and making it difficult to meet the diverse and complex application requirements. Summary of the Invention
[0007] The present invention aims to provide a visible light-infrared pedestrian re-identification method based on morphological center representation learning to solve the problem that the existing visible light-infrared pedestrian re-identification technology has low accuracy and reliability when matching visible light and infrared images, and is difficult to meet diverse and complex application requirements.
[0008] To achieve the above object, the present invention adopts the following technical solution: a visible light-infrared pedestrian re-identification method based on morphological center representation learning, comprising an appearance feature learning step and a shape feature learning step;
[0009] The appearance feature learning step is to use a two-stream network to remove non-local attention as a baseline; the two-stream network includes two parallel convolutional layers with non-shared parameters and four residual networks with shared parameters, denoted as E a ; Two parallel convolutional layers are used to process visible light and infrared images to extract shallow appearance features, and the shallow appearance features obtained are input into four residual networks; the four residual networks extract high-level semantic features f based on the shallow appearance features i m ;
[0010] The shape feature learning step is to use a shape feature learning network E s Learn the visible light and infrared human body shape features to obtain body shape features; and use infrared shape recovery method to correct the errors of the obtained body shape features;
[0011] It also includes an appearance feature enhancement step, which uses a cascaded two-stage attention network to extract appearance features related to body shape, and enhances the acquired shallow appearance features based on the appearance features.
[0012] The principles and advantages of this solution are:
[0013] In the actual application of pedestrian re-identification, in order to quickly find matching personnel information from a large number of image information based on the provided human body image, such as in actual applications such as finding lost persons, it is necessary to quickly and accurately find the target and effectively track its traces in the images obtained from a large number of road cameras. At present, most of them rely on extracting facial images to identify the target to ensure the accuracy and effectiveness of recognition, and do not use appearance too much for identification, such as clothing, head shape, etc., because it is easy to produce similar and repeated results through appearance features, resulting in low recognition. However, the area of the face is small, it is difficult to accurately capture effective facial information, and the facial state sometimes differs (such as makeup, etc.). Therefore, in the case of more obvious visible light, appearance features are also combined for recognition. Therefore, the combination of facial and appearance features can meet the needs of pedestrian re-identification. Instead of choosing a shape with lower recognizability, that is, an outer contour shape, to complete the recognition task.
[0014] However, it is found in practical applications that facial and appearance features more rely on scenes with obvious visible light during the day. As a result, in the case of weak visible light, it is difficult to meet the data volume and effective messages required for recognition. And this solution proposes to combine human shape for recognition, which effectively breaks through the limitations of data provision and improves the availability of data. Moreover, the shape mentioned in this solution is not only the outer contour of the human body itself. The outer contours of the human body itself have small differences and are difficult to provide distinctiveness. The shape features in this solution also include external structural features such as the items carried with the person. Under a certain base number, by analyzing the human shape, the inertial actions or behavioral habits of pedestrians can be captured, thereby improving the recognition degree of pedestrian features. Based on this, the formed shape features are referable and distinguishable. Therefore, this solution overcomes the technical prejudice of low referability and poor distinguishability of human shape features in the prior art, and uses the overall shape features as recognition features for training.
[0015] However, in the prior art, since visible light images are formed by reflected light, they can clearly present details such as the color and texture of pedestrians; while infrared images are formed based on the thermal radiation of objects themselves, highlighting temperature differences. This makes the shape performance of the same pedestrian vary greatly in the two modalities. For example, details such as the folds of a pedestrian's clothes are clearly visible under visible light and can assist in outlining the shape; but in infrared images, if the thermal radiation differences of the clothing materials are small, these details will be blurred, resulting in difficulties in extracting shape features. This makes the existing pedestrian re-identification technology more focused on the appearance recognition of visible light and will not use the modality of infrared images for shape recognition too much, which undoubtedly increases the recognition difficulty. And due to the limitation of the thermal radiation characteristics of infrared images, their data is also relatively single. For similar or close shape features, it is also difficult to effectively distinguish only by shape features. Therefore, it is more difficult to extract the shape features of pedestrians, and thus shape recognition will not be used too much in the technology of visible light-infrared pedestrian re-identification, which also leads to the limitations and difficulties of shape recognition in the prior art.
[0016] However, in the application of actual pedestrian re-identification technology, relying only on the appearance features of visible light for recognition, it can only be applied to scenes visible during the day and is difficult to meet the complex and diverse requirements of pedestrian re-identification. And due to the difficulty in extracting appearance in the infrared spectrum, the application effect of the existing visible light-infrared pedestrian re-identification is difficult to be further improved, and then the recognition effect cannot meet the application requirements.
[0017] This solution overcomes the problem of difficult shape extraction from infrared images. By constructing a feature learning framework focused on learning shape-centered features, it can effectively extract and learn shape features without using an auxiliary network during the inference process. An infrared shape recovery method is introduced to further correct the inaccurate infrared human shapes generated by the human parsing network, enhancing the discriminative ability of shape features, thereby improving the shape features obtained from infrared images. This method is based on the acquired appearance features and uses the appearance features as a guide to correct the extracted shape features. At the same time, to ensure the accuracy of the correction, an appearance feature enhancement step is further introduced to enhance the appearance features using the internal relationship between shape and appearance features, ensuring the accuracy and effectiveness of the appearance features, and thus improving the accuracy of the shape features to overcome the problem that shape features are difficult to be recognized and effectively extracted. Moreover, this solution successfully integrates explicit shape features into the visible-infrared pedestrian re-identification task, breaking through the technical bottleneck and providing more advanced performance for the industry development. Brief Description of the Drawings
[0018] Figure 1 It is a schematic diagram of the process framework of the present invention.
[0019] Figure 2 It is a schematic diagram of the process of the infrared shape recovery method in the present invention.
[0020] Figure 3 It is a schematic diagram of the process of the appearance feature enhancement step in the present invention.
[0021] Figure 4 It is a comparison chart of pedestrian re-identification retrieval results under different ablation settings of the present invention. Detailed Description of the Invention
[0022] The following is a further detailed description through specific embodiments:
[0023] Embodiment 1
[0024] The visible-infrared pedestrian re-identification method based on morphological center representation learning in this embodiment consists of two main branches: a shape stream and an appearance stream. In the shape stream, the feature extraction network and the infrared shape recovery work together to extract reliable shape features and enhance the discriminative ability of shape features. And through the appearance feature enhancement step, the features related to the shape are emphasized to enhance the discriminability of the appearance features. Focusing on learning shape-centered features without using an auxiliary network during the inference process, it successfully integrates explicit shape features into the visible-infrared pedestrian re-identification task, improving its accuracy and reliability. In this embodiment, as shown in the appendix Figure 1 It includes an appearance feature learning step and a shape feature learning step.
[0025] The appearance feature learning step uses a two-stream network to remove non-local attention as the baseline. In this embodiment, the two-stream network is denoted as E a , which includes two parallel convolutional layers with non-shared parameters and four residual networks with shared parameters.
[0026] In this embodiment, the two parallel convolutional layers are used to independently process visible light and infrared images to extract shallow appearance features, and the obtained shallow appearance features are input into the four residual networks of the feature extraction network E a . The four residual networks extract high-level semantic features f i m based on the shallow appearance features. In this embodiment, the high-level semantic feature f i m can be expressed as:
[0027]
[0028] In the formula, GeM represents the Generalized Mean Pooling layer; represents the input pedestrian image; i represents the i-th pedestrian image in a batch of pedestrian images; m represents the modality to which the pedestrian image belongs; vis represents the visible light modality; ir represents the near-infrared modality.
[0029] At the same time, in order to ensure the identity discriminability of the high-level semantic feature f i m , the high-level semantic feature f i m is constrained by using cross-entropy loss and weighted regularized triplet loss.
[0030] In this embodiment, the cross-entropy loss is The weighted regularized triplet loss is Among them, the weighted regularized triplet loss is used to constrain and optimize the appearance features by narrowing the distance between appearance features of the same identity and pushing away the distance between appearance features of different identities.
[0031] The shape feature learning step uses the shape feature learning network E s to learn the human body shape features of visible light and infrared to obtain the body shape feature. The way to obtain the body shape feature is: input the body shape x s,i generated in the human body parsing model into the shape feature learning network Es, and extract the body shape feature map F s,i = E s (x s,i ). Further, by performing Generalized Mean Pooling on Fs,i, the body shape feature is obtained. Among them, taking visible light as an example, the way to obtain the body shape feature is: input the body shape generated in the human body parsing model into the shape feature learning network Es Extract the visible light body shape feature map And perform global average pooling to obtain the shape feature And through the cross-entropy loss And the weighted regularization triplet loss Constrain and optimize the shape feature The acquisition method of the near-infrared body shape feature is the same as that of the visible light, and will not be elaborated here.
[0032] At the same time, an infrared shape recovery method is used to correct the error of the acquired body shape feature.
[0033] In this embodiment, the input is the infrared shape feature map Where j represents the j-th image feature, and the infrared feature map F obtained from the input image i ir,j . In this embodiment, as shown in the appendix Figure 2 The infrared shape recovery method is to use the sum of the shape feature map and the infrared feature map F i ir,j as the Query, use F i ir,j as the Key and Value, obtain the missing shape clues from the original infrared feature map through the cross-attention method and perform correction and recovery, and record the recovered infrared feature map as
[0034] It also includes an appearance feature enhancement step. In this embodiment, as shown in the appendix Figure 3 The cascaded two-stage attention network is used to enhance the appearance features related to the body shape and enhance the appearance features.
[0035] The cascaded two-stage attention network includes the following two stages
[0036] The first stage: mine the appearance features directly related to the shape; take the features output by the shape sub-network as the Query value of the first-layer self-attention mechanism, and take the features a output by E i as the Key value and Value value of the first-layer self-attention mechanism; process to obtain the result value and use it as the Query value of the second-layer self-attention mechanism.
[0037] Among them is the shape sub-network, and its structure is the same as that of the 4th residual block of the feature extraction network, and is used to encode the shape features from the appearance features a output by E
[0038] The second stage: further explore the appearance features indirectly related to body shape based on the features obtained in the first stage; take the feature F i as the Key value and Value value of the second-layer self-attention mechanism to obtain the appearance features indirectly related to body shape.
[0039] In this embodiment, it also includes generalizing and average pooling the features output in the appearance feature learning step, and the output result is denoted as and used in the shape feature learning step. In the shape feature learning step, perform shape feature propagation processing on the obtained result In this embodiment, calculate the similarity with the weight matrix ; calculate the distance between the real shape feature and to obtain the loss and to obtain the loss and calculate L kd loss; by executing the loss and the loss L kd jointly distill the body shape feature to
[0040] In this embodiment, the proposed deep learning framework focuses on learning shape-centered features without the need to use an auxiliary network during the inference process. Importantly, this solution is the first to successfully integrate explicit shape features into the visible light-infrared pedestrian re-identification task. At the same time, in this embodiment, an infrared shape recovery technology is introduced to correct the inaccurate infrared human body shape generated by the human parsing network and enhance the discriminative ability of the shape features. An appearance feature enhancement step is adopted to enhance the appearance features by utilizing the internal relationship between the shape and appearance features.
[0041] At the same time, combined with the publicly available datasets in this field (SYSU-MM01, and HITSZ-VCM), extensive experiments are carried out, and the results show that the method proposed by the present invention achieves state-of-the-art performance.
[0042] As shown in Table 1 below, it is the performance evaluation and comparison results of this solution and the current state-of-the-art pedestrian re-identification (Re-ID) method on the SYSU-MM01 dataset, including indicators such as the mean average precision value (mAP) and the cumulative match characteristic curve (CMC). Among them, "R1", "R10", and "R20" respectively represent the first, tenth, and twentieth rankings.
[0043] Table 1
[0044]
[0045] As can be seen from Table 1 above, in all search perspectives of this solution, the index R1 is 76.1%, R20 is 99.4%, and mAP is 72.6%; in the indoor search perspective, R1 is 82.4%, R20 is 99.8%, and mAP is 85.4%, indicating that it has good performance in the pedestrian re-identification task.
[0046] As shown in Table 2 below, it is the performance evaluation results of this solution and the current state-of-the-art pedestrian re-identification (Re-ID) methods on the HITSZ-VCM dataset. The comparison methods cover video-based and image-based methods.
[0047] Table 2
[0048]
[0049] From Table 2 above, it can be intuitively compared that in the video and image modes of the HITSZ-VCM dataset for this solution, in the visible light perspective, R1 is 71.2% and mAP is 52.9%; in the infrared perspective, R1 is 73.3% and mAP is 53.0%. Each index is relatively leading, and it has good performance in both visible light and infrared perspectives.
[0050] Combined with the appendix Figure 4 As shown, it shows the pedestrian re-identification retrieval results under different ablation settings on the SYSU-MM01 dataset. Among them, a, b, c, and d represent "Baseline" (baseline), "Baseline+SFP", "Baseline+SFP+ISR", and "Baseline+SFP+ISR+AFE" respectively; the green border represents the same identity, and the red border represents different identities. The figure shows the top 6 retrieval results. The left and right sides are the results under two different queries respectively. There are four sub-parts a, b, c, and d below each part, and each sub-part contains 6 pedestrian images.
[0051] From the appendix Figure 4From the results, it can be seen that as modules such as SFP, ISR, and AFE are gradually added to the baseline model, the number of images with green borders (i.e., correctly matching pedestrians with the same identity) increases. For example, under the left Query, the number of correct matches among the 6 results of a (Baseline) is relatively small, while the number of correct matches of d (Baseline+SFP+ISR+AFE) is more. This indicates that adding these modules significantly improves the hit rate of pedestrian re-identification. At the same time, by distinguishing between green and red borders, the accuracy of the model retrieval under different ablation settings can be visually seen. The more green borders, the better the model performs in retrieving pedestrians with the same identity. The red borders indicate that the model misjudges pedestrians with different identities as the target identity, and the fewer the number of red borders, the lower the misjudgment rate of the model. It can be seen that this solution effectively improves the hit rate of the model by adding modules such as SFP, ISR, and AFE, providing a visual reference basis for the optimization and improvement of the pedestrian re-identification model.
[0052] The above are only embodiments of the present invention. Specific technical solutions and / or common knowledge such as characteristics well known in the art are not described in detail herein. It should be pointed out that for those skilled in the art, without departing from the technical solution of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application shall be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.
Claims
1. A visible light-infrared pedestrian re-identification method based on morphological center representation learning, characterized by: It includes appearance feature learning step and shape feature learning step; The appearance feature learning step is to use a two-stream network to remove non-local attention as a baseline; the two-stream network includes two parallel convolutional layers with non-shared parameters and four residual networks with shared parameters, denoted as E a ; Two parallel convolutional layers are used to process visible light and infrared images to extract shallow appearance features, and the acquired shallow appearance features are input into four residual networks; Four residual networks extract high-level semantic features f based on shallow appearance features. i m ; The shape feature learning step is to use a shape feature learning network E s Learn the human body shape features in visible light and infrared to obtain body shape features; And the infrared shape recovery method is used to correct the error of the acquired body shape features; It also includes an appearance feature enhancement step, which uses a cascaded two-stage attention network to strengthen the appearance features related to body shape and enhance the appearance features.
2. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 1 is characterized in that: The high-level semantic feature f i m It is expressed as: Where GeM represents the generalized average pooling layer; represents the input pedestrian image; i represents the i-th pedestrian image; m represents the mode to which the pedestrian image belongs; vis represents the visible light mode; ir stands for near infrared modality.
3. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 1 is characterized by: The physical characteristics Including shape feature maps acquired through visible light And infrared characteristic images obtained through infrared The infrared shape recovery method is to convert the shape feature map And infrared signature and The addition result is used as Query, As the key and value, the missing shape clues are obtained from the original infrared feature map by cross attention, and the features of infrared shape recovery are obtained. Correction recovery is performed, denoted as 4. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 3 is characterized in that: The body shape feature is obtained by: s,i Input into the shape feature learning network Es to extract the body shape feature map F s,i =E s (x s,i ), and further perform generalized average pooling on Fs,i to obtain the body shape feature 5. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 4 is characterized in that: The cascaded two-stage attention network includes the following two stages: The first stage: mining appearance features directly related to shape; Output characteristics As the query value of the first layer of self-attention mechanism, take E a Output feature F i As the Key value and Value value of the first layer of self-attention mechanism; process to obtain the result value And use it as the query value of the second layer of self-attention mechanism; The second stage: Based on the features obtained in the first stage, further explore the appearance features indirectly related to body shape; i As the key and value of the second-layer self-attention mechanism, the appearance features indirectly related to the body shape are obtained.
6. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 2 is characterized by: It also includes high-level semantic features f i m Cross entropy loss and weighted regularized triplet loss are used for constraints.
7. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 5 is characterized by: The E s~ is the shape subnetwork, used to a The shape features are encoded in the output appearance features.
8. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 5 is characterized by: It also includes the shape subnetwork Output characteristics Perform generalized average pooling processing, and the output result is recorded as And used in the shape feature learning step.
9. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 8, characterized in that: It also includes the results obtained in the shape feature learning step Perform shape feature propagation processing; With the weight matrix Calculate the similarity and calculate the loss And by calculating L kd loss; loss through execution and loss L kd Body shape characteristics Distillation 10. The visible light-infrared pedestrian re-identification method based on morphological center representation learning according to claim 6, characterized in that: The cross entropy loss is The weighted regularized triplet loss is Used to constrain and optimize appearance features.