Cross-modal video monitoring pedestrian re-identification method based on text prototype guiding part alignment

Through the text prototype guides the alignment of the site, the body semantic consistency constraints are used to extract and align cross-modal features, and the problem of distribution differences between visible light images and infrared images is solved, and the accuracy and robustness of cross-modal pedestrian re-identification are improved.

CN120182889AInactive Publication Date: 2025-06-20ZHENJIANG DAJIANG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510264566.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-20
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively bridge the distribution differences between visible light images and infrared images, resulting in insufficient accuracy and robustness of cross-modal video surveillance pedestrian re-identification.

Method used

The method of guiding part alignment of text prototypes is adopted to extract part features through part-based text prototypes and align them using body semantic consistency constraints to improve the model's matching ability to cross-modal features.

Benefits of technology

It significantly improves the accuracy and robustness of cross-modal pedestrian re-identification, enhances the model's matching ability to different modes, and avoids additional noise and performance overhead caused by domain differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182889A_ABST
    Figure CN120182889A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal video monitoring pedestrian re-identification method based on text prototype guide part alignment. A text guide model is utilized to extract and align part features. The method comprises the following steps: firstly, respectively extracting corresponding feature maps from a visible light image and an infrared image by using a shallow-layer parallel deep-layer shared modal specific feature extractor; and inputting the text template which is provided with the learnable mark and is specific to the body part into a text encoder to obtain a text prototype corresponding to the body part, and taking the text prototype as text guidance for extracting body part characteristics. Thirdly, performing cross attention fusion on the text prototype and the corresponding feature map to obtain corresponding part features; and body semantic consistency constraint is applied to the part features, so that the semantic consistency of the local features is further improved. According to the method, the part features are aligned by using the text guide model, and the alignment of the semantic features with finer granularity enables the model to have stronger robustness for cross-modal and intra-modal differences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the fields of computer vision and deep learning, and more specifically, to the field of cross-modal pedestrian re-identification technology. Background Art

[0002] Pedestrian re-identification is a technology that uses computer vision technology to detect whether a specific pedestrian exists in an image or video sequence. Due to the growing demand for public safety and the increasing number of surveillance cameras, pedestrian re-identification technology has received more and more attention.

[0003] In practical applications such as intelligent surveillance and security systems, it is necessary to identify pedestrian identities all day long. However, visible light cameras are difficult to shoot at night or in environments with poor light. Therefore, only infrared cameras can be used to shoot infrared images of pedestrians to make up for the deficiencies of visible light cameras, thereby expanding the application scope and effect of the surveillance system. However, due to different imaging mechanisms, there are significant distribution differences between visible light images and infrared images. Therefore, how to bridge this distribution difference has become the main problem to be solved by cross-modal video surveillance pedestrian re-identification methods.

[0004] The generation-based approach aims to introduce an intermediate modality and uses a generative adversarial network to obtain a modality between the two modalities, thereby bridging the modality gap. However, this approach is prone to introducing additional noise, and at the same time, the information in the original image cannot be fully utilized, and effective information will be lost.

[0005] The part-alignment-based approach can provide more fine-grained recognition information, while the lack of strictly paired cross-modal image pairs hinders the performance of local alignment in identifying and locating body parts. This approach usually locates through existing body parsing models. With additional body parsing models, due to domain differences, additional noise is inevitably introduced, and at the same time, additional performance overhead will be brought.

[0006] Therefore, how to improve the model's ability to extract and align part features has become an urgent problem to be solved in cross-modal pedestrian re-identification. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a cross-modal video surveillance pedestrian re-identification method guided by text prototypes for part alignment. Using part-based text prototypes as a guide, part features are extracted, and body semantic consistency constraints are used to align part features. Using these more fine-grained semantic feature alignments makes the model more robust and generalization-capable to cross-modal and intra-modal differences.

[0008] The purpose of the present invention is achieved through the following technical solutions:

[0009] A cross-modal video surveillance pedestrian re-identification method for aligning text prototype guiding parts of the present invention includes the following steps:

[0010] Step 1: Obtain a visible light-infrared image dataset of pedestrians, extract modality-specific features from visible light images and infrared images respectively through a modality-specific feature extractor, and then obtain corresponding feature maps through a shared feature extractor.

[0011] Step 2: Input a body part-specific text template with learnable markers into a text feature extractor to obtain corresponding body part text prototypes.

[0012] Step 3: Perform cross-attention fusion on the text body part text prototypes and the feature maps.

[0013] Obtain corresponding body part features and global part features.

[0014] Step 4: Utilize body semantic consistency constraints to improve the model's ability to extract and discriminate part features, thereby enhancing the generalization ability of visible light and infrared cross-modal pedestrian re-identification.

[0015] Further, in step 1 of the present invention, a preprocessing strategy including channel random erasing and horizontal flipping is used for visible light images and infrared images to enrich the diversity of training samples. At the same time, a random channel exchange strategy is introduced for visible light images to improve the robustness to different modalities. Then, images of the two modalities are input into a modality-specific feature extractor, and finally input into a shared feature extractor to obtain corresponding feature maps.

[0016] Further, in step 2 of the present invention, the text encoder is composed of the text encoder in CLIP. The text template form is "A photo of a person’s [P]", where P is a trainable parameter, and different Ps represent different parts of the body.

[0017] Further, in step 3 of the present invention, the text prototype is used as the query, and the feature map is used as the key and value for cross-attention fusion to obtain part features. The text prototypes based on parts are added and averaged and then concatenated with the original text prototype as the input of the query for cross-attention fusion, and the feature map output from the shared feature extractor is used as the input of the key and value. After cross-attention fusion calculation, text-guided part features are obtained.

[0018] Furthermore, in step 4 of the present invention, the body semantic consistency constraint includes an inter - part discrimination loss and an intra - part alignment loss. The inter - part discrimination loss aims to enable different text tokens based on body parts to learn the semantic information of different body parts, so as to more comprehensively learn the semantic information describing body parts, thereby avoiding over - reliance on a certain part for identity discrimination. It is achieved by increasing the distance between the features of different parts of the same person in the feature space. The purpose of the intra - part alignment loss is to make the features of the same part more discriminative between different identities and more consistent between the same identities. It is realized by, for the features of the same part, increasing the distance between the center of the part features and the samples of the features of other different parts, and reducing the distance between the center of the part and the samples of the features of other parts belonging to the same part. Compared with the prior art, the beneficial effects of the present invention are:

[0019] The present invention uses a text prototype - guided model to extract part features and align them, so as to align features of different modalities at a finer - grained level, thereby making cross - modal feature matching more accurate. Among them, the text prototype is obtained from the training of a body - part - based text template with trainable parameters, so it can capture part features helpful for identity recognition. Compared with predefined part texts, the text template with learnable parameters is more flexible. The part features guided by the text prototype will not change due to modality, and the same part can be extracted for different modalities, thereby improving the model's matching ability for different modalities. Compared with the method of using existing body parsing models to divide the body, the present invention does not need to introduce additional computational complexity and can avoid additional noise caused by domain differences.

[0020] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures specifically pointed out in the written specification and the drawings.

[0021] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0022] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0023] Figure 1 It is the overall module structure diagram of a cross - modal video surveillance pedestrian re - identification method for text prototype - guided part alignment in an embodiment of the present invention;

[0024] Figure 2Flowchart of a cross-modal video surveillance pedestrian re-identification method for aligning text prototype guiding parts in an embodiment of the present invention;

[0025] Figure 3 Schematic diagram of part feature distribution in a cross-modal video surveillance pedestrian re-identification method for aligning text prototype guiding parts in an embodiment of the present invention; Detailed implementation manners

[0026] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0027] Figure 1 As shown, the overall model structure includes an image feature extractor S100, a text feature extractor S200, a cross-attention fusion module S300, and a body semantic consistency alignment module S400. Among them, the image feature extractor includes an infrared feature extractor S111, a visible light feature extractor S112, and a shared feature extractor S121, and the body semantic consistency alignment module includes an inter-part discrimination module S411 and an intra-part alignment module S41.

[0028] Figure 2 As shown is the overall flowchart. A cross-modal video surveillance pedestrian re-identification method for aligning text prototype guiding parts in this embodiment is specifically implemented as follows:

[0029] Step 1: First, obtain images of two modalities and perform data augmentation first. Preprocessing strategies including random channel erasing and horizontal flipping are used for visible light images and infrared images to enrich the diversity of training samples. At the same time, a random channel exchange strategy is introduced for visible light images to improve the robustness to different modalities. Then, the images of the two modalities after data augmentation are input into the modality-specific feature extractors, and their outputs are used as the inputs of the shared feature extractor to obtain the corresponding feature maps.

[0030] Step 2: Input the body part-specific text template with learnable markers into the text feature extractor to obtain the corresponding body part text prototypes, including the global part text prototype t containing all text prototype features p . The text encoder is composed of the text encoder in CLIP. The form of the text template is "A photo ofa person’s[P]", where P is a trainable parameter, and different Ps represent different parts of the body.

[0031] Step 3: Perform cross-attention fusion on the text body part text prototypes and the feature maps to obtain the corresponding body part features and global part features. Use the text prototypes as queries and the feature maps as keys and values to perform cross-attention fusion to obtain part features.

[0032]

[0033] First, perform a flattening operation on the feature map, and then concatenate the flattened feature map f flat and the text prototype t p with its corresponding vector center to obtain f cat and Then perform cross-attention fusion.

[0034] K c = f cat W k , V c = f cat W v ,

[0035]

[0036] where C represents the number of channels of the feature map, W q , W k , W v and W o are learnable weight matrices, D t is the dimension of the text features, D h and D o are the dimensions of the hidden layer and the output layer. where is the local feature of the k-th part, is the overall local feature representation containing all part feature information, denoted as the classification part feature.

[0037] Step 4: Use the body semantic consistency constraint to improve the model's ability to extract and discriminate part features, thereby enhancing the generalization ability of visible light and infrared cross-modal pedestrian re-identification. The body semantic consistency constraint includes the inter-part discrimination loss L ipd and the intra-part alignment loss L ipc . Among them, the goal of L ipd is to enable different text markers based on parts to learn the semantic information of different body parts, so as to more comprehensively learn the semantic information describing body parts, thereby avoiding over-reliance on a certain part for identity discrimination. Its expression is as follows:

[0038]

[0039] where N p is the number of samples of each identity in a batch, N id is the number of identities in this batch. B = N p×N id , D 1 (i, j, y a ) represents the L2 distance between the part centers of the i-th part and the j-th part with identity y within a batch, where y a represents the identity label corresponding to the a-th identity within the batch, α2 is the margin hyperparameter, and K is the number of parts. a And and represent the features of the i-th part and the j-th part with identity label y a . The purpose of L ipc is to make the features of the same part more distinguishable between different identities and more consistent between the same identities. It is expressed as:

[0040]

[0041]

[0042] where D 2 (y i , y j , k) represents the L2 distance between the k-th part within a batch in the sample with identity label y i and the part feature center with identity label y i .

[0043] In this implementation, a batch size B = 64, and the two distance hyperparameters are α1 = 0.6 and α2 = 0.3 respectively.

[0044] Figure 3 Shown is a schematic diagram of the distribution change of part features under the consistency constraint. In the initial situation, the distributions of various part features are chaotic and difficult to use for identity information discrimination. After training under the consistency constraint, the part features of different identities are distributed in different regions, and for the same part features of the same identity, different samples will gather together. It is used to perform feature alignment at the part level without considering modal differences. The present invention uses text to guide the extraction and alignment of part features, which can effectively reduce the modal differences between visible light and infrared images and significantly improve the accuracy and robustness of cross-modal pedestrian re-identification.

[0045] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and it should not be used to limit the protection scope of the present invention. All technical solutions obtained by adopting equivalent replacements or equivalent transformations fall within the protection scope of the present invention.

Claims

1. A cross-modal video surveillance pedestrian re-identification method with text prototype-guided part alignment. Its characteristics are: The following steps are involved: Step 1: Obtain a visible light-infrared image dataset of pedestrians, extract modality-specific features from the visible light image and infrared image respectively through the modality-specific feature extractor, and then obtain the corresponding feature map through the shared feature extractor. Step 2: Input the body part-specific text template with learnable tags into the text feature extractor to obtain the corresponding body part text prototype. Step 3: Cross-attention fusion of text body part prototypes and feature maps. Get the corresponding body part features and global part features. Step 4: Use body semantic consistency constraints to improve the model's ability to extract and discriminate part features, thereby improving the generalization ability of visible light and infrared cross-modal pedestrian re-identification.

2. The method for cross-modal video surveillance pedestrian re-identification based on text prototype-guided part alignment according to claim 1 is characterized in that: In step 1, in the dataset, preprocessing strategies including random channel erasing and horizontal flipping are used for visible light images and infrared images, and a random channel swapping strategy is introduced for visible light images to improve the robustness to different modalities. The images of the two modalities are then input into the modality-specific feature extractors and finally into the shared feature extractor to obtain the corresponding feature maps.

3. The method for cross-modal video surveillance pedestrian re-identification based on text prototype-guided part alignment according to claim 1 is characterized in that: In step 1, the modality-specific feature extractor includes an infrared feature extractor and a visible light feature extractor. The infrared feature extractor extracts modality-specific features from infrared images, and the visible light feature extractor extracts modality-specific features from visible light images. Both the visible light feature extractor and the infrared feature extractor are composed of the shallow convolution part of ResNet-50 in the multimodal model CLIP (Contrastive Language-Image Pretraining). The shared feature extractor extracts modality-shared features and is composed of the deep residual block part of ResNet-50 in CLIP. After different modality images are input, the feature map f is finally obtained from the output of the shared feature extractor. map .

4. The method for cross-modal video surveillance pedestrian re-identification based on text prototype-guided part alignment according to claim 1 is characterized in that: Step 2: The text encoder is constructed from the text encoder in CLIP. The text template is in the form of "A photo of a person's [P]", where P is a trainable parameter and different P represents different parts of the body.

5. The method for cross-modal video surveillance pedestrian re-identification based on text prototype-guided part alignment according to claim 1 is characterized in that: In step 3, the text prototype is used as the query, and the feature map is used as the key and value to perform cross-attention fusion to obtain the part feature. The part-based text prototypes are added and averaged, and then connected with the original text prototype as the input of the cross-attention fusion query. The feature map output from the shared feature extractor is used as the key and value input, and after cross-attention fusion calculation, the text-guided part feature is obtained.

6. The method for cross-modal video surveillance pedestrian re-identification based on text prototype-guided part alignment according to claim 1, characterized in that: In step 4, the body semantic consistency constraint includes inter-part discrimination loss and same-part alignment loss. The goal of inter-part discrimination loss is to enable different text labels based on parts to learn semantic information of different body parts, so as to learn semantic information describing body parts more completely, thereby avoiding over-reliance on a certain part for identity discrimination. The purpose of same-part alignment loss is to make the features of the same part more distinguishable between different identities, and at the same time, the features of the same part are more consistent between the same identity. It is implemented by increasing the distance between the center of the feature of the same part and other different part feature samples, and shortening the distance between the center of the part and other part feature samples belonging to the same part.