Cross-modal pedestrian re-identification method based on channel perception collaborative feature enhancement

Through the channel-aware collaborative feature enhancement method, we deeply explore the subtle features within the channel and learn the semantic relationship between channels, which solves the problem of insufficient feature extraction in cross-modal pedestrian re-identification and achieves more efficient all-weather pedestrian recognition.

CN120708270APending Publication Date: 2025-09-26CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510377438.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing cross-modal person re-identification methods have shortcomings in extracting robust identity recognition features, especially in complex scenarios, where recognition performance is limited. They also ignore the mining of intermediate-granularity channel-level features, resulting in poor recognition results.

Method used

A cross-modal person re-identification method based on channel-aware collaborative feature enhancement is designed. The channel-level feature mining module is used to deeply mine subtle features within the channel. The inter-channel semantic relationship learning module is combined to improve the depth and breadth of feature representation. The adaptive feature fusion module is used to optimize the final feature representation.

Benefits of technology

It significantly improves the accuracy and robustness of pedestrian re-identification, enables effective all-weather pedestrian monitoring in complex scenarios, and improves the precision of feature extraction and recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HDA0005333479620000011
    Figure HDA0005333479620000011
  • Figure HDA0005333479620000012
    Figure HDA0005333479620000012
  • Figure HDA0005333479620000021
    Figure HDA0005333479620000021
Patent Text Reader

Abstract

The invention discloses a cross-modal pedestrian re-identification method based on channel perception collaborative feature enhancement. Fine-grained modal sharing feature extraction is performed from a channel level. Specifically, a channel-level feature mining module is designed to mine features directly from a channel level to capture fine features within each channel. An inter-channel semantic relationship learning module is designed, focuses on learning the relationship of semantic information between different channels, and works in parallel with a channel-level feature mining module, so that the depth and breadth of feature representation are improved. And then designing an adaptive feature fusion module for intelligently processing features from the channel-level feature mining module and the inter-channel semantic relationship learning module, and optimizing final feature representation through an adaptive fusion mechanism. According to the method, comprehensive evaluation is carried out on three widely-used cross-modal data sets, and the effectiveness of the proposed method is proved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a method for visible light and infrared cross-modal pedestrian re-identification, specifically a cross-modal pedestrian re-identification method based on channel perception collaborative feature enhancement. Background Art

[0002] Person re-identification (ReID) aims to match images or video sequences of the same person captured by multiple cameras to determine the specific pedestrian's movement trajectory. This is of great significance to the development of fields such as intelligent security and criminal investigation. However, person re-identification faces numerous challenges, such as changes in viewing angle, pedestrian posture, clothing, occlusion, and motion blur. Although relevant research has made significant progress in recent years, most previous studies have focused on visible light images captured in well-lit environments.

[0003] Given that visible light images cannot capture sufficient information in low-light environments, and that many crimes occur at night, new cameras are now equipped with both visible light and infrared capture modes. Against this backdrop, the field of cross-modal person re-identification (VI-ReID) has emerged. This field focuses on matching visible light and infrared data of the same person captured by different cameras to achieve effective all-weather monitoring.

[0004] Compared with traditional single-modal pedestrian re-identification, cross-modal pedestrian re-identification has three major challenges: first, there are significant differences between visible light and infrared modalities, and it is difficult to align cross-modal features related to identity recognition; second, infrared images are seriously affected by lighting and cannot utilize color information, which is very important for identity recognition; third, the shooting data has a large time span and pedestrians' clothing changes frequently, which increases the difficulty of extracting robust features.

[0005] In the field of cross-modal person re-identification, modal differences are often mitigated by effectively extracting rich cross-modal shared features. Although existing methods have made many efforts, there are still obvious deficiencies in feature extraction and modal fusion, which limits their recognition performance in complex and changing real-world scenarios. Current representation learning-based methods usually adopt fine-grained pixel-level methods and coarse-grained global feature extraction methods to obtain diverse cross-modal shared features. However, these methods generally ignore the mining of intermediate-granularity channel-level features, resulting in poor performance of the model in complex scenarios. Some methods focus too much on fine-grained pixel-level features. Although they can capture rich detailed information, they have defects in overall semantic understanding; while other methods focus too much on coarse-grained global features and tend to ignore key local details, which seriously affects the recognition effect.

[0006] To address these issues, we propose a cross-modal person re-identification method based on channel-aware collaborative feature enhancement. This method attributes semantic differences between modalities to differences at the channel level and attempts to indirectly improve identification performance by optimizing semantic consistency between channels. Furthermore, we deeply explore complex channel feature relationships and extract intermediate-granularity channel-level features to effectively address modality differences. Summary of the Invention

[0007] The proposed cross-modal pedestrian re-identification method based on channel-aware collaborative feature enhancement effectively extracts cross-modal shared features through innovative network structure design, overcomes the shortcomings of existing technologies in extracting robust identity recognition features, improves the accuracy of pedestrian re-identification in complex scenarios, and achieves more efficient and stable all-weather pedestrian monitoring.

[0008] The present invention provides a cross-modal person re-identification method with channel-aware collaborative feature enhancement, the method comprising:

[0009] S1: Obtain pedestrian images containing different modalities (visible light and infrared) and their corresponding identity labels from relevant datasets, ensuring that the number of visible light images and infrared images of pedestrians of each identity is the same; preprocess the relevant datasets to expand the training dataset and increase the diversity of the dataset;

[0010] S2: Design a channel-level feature mining module to perform feature mining on the input image at the channel level, capturing subtle features within each channel and providing fine-grained information for subsequent feature extraction;

[0011] S3: Builds an inter-channel semantic relationship learning module, which works in parallel with the channel-level feature mining module and focuses on learning the relationship between semantic information between different channels. This improves the depth and breadth of feature representation and enhances the understanding of pedestrian characteristics.

[0012] S4: Build an adaptive feature fusion module to intelligently fuse the features output by the channel-level feature mining module and the inter-channel semantic relationship learning module, and optimize the final feature representation through the adaptive fusion mechanism to make it more representative and robust.

[0013] S5: A channel-aware collaborative feature enhancement network model constructed based on the visible light image and infrared image training;

[0014] S6: Using the constructed network model, feature extraction is performed on the visible light and infrared images in the test set to obtain pedestrian features of the corresponding modality. The pedestrian features of one modality are then used as the query set and the pedestrian features of the other modality as the image gallery. The similarity between each pedestrian feature in the query set and each pedestrian feature in the image gallery is calculated and ranked, and the cross-modal pedestrian re-identification results are obtained based on the ranking results.

[0015] In S1, we select datasets commonly used in the field of cross-modal person re-ID, such as SYSU-MM01, RegDB, and LLCM, from which we obtain images of pedestrians of several identities and their corresponding identity labels. These datasets contain abundant visible light and infrared pedestrian image samples, which can provide sufficient data for model training and testing.

[0016] In S2, the channel-level feature mining module uses convolutional attention operations such as triple attention mechanism and efficient channel attention mechanism to deeply explore the subtle features in each channel;

[0017] In S3, the inter-channel semantic relationship learning module uses the SENet channel attention mechanism and collaborative attention to effectively capture the complex semantic associations between different channels when learning inter-channel semantic relationships, further improving the feature representation capability and enhancing the model's adaptability to complex scenarios.

[0018] In S4, the adaptive fusion mechanism of the adaptive feature fusion module can be implemented through learnable weight parameters, and the fusion ratio is dynamically adjusted according to the importance of different features to achieve the optimal feature fusion effect;

[0019] In S5, the channel-aware collaborative feature enhancement network model includes a channel-level feature mining module, an inter-channel semantic relationship learning module, and an adaptive feature fusion module;

[0020] In S6, when calculating similarity, common metrics such as cosine similarity and Euclidean distance are used to measure the similarity between cross-modal pedestrian features. At the same time, the parameters of the similarity calculation method can be adjusted according to the actual application scenario to optimize the accuracy and reliability of the re-identification results;

[0021] A dual-branch ResNet-50 is used as the backbone network to extract features from visible light images and infrared images.

[0022] The ResNet-50 network has five stages. The first stage is usually the initial convolutional layer and pooling layer. The first two stages of the ResNet-50 network (stages 2 and 3) are used as weight sharing to extract shallow common features of different modal data.

[0023] The channel-aware collaborative feature enhancement network model includes a channel-level feature mining module, an inter-channel semantic relationship learning module, and an adaptive feature fusion module. The fourth and fifth stages of the ResNet-50 network are set as the modality-specific information mining part. When processing visible light images and infrared images, the network parameter weights of these two stages are not shared to better capture the unique features of different modal data.

[0024] The features extracted by the backbone network are simultaneously input into two parallel modules;

[0025] In S2, the channel-level feature mining module captures subtle features within each channel and deeply mines the local information of a single channel to ensure that the uniqueness of each channel is fully reflected. The specific processing flow is as follows:

[0026] Assume that the input feature map X∈R B×C×H×W , where B represents the batch size, C represents the number of input channels, H and W represent the feature map height and width respectively. The input feature map (B, C, H, W) is reshaped to (H, W, B, C), and self-attention is calculated separately for the RGB and IR feature maps using multi-head attention.

[0027] MultiHead(Q,K,V)=Concat(head1,...,head h )W o (1)

[0028] in, is the learned weight matrix, W O is the weight matrix of the output linear transformation, and h represents the number of heads. The self-attention output is added to the input, and then layer normalization is performed to stabilize the training process.

[0029] X′=LayerNorm(X+MultiHead(X)) (2)

[0030] The features are further processed through a feedforward network (FFN), consisting of two linear transformation layers and a ReLU activation function in between. The format X suitable for the attention mechanism is first processed through Triplet Attention, which optimizes feature representation from different perspectives using three attention mechanisms: spatial, channel, and token.

[0031] The Spatial Attention (SA) module reduces the number of channels to 1 through a 7×7 convolutional layer and uses a Sigmoid activation function to generate a spatial weight map. This weight map emphasizes important spatial locations in the image, which helps improve the model's accuracy in localizing the target object.

[0032] X sa =σ(Conv2d(X)) (1)

[0033] Here, σ is the Sigmoid activation function, which is used to limit the output to the range [0, 1].

[0034] The channel attention (CA) module first uses adaptive average pooling to compress each feature map into a scalar value. It then uses two convolutional layers (the first reduces the number of channels to 1 / 16 of the original number, and the second restores it back to the original number) and a ReLU activation function to model the dependencies between channels. Finally, a sigmoid activation function is used to generate channel weights. By adjusting the importance weight of each channel, channels that carry more useful information receive more attention, thereby reducing dependence on irrelevant channels.

[0035] X ca =σ(Conv2d(ReLU(Conv2d(AdaptiveAvgPool2d(X))))) (2)

[0036] The Token Attention (TA) module uses two consecutive 3×3 convolutional layers plus a ReLU activation function to learn more complex local patterns and uses a Sigmoid activation function to generate token-level attention weights. This helps further refine local features, capture subtle changes at each position, and enhance understanding of local details.

[0037] X ta =σ(Conv2d(ReLU(Conv2d(X)))) (3)

[0038] Then, the results of the above three attention modules are added together and multiplied with the original input feature map X to obtain the feature map X′ that integrates the three attention information.

[0039] X′=X·(X sa +X ca +X ta ) (4)

[0040] Next, the fused feature map X′ is input into the ECA (Efficient Channel Attention) module. By learning the importance weights of each channel, the model can focus on feature channels that are more helpful for the current task while maintaining low computational complexity. It first compresses the feature map into a scalar value for each channel through global average pooling to form a one-dimensional vector. This vector is then processed through a one-dimensional convolutional layer with a configurable kernel size to capture inter-channel dependencies at different scales. Finally, the output passes through the Sigmoid activation function, then expands back to its original shape and multiplied by the previously obtained X′ to obtain the final output feature map X″.

[0041] X″=expand(σ(Conv1d(squeeze(AdaptiveAvgPool2d(X′))))) (5)

[0042] Squeeze and expand represent removing redundant dimensions and restoring the original dimensions, respectively. In this way, the CLFM module not only highlights important channel features but also effectively suppresses noise and irrelevant information. This design helps the model better focus on identity-related information, thereby improving the consistency and robustness of cross-modal feature representation in the VI-ReID task.

[0043] In S3, the inter-channel semantic relationship learning module learns the semantic relationships between different channels, focusing on understanding the interactions between channels and the global structure, thereby improving the depth and breadth of feature representation. The specific processing flow is as follows;

[0044] The feature map X∈R B×C×H×W The SE block is input to capture global semantic relationships. The SE block captures global information and adaptively reweights the importance of each channel to highlight important features. Through the global average pooling operation, the input feature map X∈R B×C×H×W Compressed to Z∈R B×C×1×1 , then two fully connected layers (or convolutional layers) are used to generate the weights of each channel and converted to values ​​in the range (0, 1) through the Sigmoid activation function.

[0045] F se (Z)=σ(FC(δ(FC1(Z)))) (6)

[0046] Where σ represents the Sigmoid function and δ represents the ReLU activation function. Apply the weights generated by the SE block to the original input feature map X:

[0047] X′=X⊙F se (Z) (7)

[0048] Then, Coordinate Attention (CoordAtt) is used to capture the spatial structure of the feature map and further refine the feature representation. Adaptive average pooling is used to extract features along the height and width dimensions respectively, obtaining two one-dimensional feature vectors X h and X w , X h and X w Concatenate and transform through a shared convolutional layer, and then activate with the h-swish activation function:

[0049] Y=h-swish(Conv([X h ;X w ])) (8)

[0050] The output is again split into two parts, corresponding to the attention map A in the height and width directions respectively h and A w , and apply the Sigmoid activation function:

[0051] A h =σ(Conv h (Y)) A w =σ(Conv w (Y)) (9)

[0052] The attention weights in the height and width directions are obtained and multiplied with the original features, thereby enhancing the representation ability of the features while retaining the spatial position information.

[0053] X isrl =X′⊙A h ⊙A w (10)

[0054] In S4, the features processed by the parallel modules are then input into the adaptive feature fusion module. The adaptive adjustment mechanism ensures that the features from the parallel modules can be effectively integrated to form a more comprehensive, robust, and efficient feature representation. The final output of the network model is as follows:

[0055] X final =α·X clfm +(1-α)X isrl (11)

[0056] where X final Represents the output feature map after fusion. α is a learnable parameter with a value range of [0,1], which is used to control the two input feature maps X clfm and X isrl fusion ratio.

[0057] In S6, the loss function is used to optimize the network and classification results, and the similarity of the features output by the network model is measured. When calculating the similarity, common measurement methods such as cosine similarity and Euclidean distance are used to measure the similarity between cross-modal pedestrian features. At the same time, the parameters of the similarity calculation method can be adjusted according to the actual application scenario to optimize the accuracy and reliability of the re-identification results.

[0058] The beneficial effects of the present invention are:

[0059] 1. Through the channel-level feature mining module, the present invention deeply mines features from the channel level and effectively captures the subtle features within each channel, making the model more accurate in grasping the details of pedestrian images. In subsequent feature extraction, based on this fine-grained information, it can generate feature representations that are highly sensitive to the local features of pedestrians, significantly improving the precision of feature extraction.

[0060] 2. This invention designs an inter-channel semantic relationship learning module, focusing on learning the association between semantic information between different channels. In this way, it broadens and deepens the dimension of feature representation, enabling the model to understand pedestrian features from multiple angles and extract more distinctive and comprehensive pedestrian features, greatly enhancing the depth and breadth of the model's understanding of pedestrian features.

[0061] 3. The present invention constructs a parallel working mechanism. The channel-level feature mining module goes deep into the channel to extract fine-grained local features; the inter-channel semantic relationship learning module grasps the semantic connection between different channels from a macro perspective. The outputs of the two complement each other, avoiding the one-sidedness of feature extraction by a single module, constructing a more comprehensive pedestrian feature extraction system, and greatly enhancing the accuracy of the model in pedestrian identity recognition.

[0062] 4. This invention uses an adaptive fusion mechanism to intelligently integrate the features output by the channel-level feature mining module and the inter-channel semantic relationship learning module, allowing the model to better adapt to the characteristics of different modal data such as visible light and infrared, and optimize the final feature representation, so that it still has strong robustness in complex and changing environments, significantly improving the accuracy of pedestrian feature extraction.

[0063] 5. This paper conducts a comprehensive evaluation on three widely used cross-modal datasets, and the results fully demonstrate the effectiveness of this method in the cross-modal pedestrian re-identification task. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 is a flow chart of the present invention;

[0065] Figure 2 Schematic diagram of the network structure of the present invention.

[0066] Figure 3 Schematic diagram of the visualization results of the present invention. DETAILED DESCRIPTION

[0067] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0068] like Figure 1 As shown, an embodiment of the present invention provides a cross-modal pedestrian re-identification method based on channel-aware collaborative feature enhancement, which includes the following steps:

[0069] (1) Obtain pedestrian images of different modalities (visible light and infrared) and their corresponding identity labels from relevant datasets, ensuring that the number of visible light images and infrared images in the pedestrian images of each identity is the same; preprocess the relevant datasets (random horizontal flipping and random erasing techniques) to expand the training dataset and increase the diversity of the dataset;

[0070] (2) A dual-branch ResNet-50 was constructed as the backbone network. The first two stages (stage 1 and stage 2) of the ResNet-50 network were set as the modality-specific information mining parts. When processing visible light images and infrared images, the network parameter weights of these two stages were not shared, so as to capture the unique features of different modal data respectively. The third, fourth and fifth stages of the ResNet-50 network were used as the weight-sharing parts to extract the common features of different modal data.

[0071] (3) The features extracted by the backbone network are simultaneously input into two parallel modules. The channel-level feature mining module captures the subtle features within each channel and deeply explores the local information of a single channel to ensure that the uniqueness of each channel is fully reflected. The inter-channel semantic relationship learning module learns the semantic relationship between different channels, focusing on understanding the interaction and global structure between channels, thereby improving the depth and breadth of feature representation.

[0072] (4) The features processed by the parallel module are then input into the adaptive feature fusion module. Through the adaptive adjustment mechanism, it ensures that the features from the parallel modules can be effectively integrated to form a more comprehensive, robust and efficient feature representation.

[0073] (5) Use the loss function to optimize the network and classification results, measure the similarity of the features output by the network model, and output the matching results.

[0074] The experimental analysis is as follows:

[0075] 1. Implementation details:

[0076] This paper was implemented using the PyTorch framework, with training and evaluation performed on a Tesla P100 GPU. We used a ResNet-50 pre-trained on the ImageNet dataset as the backbone network. Input images were resized to 3×384×144.

[0077] 2. Sampling strategy:

[0078] In each mini-batch, we randomly select 4 VIS images and 4 IR images from 6 identities for training. During the test phase, only modality-shared features are used to evaluate performance. In addition, we employ random horizontal flipping and random erasing techniques during the training phase.

[0079] 3. Training parameter settings:

[0080] The model was trained for a total of 150 epochs. The learning rate used a warm-up strategy with the initial learning rate set to 0.1. It was increased to 1×10-1 using the warm-up strategy after 10 epochs, decayed to 1×10-2 after 20 epochs, and decayed to 1×10-3 and 1×10-4 after 60 and 120 epochs respectively until the 150th epoch.

Claims

1. A cross-modal person re-identification method based on channel-aware collaborative feature enhancement, characterized by: The specific steps of the method are as follows: Step 1: Use the dual-branch ResNet-50 as the backbone network to extract features from visible light images and infrared images; In step 2, the features extracted by the backbone network are simultaneously input into two parallel modules. The channel-level feature mining module captures subtle features within each channel, deeply exploring the local information of a single channel to ensure that the uniqueness of each channel is fully reflected. The inter-channel semantic relationship learning module learns the semantic relationships between different channels, focusing on understanding the interactions between channels and the global structure, thereby improving the depth and breadth of feature representation. In step 3, the features processed by the parallel modules are then input into the adaptive feature fusion module. Through the adaptive adjustment mechanism, the features from the parallel modules can be effectively integrated to form a more comprehensive, robust and efficient feature representation. Step 4: Use the loss function to optimize the network and classification results, measure the similarity of the features output by the network model, and output the matching results.

2. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The specific process of step 1 is as follows: Step 101: The ResNet-50 network has five stages. The first stage is usually the initial convolutional layer and the pooling layer. The first two stages (stages 2 and 3) of the ResNet-50 network are used as weight sharing to extract shallow common features of data from different modalities. In step 102, the fourth and fifth stages of the ResNet-50 network are set as the modality-specific information mining part. When processing visible light images and infrared images, the network parameter weights of these two stages are not shared to better capture the unique features of different modal data.

3. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The specific process of step 2 is as follows: Between the fourth and fifth stages, the channel-level feature mining module and the inter-channel semantic relationship learning module are designed in parallel. The features extracted by the backbone network are input into the two parallel modules at the same time. Among them, the channel-level feature mining module performs feature mining on the input image at the channel level when performing channel-level feature mining, capturing subtle features within each channel and providing fine-grained information for subsequent feature extraction, while the inter-channel semantic relationship learning module focuses on learning the relationship between semantic information between different channels, thereby improving the depth and breadth of feature representation and enhancing the understanding of pedestrian characteristics.

4. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The processing flow of the channel-level feature mining module in step 2 is as follows: Assume that the input feature map X∈R B×C×H×W , where B represents the batch size, C represents the number of input channels, H and W represent the feature map height and width respectively. The input feature map (B, C, H, W) is reshaped into (H, W, B, C), and multi-head attention is used to perform self-attention calculations on the RGB and IR feature maps respectively. MultiHead(Q,K,V)=Concat(head1,...,head h )W o (1) Among them, head i =Attention(QW i Q ,KW i K ,VW i V ). W i Q ,W i K ,W i V is the learned weight matrix, W O is the weight matrix of the output linear transformation, and h represents the number of heads. The self-attention output is added to the input, and then layer normalization is performed to stabilize the training process; X′=LayerNorm(X+MultiHead(X)) (2) The features are further processed through a feedforward network (FFN), consisting of two linear transformation layers and a ReLU activation function in the middle. The format X suitable for the attention mechanism is first processed by TripletAttention. This module uses three attention mechanisms: spatial, channel, and token, to optimize the feature representation from different perspectives. The Spatial Attention (SA) module reduces the number of channels to 1 through a 7×7 convolutional layer and generates a spatial weight map using a Sigmoid activation function. This weight map emphasizes important spatial locations in the image, which helps improve the model's accuracy in localizing the target object. X sa =σ(Conv2d(X)) (3) Where σ is the Sigmoid activation function, which is used to limit the output to the range of [0,1]; The channel attention (CA) module first uses adaptive average pooling to compress each feature map into a scalar value. It then uses two convolutional layers (the first reduces the number of channels to 1 / 16 of the original number, and the second restores it back to the original number) and a ReLU activation function to model the dependencies between channels. Finally, a sigmoid activation function is used to generate channel weights. By adjusting the importance weight of each channel, it ensures that channels that carry more useful information receive more attention, thereby reducing dependence on irrelevant channels. X ca =σ(Conv2d(ReLU(Conv2d(AdaptiveAvgPool2d(X))))) (4) The Token Attention (TA) module uses two consecutive 3×3 convolutional layers with ReLU activation functions to learn more complex local patterns and uses a Sigmoid activation function to generate token-level attention weights. This helps further refine local features, capture subtle changes at each position, and enhance understanding of local details. X ta =σ(Conv2d(ReLU(Conv2d(X)))) (5) Then, the results of the above three attention modules are added together and multiplied with the original input feature map X to obtain the feature map X′ that integrates the three attention information; X′=X·(X sa +X ca +X ta ) (6) Next, the fused feature map X′ is input into the Efficient Channel Attention (ECA) module. By learning the importance weights of each channel, the model can focus on feature channels that are more helpful for the current task while maintaining low computational complexity. It first compresses the feature map into a scalar value for each channel through global average pooling, forming a one-dimensional vector. This vector is then processed by a one-dimensional convolutional layer with a configurable kernel size to capture inter-channel dependencies at different scales. Finally, the output passes through a Sigmoid activation function, then expands back to its original shape and multiplies it with the previously obtained X′ to obtain the final output feature map X″. X″=expand(σ(Conv1d(squeeze(AdaptiveAvgPool2d(X′))))) (7) Among them, squeeze and expand represent removing redundant dimensions and restoring to the original dimensions, respectively. In this way, the CLFM module can not only highlight important channel features, but also effectively suppress noise and irrelevant information. This design helps the model better focus on identity recognition-related information, thereby improving the consistency and robustness of cross-modal feature expression in the VI-ReID task.

5. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The processing flow of the inter-channel semantic relationship learning module in step 2 is as follows: The feature map X∈R B×C×H×W The purpose of the SE block is to capture global semantic relationships. The SE block captures global information and adaptively reweights the importance of each channel to highlight important features. Through the global average pooling operation, the input feature map X∈R B×C×H×W Compressed to Z∈R B×C×1×1 , then use two fully connected layers (or convolutional layers) to generate the weights of each channel and convert them into values ​​in the range of (0,1) through the Sigmoid activation function; F se (Z)=σ(FC(δ(FC1(Z)))) (8) Where σ represents the Sigmoid function and δ represents the ReLU activation function. The weights generated by the SE block are applied to the original input feature map X: X′=X⊙F se (Z) (9) Then, collaborative attention is used to capture the spatial structure of the feature map and further refine the feature representation. Adaptive average pooling is used to extract features along the height and width dimensions respectively, obtaining two one-dimensional feature vectors X h and X w , X h and X w Concatenate and transform through a shared convolutional layer, and then activate with the h-swish activation function: Y=h-swish(Conv([X h ;X w ])) (10) The output is again split into two parts, corresponding to the attention map A in the height and width directions respectively h and A w , and apply the Sigmoid activation function: TO h =σ(Conv h (ALREADY w =σ(Conv w (Y)) (11) Obtain the attention weights in the height and width directions, and then multiply them with the original features, thereby enhancing the representation ability of the features while retaining the spatial position information; X isrl =X′⊙A h ⊙A w (12) In S4, the features processed by the parallel modules are then input into the adaptive feature fusion module. The adaptive adjustment mechanism ensures that the features from the parallel modules can be effectively integrated to form a more comprehensive, robust, and efficient feature representation. The final output of the network model is as follows: X final =α·X clfm +(1-a)X isrl (13) where X final Represents the output feature map after fusion, α is a learnable parameter with a value range of [0,1], which is used to control the two input feature maps X clfm and X isrl fusion ratio.

6. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The final output of the parallel module in step 2 is as follows: X p =X clfm +X isrl (14) 7. The cross-modal person re-identification method based on channel-aware collaborative feature enhancement according to claim 1 is characterized in that: The specific process in step 3 is as follows: Considering that simply adding the final output information of the parallel modules is not the optimal fusion method, the features processed by the parallel modules are input into the adaptive feature fusion module. Through the adaptive adjustment mechanism, it is ensured that the features from the parallel modules can be effectively integrated to form a more comprehensive, robust and efficient feature representation.