Pedestrian re-identification method and system in subway hub large passenger flow crowd monitoring

By constructing a feature diversity module, an edge information difference bridging module, and a shallow feature injection module, the problem of low cross-camera recognition accuracy for pedestrian re-identification in high-passenger-flow scenarios in subway hubs was solved, achieving improved accuracy for cross-modal pedestrian re-identification and enhancing subway operation and emergency management capabilities.

CN121904677APending Publication Date: 2026-04-21CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-02
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In subway hubs with high passenger flow, existing pedestrian re-identification methods have low recognition accuracy under complex lighting, large occlusion, and multi-view conditions, making it difficult to achieve continuous trajectory tracking and behavior monitoring across cameras, and the matching effect between infrared and visible light image modalities is poor.

Method used

A feature diversity module is constructed for image expansion and stitching, an edge information difference bridging module captures local edge intensity deviations and eliminates orientation specificity, an edge information-driven shallow feature injection module enhances deep semantic representation, and an overall network loss function is constructed to improve feature robustness.

Benefits of technology

It improves the accuracy of cross-modal pedestrian re-identification, thereby enhancing subway operation safety, passenger flow organization efficiency, and emergency management capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121904677A_ABST
    Figure CN121904677A_ABST
Patent Text Reader

Abstract

The invention discloses a pedestrian re-identification method and system in subway hub large passenger flow crowd monitoring, and the system comprises an edge computing device and a multi-mode camera. A feature diversity module, an edge information difference bridging module, an edge information driven shallow feature injection module and a multi-quality perception triple loss module which are connected in sequence are deployed on the edge computing equipment; the method comprises the following steps: constructing a feature diversity module to carry out feature expansion and splicing of spatial dimension and channel dimension on a visible light image and an infrared image respectively; constructing an edge information difference bridging module; the shallow feature injection module based on edge information driving injects the information subjected to edge information difference processing into deep semantic representation to realize collaborative enhancement of cross-level geometric information and semantic features; and constructing a network overall loss function. According to the method and the system, the consistency of different modal edge information can be improved, the characterization force and the robustness of features are enhanced, and the cross-modal pedestrian re-identification precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for pedestrian re-identification in monitoring large passenger flows in subway hubs, belonging to the field of pedestrian re-identification technology. Background Technology

[0002] To effectively monitor passenger flow within subway hubs and supervise personnel movement and operational behavior, numerous video cameras are deployed in key areas such as platforms, passageways, and turnstiles to achieve real-time video transmission and the identification and alarm of abnormal behavior. However, subway hub environments are typically characterized by complex spaces, uneven lighting, and dense crowds. Coupled with significant fluctuations in passenger flow at different times, severe occlusion of individuals, and high similarity in physical features, images suffer from occlusion, motion blur, and perspective differences, severely impacting the accuracy of pedestrian detection and recognition. This poses numerous challenges to traditional pedestrian re-identification methods in high-passenger-flow subway scenarios. Furthermore, existing recognition research is largely based on a single camera's independent viewpoint, making it difficult to achieve continuous tracking and trajectory reconstruction across cameras. Therefore, developing pedestrian re-identification methods and systems suitable for complex lighting, significant occlusion, and multi-viewpoint conditions, based on the existing subway video surveillance network, to achieve continuous tracking of personnel trajectories and behavioral monitoring across cameras, is a crucial technological prerequisite for improving subway operational safety, passenger flow organization efficiency, and emergency management capabilities.

[0003] Pedestrian re-identification (Re-ID) is an important technology in image retrieval, aiming to match pedestrian images across non-overlapping camera views based on feature similarity. This technology plays a crucial role in monitoring large passenger flows in subway hubs, enabling continuous tracking of pedestrian trajectories across cameras, thus supporting passenger flow analysis, behavior monitoring, and emergency command. While existing Re-ID methods have achieved good results under normal lighting conditions, in the complex indoor environment of subways, factors such as crowd obstruction and sudden changes in light can easily affect the quality of visible light images, leading to a decline in recognition performance. In emergencies such as fires, the presence of dense smoke further hinders effective monitoring by visible light cameras. The intervention of infrared cameras can penetrate smoke and capture thermal radiation information, providing an important supplement for cross-modal pedestrian matching. However, due to the significant spectral differences between visible light and infrared images, achieving effective matching between the two modalities remains a technical challenge. Current methods to address this challenge mainly include feature engineering methods and generative methods. However, both methods have the following drawbacks: While two-stream network architectures with shared parameters are widely used to extract modality-invariant features, their inherent drawback lies in the fact that the rigid constraints of parameter sharing may inhibit the effective expression of modality-specific discriminative information. These methods, by forcing visible and infrared branches to share backbone network parameters, can improve feature alignment efficiency, but they struggle to adapt to the inherent spectral and texture differences between different modalities. Especially during feature decoupling, mechanisms such as orthogonal constraints or adversarial learning can easily lead to excessive suppression of modality-specific features within the shared parameter framework, resulting in the loss of identity-related details. Furthermore, in terms of feature alignment, shared parameter networks have limited ability to model shallow local features, making it difficult to accurately capture complex nonlinear mapping relationships across modalities. This results in incomplete feature distribution alignment, affecting the model's generalization ability in complex scenarios. While generative methods mitigate modal differences through image transformation, their shared generative networks often struggle to accommodate the unique properties of different modalities, leading to structural distortions or blurred details in the generated images. These methods frequently employ a unified generator for bidirectional modal transformation; however, due to the fundamental differences in physical imaging mechanisms between visible light and infrared images, the shared network parameters cannot adequately adapt to the texture characteristics and semantic distributions of different modalities, easily introducing artifacts or noise. Furthermore, generative methods rely on complex cross-modal transformation networks, which have a large number of parameters, high computational costs, and whose generation quality is significantly affected by the distribution balance of training data. In dynamic and ever-changing subway monitoring scenarios, the shared generative network's insufficient adaptability to complex conditions such as lighting changes and occlusion limits its feasibility for deployment in real-time systems. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for pedestrian re-identification in the monitoring of large passenger flow in subway hubs. This method and system can realize the interaction between dense shallow edge information and deep semantics, improve the consistency of edge information of different modalities, enhance the representation power and robustness of features, improve the accuracy of cross-modal pedestrian re-identification, and improve the subway operation safety, passenger flow organization efficiency and emergency management level.

[0005] To achieve the above objectives, the present invention provides a method for pedestrian re-identification in large passenger flow monitoring in subway hubs, comprising the following steps: S1. By constructing a feature diversity module, the visible light image and infrared image are respectively extended and stitched in the spatial dimension and channel dimension. S2. Construct an edge information difference bridging module, which captures local edge intensity deviations and eliminates direction-specific three-dimensional edge response intensity to quantify edge differences, and uses energy coefficients to bridge information. S3. The shallow feature injection module driven by edge information injects the information after edge information difference processing into the deep semantic representation to achieve synergistic enhancement of cross-level geometric information and semantic features. S4. Construct the overall network loss function, where the quality-aware triplet loss supervises and constrains the quality of edge-aware features by perceiving the quality of sample quality, pushing away the feature center distribution of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.

[0006] Furthermore, the specific process of S1 is as follows: S1.1, Visible light image Expanded into spatial-level images With channel-level images ; Single-channel infrared image Expanded into spatial-level images With channel-level images ; S1.2 Feature extractor via shared parameter network G Feature mapping is performed on the expanded multi-dimensional data, and channel concatenation is then executed as follows: ; in, This represents the joint feature tensor output by layer 0 of the network. These are the feature components of the visible light spatial dimension, the visible light channel dimension, the infrared spatial dimension, and the infrared channel dimension, respectively.

[0007] Furthermore, the specific process of S2 is as follows: The module is used to systematically mitigate numerical scale bias in the edge information of shallow networks. Specifically, for a given feature map... Three sets of directional convolutions are introduced, each containing a directional mask and weights for the grouped convolutions. ,in m Indicates direction, g Indicates the number of groups; the orientation mask includes vertical, horizontal, and diagonal lines, each defined as follows: ; in, To represent two diagonal masks, denoted uniformly as... ; Automatic broadcasting will connect each mask to its corresponding Multiply, A mask for vertical, horizontal, and two diagonal lines. The local response along each direction is calculated as follows: ; in, This indicates that element-wise multiplication is broadcast along the spatial dimension; the mask is used to ensure that the output highlights edge responses and effectively characterizes the relative changes of local pixels in their directional neighborhood. The edge difference term is defined as: ; ; in, These represent the adaptive weight coefficients of the masks in the horizontal, vertical, and diagonal directions, respectively, where k represents the feature dimension. These represent the grouped convolution weights in the horizontal, vertical, and diagonal directions, respectively. This represents the adaptive parameters, i.e., the spatial summation of the directional convolution kernels. Further explanation: Assume the sample number is n The output channel element is o ,but Represented as: ; in, This indicates the number of input channels in each group. r For local channel indexing, g ( o () is a grouping function used to determine the current output channel. o The group index to which it belongs; It is shared in spatial location to map edge responses from different modalities to a unified energy metric space; The dot product yields the weighted energy at each location along the channel direction, reflecting the offset caused by the inherent local intensity deviation between different modes; edge information differences contain identity-related information, and an energy coefficient is introduced to ensure that useful information is retained while suppressing bias. ; in, It is a convolutional layer whose weights and biases are determined by... Derivation; Assume that the desired directional response of a given output channel can be expressed as the sum of edge information and bias term: ; Since the deviation is an estimate, Cannot be precisely separated ,Right now: ; in, Representing edge information The energy coefficient, Indicates the deviation term The energy coefficient, when Approximately 1 and When smaller, Strongly suppressed, representing stable structural information Then it can be preserved. The edge residuals are considered calibrated; the edge information difference bridging module is applied to the shallow layers of the network through the residual terms: ; in As an adjustable parameter, Introduced as a correction term with a small weight, it preserves the main structural information while making Overlay This improves the overall quality of the features.

[0008] Furthermore, the specific process of S3 is as follows: S3.1 Map features from different levels to a unified embedding space, and then apply multi-scale deep convolution to transform local representations into prior information with rich receptive fields: ; Where z is a set of different convolution kernel sizes (3×3, 5×5 and 7×7), and Z represents the effect of the convolution kernel.

[0009] S3.2. Utilize features rich in edge information to guide semantic learning and construct an edge-aware association matrix with edge geometric similarity. : ; in, For grouped convolutional layers, The temperature factor is used; to reduce semantic interference, a dual-branch Top-K sparsity strategy is introduced, along with... Geometric guidance is implemented in the feature direction, starting only from deep features. In-process polymerization and Most relevant Context token : ; ; in, To control Sparsity is a scalar parameter. The set of indices obtained from the Top-K operations. These represent the indices of the "multi-head" dimension in the attention mechanism, and the attention correlation matrix, respectively. Row index and attention association matrix column index, I This is an indicator function that returns 1 when the condition is true and 0 otherwise; sparse branches are fused using learnable weights. ; in, These are learnable scalar weights; By leveraging features rich in edge information as prior knowledge, the contribution strength of sparse semantic context is dynamically balanced. Through stacking EDI modules, shallow information with geometric priors is gradually injected into deep semantic representations, achieving synergistic enhancement of cross-level geometric information and semantic features.

[0010] Furthermore, the overall network loss function in S4 includes identity recognition loss. and quality-perceived triplet loss Traditional triplet loss methods assume that all samples have equal importance, ignoring differences in sample quality and difficulty. This can lead to noisy or low-quality samples dominating the optimization process, causing convergence instability and decreased discriminative power. Quality-aware triplet loss, on the other hand, is used to perceive sample quality, and its formal definition is: ; Where N is the batch size. Let represent the feature vectors of the anchor sample, positive sample, and negative sample of the i-th triplet, respectively; Represents Euclidean distance. margin This is a preset distance threshold. This is an adaptive weighting coefficient. This loss allows harder samples with smaller inter-class differences to receive higher weights, guiding the model to focus on more challenging sample relationships. To further integrate quality perception into the optimization process, a dynamic modulation factor is introduced. Adjust the weight sensitivity based on the structural reliability of the triplet: ; in It is the cardinality of the quality perception factor. It is a scaling constant. This design quantifies the quality of feature relationships between anchor samples, positive samples, and negative samples. Through this design, the model dynamically adjusts the weighting function for triples with reliable discriminative feature relationships, making the weighting function more sensitive to difficult but high-quality samples; conversely, it reduces the impact of unstable or noisy triples. QATL thus achieves a self-balancing optimization between hard sample mining and quality-driven learning, ensuring that the feature optimization process is guided by structurally consistent and semantically reliable samples.

[0011] The overall network loss function is obtained as follows: ; in, It is a cross-entropy loss function: ; Where i is the category number. Let N be the label of the i-th category, and N be the number of images input to the model in the current iteration.

[0012] The present invention also provides a pedestrian re-identification system for monitoring large passenger flow in subway hubs, including an edge computing device and a multimodal camera. The edge computing device is equipped with a feature diversity module, an edge information difference bridging module, an edge information-driven shallow feature injection module, and a multi-quality perception triplet loss module connected in sequence. The multimodal camera is used to capture streaming images and transmit the captured images to the edge computing device; The edge computing device is used to receive streaming images and process the images through the deployed feature diversity module, edge information difference bridging module, edge information-driven shallow feature injection module, and multi-quality perception triplet loss module. The feature diversity module is used to enhance the features of the input visible light and infrared images, copy and expand the single-channel infrared data into three channels, construct a tensor format consistent with the RGB input, extract spatial information through standard convolution kernels in the spatial dimension, and obtain channel information through channel normalization in the channel dimension. The edge information difference bridging module obtains the local edge strength deviation of the feature through the directional mask capture feature diversity module, then eliminates the directional specificity while retaining the edge response strength to quantify the edge difference, and performs information bridging through the energy coefficient; The edge information-driven shallow feature injection module enhances the density of local edge information purified by the edge information difference bridging module through a parallel sparsification mechanism, and selectively injects deep semantic representations to promote the interaction between shallow and deep edge perception features. The quality-perceived triplet loss supervises and constrains the quality of edge-perceived features by perceiving the quality of samples, pushing away the feature centers of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.

[0013] This invention constructs a feature diversity module to extend and stitch features in the spatial and channel dimensions of visible light and infrared images, respectively. It also constructs an edge information difference bridging module, which captures local edge strength deviations through directional masks and eliminates directional specificity in three dimensions while preserving edge response intensity to quantify edge differences. Information bridging is achieved through energy coefficients. Based on an edge information-driven shallow feature injection module, the edge information difference-processed information is injected into deep semantic representations, achieving synergistic enhancement of cross-level geometric and semantic features. A network overall loss function is constructed, and the quality of edge-perceived features is supervised and constrained by the quality of perceptual samples, pushing away the feature center distribution of different modalities in the same subspace while simultaneously bringing closer feature centers of different dimensions within the same modality. This invention achieves interaction between dense shallow edge information and deep semantics, improves the consistency of edge information across different modalities, enhances feature representation power and robustness, improves cross-modal pedestrian re-identification accuracy, and ultimately improves subway operation safety, passenger flow organization efficiency, and emergency management level. Attached Figure Description

[0014] Figure 1 This is a flowchart of the identification method of the present invention; Figure 2 This is a schematic diagram of the workflow of the feature diversity module of the present invention; Figure 3 This is a schematic diagram of the workflow of the edge information difference bridging module of the present invention; Figure 4 This is a schematic diagram illustrating the workflow of the overall network loss function of this invention; Figure 5 This is a schematic diagram of the identification system of the present invention. Detailed Implementation

[0015] The invention will now be further described with reference to the accompanying drawings.

[0016] like Figure 1 As shown, a method for pedestrian re-identification in monitoring large passenger flows at subway hubs includes the following steps: S1. By constructing a feature diversity module, the visible light image and infrared image are respectively extended and stitched in the spatial dimension and channel dimension. S2. Construct an edge information difference bridging module, which captures local edge intensity deviations and eliminates direction-specific three-dimensional edge response intensity to quantify edge differences, and uses energy coefficients to bridge information. S3. The shallow feature injection module driven by edge information injects the information after edge information difference processing into the deep semantic representation to achieve synergistic enhancement of cross-level geometric information and semantic features. S4. Construct the overall network loss function, where the quality-aware triplet loss supervises and constrains the quality of edge-aware features by perceiving the quality of sample quality, pushing away the feature center distribution of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.

[0017] like Figure 2 As shown, the specific process of S1 is as follows: S1.1, Visible light image Expanded into spatial-level images With channel-level images ; Single-channel infrared image Expanded into spatial-level images With channel-level images ; S1.2 Feature extractor via shared parameter network G Feature mapping is performed on the expanded multi-dimensional data, and channel concatenation is then executed as follows: ; in, This represents the joint feature tensor output by layer 0 of the network. These are the feature components of the visible light spatial dimension, the visible light channel dimension, the infrared spatial dimension, and the infrared channel dimension, respectively.

[0018] like Figure 3 As shown, the specific process of S2 is as follows: The module is used to systematically mitigate numerical scale bias in the edge information of shallow networks. Specifically, for a given feature map... Three sets of directional convolutions are introduced, each containing a directional mask and weights for the grouped convolutions. ,in m Indicates direction, g Indicates the number of groups; the orientation mask includes vertical, horizontal, and diagonal lines, each defined as follows: ; in, To represent two diagonal masks, denoted uniformly as... ; Automatic broadcasting will connect each mask to its corresponding Multiply, A mask for vertical, horizontal, and two diagonal lines. The local response along each direction is calculated as follows: ; in, This indicates that element-wise multiplication is broadcast along the spatial dimension; the mask is used to ensure that the output highlights edge responses and effectively characterizes the relative changes of local pixels in their directional neighborhood. The edge difference term is defined as: ; ; in, These represent the adaptive weight coefficients of the masks in the horizontal, vertical, and diagonal directions, respectively, where k represents the feature dimension. These represent the grouped convolution weights in the horizontal, vertical, and diagonal directions, respectively. This represents the adaptive parameters, i.e., the spatial summation of the directional convolution kernels. By degenerating into a 1×1 convolution weight, spatial directionality is eliminated, retaining only the energy amplitude. Further explanation: Assume the sample number is n The output channel element is o ,but Represented as: ; in, This indicates the number of input channels in each group. r For local channel indexing, g ( o () is a grouping function used to determine the current output channel. o The group index to which it belongs; It is shared in spatial location to map edge responses from different modalities to a unified energy metric space; The dot product yields the weighted energy at each location along the channel direction, reflecting the offset caused by the inherent local intensity deviation between different modes; edge information differences contain identity-related information, and an energy coefficient is introduced to ensure that useful information is retained while suppressing bias. ; in, It is a convolutional layer whose weights and biases are determined by... Derivation; Assume that the desired directional response of a given output channel can be expressed as the sum of edge information and bias term: ; Since the deviation is an estimate, Cannot be precisely separated ,Right now: ; in, Representing edge information The energy coefficient, Indicates the deviation term The energy coefficient, when Approximately 1 and When smaller, Strongly suppressed, representing stable structural information It was preserved. The edge residuals are considered calibrated; the edge information difference bridging module is applied to the shallow layers of the network via residual terms: ; in , As an adjustable parameter, Introduced as a correction term with a small weight, it preserves the main structural information while making Overlay This improves the overall quality of the features.

[0019] The specific process of S3 is as follows: S3.1 Map features from different levels to a unified embedding space, and then apply multi-scale deep convolution to transform local representations into prior information with rich receptive fields: ; Where z is a set of different convolution kernel sizes (3×3, 5×5 and 7×7), and Z represents the effect of the convolution kernel.

[0020] S3.2. Utilize features rich in edge information to guide semantic learning and construct an edge-aware association matrix with edge geometric similarity. : ; in For grouped convolutional layers, The temperature factor is used; to reduce semantic interference, a dual-branch Top-K sparsity strategy is introduced, along with... Geometric guidance is implemented in the feature direction, starting only from deep features. In-process polymerization and Most relevant Context token : ; ; in, To control Sparsity is a scalar parameter. The set of indices obtained from the Top-K operations. These represent the indices of the "multi-head" dimension in the attention mechanism, and the attention correlation matrix, respectively. Row index and attention association matrix column index, IThis is an indicator function that returns 1 when the condition is true and 0 otherwise; sparse branches are fused using learnable weights. ; in, These are learnable scalar weights; By leveraging features rich in edge information as prior knowledge, the contribution strength of sparse semantic context is dynamically balanced. Through stacking EDI modules, shallow information with geometric priors is gradually injected into deep semantic representations, achieving synergistic enhancement of cross-level geometric information and semantic features.

[0021] like Figure 4 As shown, the overall network loss function in S4 includes identity recognition loss. and quality-perceived triplet loss Traditional triplet loss methods assume that all samples have equal importance, ignoring differences in sample quality and difficulty. This can lead to noisy or low-quality samples dominating the optimization process, causing convergence instability and decreased discriminative power. Quality-aware triplet loss, on the other hand, is used to perceive sample quality, and its formal definition is: ; Where N is the batch size. Let represent the feature vectors of the anchor sample, positive sample, and negative sample of the i-th triplet, respectively; Represents Euclidean distance. margin For the preset distance threshold, This is an adaptive weighting coefficient. This loss allows harder samples with smaller inter-class differences to receive higher weights, guiding the model to focus on more challenging sample relationships. To further integrate quality perception into the optimization process, a dynamic modulation factor is introduced. Adjust the weight sensitivity based on the structural reliability of the triplet: ; in, It is the cardinality of the quality perception factor. It is a scaling constant. This design quantifies the quality of feature relationships between anchor samples, positive samples, and negative samples. Through this design, the model dynamically adjusts the weighting function for triples with reliable discriminative feature relationships, making the weighting function more sensitive to difficult but high-quality samples; conversely, it reduces the impact of unstable or noisy triples. QATL thus achieves a self-balancing optimization between hard sample mining and quality-driven learning, ensuring that the feature optimization process is guided by structurally consistent and semantically reliable samples.

[0022] The overall network loss function is obtained as follows: ; in, It is a cross-entropy loss function: ; Where i is the category number. Let N be the label of the i-th category, and N be the number of images input to the model in the current iteration.

[0023] like Figure 5 As shown, a pedestrian re-identification system for monitoring large passenger flows in a subway hub includes an edge computing device and a multimodal camera. The edge computing device is equipped with a feature diversity module, an edge information difference bridging module, an edge information-driven shallow feature injection module, and a multi-quality perception triplet loss module, which are connected in sequence. The multimodal camera is used to capture streaming images and transmit the captured images to the edge computing device; The edge computing device is used to receive streaming images and process the images through the deployed feature diversity module, edge information difference bridging module, edge information-driven shallow feature injection module, and multi-quality perception triplet loss module. The feature diversity module is used to enhance the features of the input visible light and infrared images, copy and expand the single-channel infrared data into three channels, construct a tensor format consistent with the RGB input, extract spatial information through standard convolution kernels in the spatial dimension, and obtain channel information through channel normalization in the channel dimension. The edge information difference bridging module obtains the local edge strength deviation of the feature through the directional mask capture feature diversity module, then eliminates the directional specificity while retaining the edge response strength to quantify the edge difference, and performs information bridging through the energy coefficient; The edge information-driven shallow feature injection module enhances the density of local edge information purified by the edge information difference bridging module through a parallel sparsification mechanism, and selectively injects deep semantic representations to promote the interaction between shallow and deep edge perception features. The quality-perceived triplet loss supervises and constrains the quality of edge-perceived features by perceiving the quality of samples, pushing away the feature centers of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.

[0024] Example: A cross-modal parameter-sharing neural network based on the ResNet 50 backbone. Structurally, this model is divided into five stages (Stage 0–Stage 4): Stage 0 employs independent convolutional paths, without parameter sharing, to extract low-level modality-specific features from visible light and infrared images separately, enhancing the perception of input information; Stages 1 to 4 are constructed as a unified parameter-sharing module, utilizing the residual structure and cross-layer connection mechanism of ResNet 50 to extract consistent high-level semantic information, thereby achieving cross-modal feature alignment. During training, images from both modalities are first processed through the independent branches of Stage 0 for preliminary feature adaptation, and then input into the subsequent shared network for deep representation learning. This design effectively achieves unified extraction of high-level semantic features while preserving modality-specific information, exhibiting good scalability and cross-modal matching performance.

[0025] The training setup was as follows: the backbone network used ResNet50 pre-trained on ImageNet; the training and evaluation datasets were SYSU-MM01, RegDB, and LLCM; the framework was PyTorch, and the computing hardware was an NVIDIA 4090D GPU. The input image size was set to 384×192, and zero-padding, multi-scale transformation, and random erasure were applied for data augmentation. Each training batch was sampled from 6 random identities, with each identity containing 4 RGB and 4 IR images. The SGD optimizer was used with a momentum parameter of 0.9 and a weight decay of 5×10⁻⁻⁶. 4 The learning rate was reduced to one-tenth of its original value at rounds 30, 90, and 120, for a total of 150 training rounds.

[0026] Table 1 Comparison methods for the SYSU-MM01 dataset As shown in Table 1, experimental results on the large-scale SYSU-MM01 dataset demonstrate that the proposed method exhibits superior performance in both Rank-k and mAP metrics. In All-Search mode, the proposed method achieves a Rank-1 accuracy of 78.56% and an mAP of 74.76%, ranking best among all state-of-the-art methods. In Indoor-Search mode, the Rank-1 accuracy is 82.97%, and the mAP is 85.57%, reaching the optimal or near-optimal level. This invention focuses on comparing several methods involving modality-specific features. For example, CSVI, through a cross-modal center weight generation module and a segmentation decoder, implicitly constructs an interaction mechanism between RGB and infrared images while extracting more modality-shared information; MRCN proposes a modality compensation module to distill modality-related features to supplement the feature representation of another modality; and FI introduces a dynamic aggregation module to project bimodal features onto each other. Compared with these methods, the proposed method achieves superior performance. These methods typically rely on explicit modal interactions, which can lead to performance degradation during the testing phase. In contrast, this paper identifies modality-specific components from shared features and improves feature quality by progressively removing these components, without introducing any explicit cross-modal interaction mechanisms.

[0027] Table 2 Comparison Methods for RegDB Datasets Table 3 Comparison methods for LLCM datasets As shown in Tables 2 and 3, the results on the RegDB and LLCM datasets further demonstrate that EAPINet has good generalization ability. Specifically, on RegDB, EAPINet achieves a Rank-1 accuracy of 91.4% and an mAP of 83.6% in V2I mode, and 88.8% and 80.1% respectively in I2V mode. On the LLCM dataset, the Rank-1 and mAP are 64.8% and 66.2% in V2I mode, and 55.2% and 61.0% in I2V mode. On RegDB, the method of this invention achieves the best or second-best results. Although the CSDN method injects textual knowledge into image features by leveraging CLIP's image-text alignment capability, thus achieving higher mAP and other metrics than the proposed method, the proposed method has an advantage in training complexity, as it does not require an additional text information alignment stage. It is worth noting that the performance improvement of the proposed method on the LLCM dataset is relatively limited. This is mainly due to the presence of low illumination and strong infrared noise in the dataset itself, which weakens the stability and discriminative power of edge features, resulting in relatively small cross-modal edge differences to be corrected and limited high-quality edge priors that can be injected. Nevertheless, the method of this invention still alleviates subtle modal biases to some extent and improves feature compactness.

Claims

1. A method for pedestrian re-identification in monitoring large passenger flows at subway hubs, characterized in that, Includes the following steps: S1. By constructing a feature diversity module, the visible light image and infrared image are respectively extended and stitched in the spatial dimension and channel dimension. S2. Construct an edge information difference bridging module, which captures local edge intensity deviations and eliminates direction-specific three-dimensional edge response intensity to quantify edge differences, and uses energy coefficients to bridge information. S3. The shallow feature injection module driven by edge information injects the information after edge information difference processing into the deep semantic representation to achieve synergistic enhancement of cross-level geometric information and semantic features. S4. Construct the overall network loss function, where the quality-aware triplet loss supervises and constrains the quality of edge-aware features by perceiving the quality of sample quality, pushing away the feature center distribution of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.

2. The method for pedestrian re-identification in large passenger flow monitoring of subway hubs according to claim 1, characterized in that, The specific process of S1 is as follows: S1.1, Visible light image Expanded into spatial-level images With channel-level images ; Single-channel infrared image Expanded into spatial-level images With channel-level images ; S1.2 Feature extractor via shared parameter network G Feature mapping is performed on the expanded multi-dimensional data, and channel concatenation is then executed as follows: ; in, This represents the joint feature tensor output by layer 0 of the network. These are the feature components of the visible light spatial dimension, the visible light channel dimension, the infrared spatial dimension, and the infrared channel dimension, respectively.

3. The method for pedestrian re-identification in large passenger flow monitoring of subway hubs according to claim 1 or 2, characterized in that, The specific process of S2 is as follows: For a given feature map Three sets of directional convolutions are introduced, each containing a directional mask and weights for the grouped convolutions. ,in m Indicates direction, g Indicates the number of groups; the orientation mask includes vertical, horizontal, and diagonal lines, each defined as follows: ; in, To represent two diagonal masks, denoted uniformly as... ; Automatic broadcasting will connect each mask to its corresponding Multiply, A mask for vertical, horizontal, and two diagonal lines. The local response along each direction is calculated as follows: ; in, This indicates that element-wise multiplication is broadcast along the spatial dimension; the mask is used to ensure that the output highlights edge responses, and the edge difference term is defined as: ; ; in, These represent the adaptive weight coefficients of the masks in the horizontal, vertical, and diagonal directions, respectively, where k represents the feature dimension. These represent the grouped convolution weights in the horizontal, vertical, and diagonal directions, respectively. This represents the adaptive parameters, i.e., the spatial summation of the directional convolution kernels. Further explanation: Assume the sample number is n The output channel element is o ,but Represented as: ; in, This indicates the number of input channels in each group. r For local channel indexing, g ( o () is a grouping function used to determine the current output channel. o The group index to which it belongs; It is shared in spatial location to map edge responses from different modalities to a unified energy metric space; The dot product yields the weighted energy at each location along the channel direction, reflecting the offset caused by the inherent local intensity deviation between different modes; edge information differences contain identity-related information, and an energy coefficient is introduced to ensure that useful information is retained while suppressing bias. ; in, It is a convolutional layer whose weights and biases are determined by... Derivation; Assume the desired directional response of a given output channel is represented as the sum of edge information and bias term: ; Since the deviation is an estimate, Cannot be precisely separated ,Right now: ; in, Representing edge information The energy coefficient, Indicates the deviation term The energy coefficient, when Approximately 1 and When smaller, Strongly suppressed, representing stable structural information Then it can be preserved. The edge residuals are considered calibrated; the edge information difference bridging module is applied to the shallow layers of the network through the residual terms: ; in, This is an adjustable parameter; Introduced as a correction term with a small weight, it preserves the main structural information while making Overlay This improves the overall quality of the features.

4. The pedestrian re-identification method for monitoring large passenger flows in subway hubs according to claim 3, characterized in that, The specific process of S3 is as follows: S3.1 Map features from different levels to a unified embedding space, and then apply multi-scale deep convolution to transform local representations into prior information with rich receptive fields: ; Where z is a set of different convolution kernel sizes, and Z represents the effect of the convolution kernel; S3.

2. Utilize features rich in edge information to guide semantic learning and construct an edge-aware association matrix with edge geometric similarity. : ; in, For grouped convolutional layers, The temperature factor is used; to reduce semantic interference, a dual-branch Top-K sparsity strategy is introduced, along with... Geometric guidance is implemented in the feature direction, starting only from deep features. In-process polymerization and Most relevant Context token : ; ; in, To control Sparsity is a scalar parameter. The set of indices obtained from the Top-K operations. These represent the indices of the "multi-head" dimension in the attention mechanism, and the attention correlation matrix, respectively. Row index and attention association matrix column index, I This is an indicator function that returns 1 when the condition is true and 0 otherwise; sparse branches are fused using learnable weights. ; in, These are learnable scalar weights; By utilizing features rich in edge information as prior knowledge, the contribution intensity of semantically sparse context is dynamically balanced.

5. The pedestrian re-identification method for monitoring large passenger flows in subway hubs according to claim 3, characterized in that, The overall network loss function in S4 includes identity recognition loss. and quality-perceived triplet loss Among them, the quality-perceived triplet loss is used to perceive sample quality, and its formal definition is: ; Where N is the batch size. Let represent the feature vectors of the anchor sample, positive sample, and negative sample of the i-th triplet, respectively; Represents Euclidean distance. margin The preset distance threshold; These are adaptive weighting coefficients; To further integrate quality perception into the optimization process, a dynamic modulation factor is introduced. Adjust the weight sensitivity based on the structural reliability of the triplet: ; in It is the cardinality of the quality perception factor. It is a scaling constant. This is used to quantify the quality of the feature relationships between anchor samples, positive samples, and negative samples; the overall network loss function is obtained as follows: ; in, It is a cross-entropy loss function: ; Where i is the category number, Let N be the label of the i-th category, and N be the number of images input to the model in the current iteration.

6. A system for pedestrian re-identification in large passenger flow monitoring of a subway hub as described in any one of claims 1 to 5, comprising an edge computing device and a multimodal camera, characterized in that, The edge computing device is equipped with a feature diversity module, an edge information difference bridging module, an edge information-driven shallow feature injection module, and a multi-quality-aware triplet loss module that are connected in sequence. The multimodal camera is used to capture streaming images and transmit the captured images to the edge computing device; The edge computing device is used to receive streaming images and process the images through the deployed feature diversity module, edge information difference bridging module, edge information-driven shallow feature injection module, and multi-quality perception triplet loss module. The feature diversity module is used to enhance the features of the input visible light and infrared images, copy and expand the single-channel infrared data into three channels, construct a tensor format consistent with the RGB input, extract spatial information through standard convolution kernels in the spatial dimension, and obtain channel information through channel normalization in the channel dimension. The edge information difference bridging module obtains the local edge strength deviation of the feature through the directional mask capture feature diversity module, then eliminates the directional specificity while retaining the edge response strength to quantify the edge difference, and performs information bridging through the energy coefficient; The edge information-driven shallow feature injection module enhances the density of local edge information purified by the edge information difference bridging module through a parallel sparsification mechanism, and selectively injects deep semantic representations to promote the interaction between shallow and deep edge perception features. The quality-perceived triplet loss supervises and constrains the quality of edge-perceived features by perceiving the quality of samples, pushing away the feature centers of different modalities in the same subspace, while bringing closer the feature centers of different dimensions of the same modality.