Pedestrian identification method and system based on multi-modal video data

By extracting features from multimodal video data and modeling dynamic graph structures, combined with graph attention fusion networks, the problem of recognition accuracy in complex environments of single-modal pedestrian recognition systems is solved, achieving high accuracy and robustness in pedestrian recognition under multimodal data.

CN120954046APending Publication Date: 2025-11-14GUANGZHOU POWER SUPPLY BUREAU GUANGDONG POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511022399.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Existing single-modal pedestrian recognition systems are easily affected by factors such as changes in lighting, occlusion, changes in viewing angle, and background interference in complex environments, leading to decreased recognition accuracy and misidentification. Furthermore, existing multimodal pedestrian recognition methods fail to effectively utilize pedestrian feature relationships across videos and scenes, limiting the system's performance in large-scale complex video surveillance environments.

Method used

By acquiring a multimodal video dataset, including RGB frame sequences, IR frame sequences, depth data, and pedestrian attribute data, visual, infrared, depth, and attribute features are extracted using ResNet50, CNN, 3D-CNN, and MLP neural networks. A dynamic graph structure is constructed, and feature fusion is performed using a graph attention fusion network to finally generate pedestrian recognition results.

Benefits of technology

It significantly improves the accuracy and robustness of pedestrian recognition, can accurately model the spatiotemporal relationships between pedestrians in complex environments, overcomes feature space differences, enhances feature representation capabilities, and improves the accuracy and stability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120954046A_ABST
    Figure CN120954046A_ABST
Patent Text Reader

Abstract

The invention provides a pedestrian recognition method and system based on multi-modal video data, and the method comprises the steps: obtaining a multi-modal data set, and inputting the multi-modal data set into a preset feature extraction model, enabling the feature extraction model to use neural networks of different structures to extract features of different modes of each pedestrian from the multi-mode data set, carrying out the alignment fusion of each feature, and outputting a multi-mode feature vector of each pedestrian; constructing a dynamic graph structure according to each multi-modal feature vector; inputting the dynamic graph structure into a preset graph attention fusion network so as to enable the graph attention fusion network to perform feature fusion based on a multi-head attention mechanism and a topological relation among the nodes, and obtaining a fusion feature matrix of the dynamic graph structure; and inputting the fused feature matrix into a preset inference engine, so that the inference engine generates a corresponding pedestrian recognition result according to the fused feature matrix, and the accuracy of pedestrian recognition is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of deep learning and pedestrian recognition technology, and in particular to a pedestrian recognition method and system based on multimodal video data. Background Technology

[0002] With the increasing demand for video surveillance systems in fields such as urban public safety management, intelligent transportation, and business behavior analysis, video-based pedestrian recognition technology has become an important research direction in computer vision and artificial intelligence. Traditional pedestrian recognition methods mainly rely on single-modal information in images, such as human feature extraction and matching based on visible light or infrared images. These methods typically use convolutional neural networks (CNNs) to extract pedestrian image features and determine pedestrian identity through metric learning or classification models. However, single-modal pedestrian recognition systems have many problems in practical applications. For example, they are easily affected by environmental factors such as changes in lighting, occlusion, changes in viewing angle, and background interference, which can lead to decreased recognition accuracy and false recognition.

[0003] In recent years, multimodal pedestrian recognition technology has gained increasing attention to improve the robustness and accuracy of pedestrian recognition systems in complex environments. Multimodal pedestrian recognition integrates heterogeneous data from multiple sources (such as visible light, infrared, depth maps, video temporal information, and pedestrian attribute labels) to comprehensively characterize pedestrian identity features, effectively mitigating the limitations of single-modal recognition schemes in specific environments. Common multimodal fusion methods include feature-level fusion, decision-level fusion, and model-level joint learning methods. However, most existing multimodal pedestrian recognition methods rely on static feature fusion, lacking modeling of the potential relationships between pedestrians and failing to fully utilize pedestrian feature relationships across videos and scenes, thus limiting the system's performance in large-scale, complex video surveillance environments. Meanwhile, Graph Neural Networks (GNNs), due to their ability to model non-Euclidean structured data, have achieved significant results in social network analysis, recommender systems, and traffic prediction, and are gradually being introduced into pedestrian re-identification and pedestrian video recognition tasks. Graph Neural Networks (GNNs) can improve recognition accuracy and system robustness by constructing similarity graphs between pedestrian samples and mining contextual relationships and potential identity associations between samples. However, current graph network-based pedestrian recognition methods are mostly focused on single-modal or static image scenes, with limited research on graph structure modeling and joint feature learning for multimodal video data. A unified and effective multimodal pedestrian video recognition scheme has not yet been formed. Summary of the Invention

[0004] To address the aforementioned technical problems, this application provides a pedestrian recognition method and system based on multimodal video data, thereby improving the accuracy of pedestrian recognition.

[0005] In a first aspect, embodiments of this application provide a pedestrian recognition method based on multimodal video data, including:

[0006] Obtain a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data;

[0007] The multimodal dataset is input into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian.

[0008] A dynamic graph structure is constructed based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian.

[0009] The dynamic graph structure is input into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node to obtain the fusion feature matrix of the dynamic graph structure.

[0010] The fused feature matrix is ​​input into a preset inference engine so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include spatiotemporal trajectory prediction results, action state prediction results, and identity matching results for each pedestrian.

[0011] This application provides a pedestrian recognition method based on multimodal video data. It extracts features by fusing multimodal video data, constructs a dynamic graph structure based on the extracted multimodal feature vectors, and finally performs pedestrian recognition based on the dynamic graph structure and graph attention mechanism to generate corresponding pedestrian recognition results. Existing pedestrian recognition systems mostly rely on single-modal information such as visible light or infrared images, which are easily affected by complex environmental factors such as changes in illumination, occlusion, pose changes, camera perspective differences, and background interference, leading to a significant decrease in recognition accuracy and failing to meet the high stability and accuracy requirements of complex video surveillance environments. This application, however, combines multimodal video data to construct a dynamic graph structure, which can accurately model the spatiotemporal relationships between pedestrians, significantly improving the accuracy and robustness of pedestrian recognition in complex environments. Through the aggregation mechanism of the graph attention fusion network, the correlation between different modal features is automatically learned during the interaction process of each node, effectively overcoming the feature space difference problem, improving multimodal fusion performance, enhancing feature representation capabilities, and thus improving the accuracy of pedestrian recognition.

[0012] Furthermore, the acquisition of the multimodal dataset includes:

[0013] Several visible light video streams within a preset time period are acquired using several visible light cameras, and the RGB frame sequence data is extracted from each of the visible light video streams.

[0014] Several infrared video streams within a preset time period are acquired using several infrared cameras, and the IR frame sequence data is extracted from each of the infrared video streams.

[0015] Several depth map sequences within a preset time period are acquired using several depth cameras, and the depth data is extracted from each of the depth map sequences.

[0016] Each of the visible light video streams is input into a preset YOLOv8 model so that the YOLOv8 model generates pedestrian attribute data for each pedestrian.

[0017] The multimodal dataset is constructed by combining the RGB frame sequence data, the IR frame sequence data, the depth data, and the pedestrian attribute data.

[0018] This application provides a method for acquiring a multimodal dataset. It involves deploying several visible light cameras, several infrared cameras, and several depth cameras to acquire RGB frame sequence data, IR frame sequence data, and depth data within a preset time period. Additionally, a YOLOv8 model is used to perform preliminary identification of the visible light video stream, obtaining pedestrian attribute data for each pedestrian, thus completing the construction of the multimodal dataset. In this application embodiment, multi-source data acquisition (visible light, infrared, and depth cameras) ensures data diversity and adapts to different lighting and occlusion scenarios. The attribute data generated by YOLOv8 supplements semantic information, enhancing the comprehensiveness of pedestrian feature description. Specifically, this application embodiment considers the reduced performance of single-image modality recognition when visible light imaging quality is poor, such as at night or in inclement weather. Therefore, IR frame sequence data and depth data are introduced to compensate for missing, interfered, or abnormal modalities by relying on information from other modalities, thereby improving the accuracy of pedestrian recognition.

[0019] In one possible implementation, the feature extraction model uses neural networks with different structures to extract visual features, infrared features, depth features, and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses these features to output a multimodal feature vector for each pedestrian, including:

[0020] Several first feature vectors are extracted from the RGB frame sequence data in the multimodal dataset using a preset ResNet50 network to construct the visual features of each pedestrian.

[0021] Several second feature vectors are extracted from the IR frame sequence data in the multimodal dataset using a pre-set CNN convolutional network to construct the infrared features of each pedestrian.

[0022] Several third feature vectors are extracted from the depth data in the multimodal dataset using a pre-defined 3D-CNN convolutional network to construct the depth features of each pedestrian.

[0023] Several encoded attributes are extracted from the pedestrian attribute data in the multimodal dataset using a preset MLP encoding model to construct the attribute features of each pedestrian.

[0024] The visual features, infrared features, depth features, and attribute features of each pedestrian are sequentially subjected to feature concatenation, feature dimensionality reduction, and regularization operations to obtain a multimodal feature vector of each pedestrian with a preset dimension.

[0025] This application provides a feature extraction method that uses dedicated neural networks ResNet50, CNN, 3D-CNN, and MLP to extract different modal features based on the characteristics of each modal feature, ensuring the professionalism and independence of each modal feature. At the same time, the extracted features are aligned and fused to reduce data redundancy between different modal features and improve the accuracy and computational efficiency of subsequent pedestrian recognition.

[0026] In one possible implementation, constructing the dynamic graph structure based on each of the multimodal feature vectors includes:

[0027] Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian;

[0028] The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed.

[0029] Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge.

[0030] Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established;

[0031] The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix.

[0032] Furthermore, the step of calculating the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge based on each of the multimodal feature vectors includes:

[0033] Based on each of the multimodal feature vectors, the cosine similarity between the two nodes corresponding to each of the connecting edges is calculated, and the cosine similarity includes visual feature cosine similarity and infrared feature cosine similarity.

[0034] Calculate the visual similarity weight of each connecting edge based on the cosine similarity of each edge;

[0035] Optical flow estimation is performed on the RGB frame sequence data, and the cosine value of the angle between the two optical flow vectors corresponding to the two nodes of each connection edge is calculated;

[0036] Calculate the motion consistency weight of each connecting edge based on the cosine value of each included angle;

[0037] Based on each of the multimodal feature vectors, calculate the attribute feature matching degree between the two nodes of each connection edge;

[0038] The attribute association weight of each connection edge is calculated based on the matching degree of each attribute feature.

[0039] This application provides a method for constructing a dynamic graph structure. It establishes corresponding node feature matrices and weighted adjacency matrices using multimodal feature vectors, thereby constructing the dynamic graph structure. Existing pedestrian recognition methods mostly process individual pedestrian samples independently, ignoring the potential spatial, temporal, or attribute-based relationships between pedestrians and failing to effectively utilize contextual information between pedestrians, resulting in insufficient discrimination capability in large-scale video surveillance scenarios. However, this application, in constructing the weighted adjacency matrix, determines the spatiotemporal adjacency relationship between nodes by determining the time interval between each pedestrian appearing in the same camera. Furthermore, it calculates the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge to capture the complex interactions and dynamic changes between pedestrians, enhancing the accuracy of edge connections. This means that in subsequent pedestrian recognition, this embodiment does not predict the individual movement trajectory and state of a pedestrian in isolation, but fully considers the mutual influence and interaction between pedestrians, optimizing pedestrian interaction modeling and improving the accuracy of pedestrian recognition.

[0040] In one possible implementation, the graph attention fusion network performs feature fusion based on a multi-head attention mechanism and the topological relationships between nodes to obtain a fusion feature matrix of the dynamic graph structure, including:

[0041] Based on the topological relationship between each node, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features;

[0042] Each of the aggregated features is normalized by layer convolution and projected onto a preset dimension to obtain the corresponding normalized aggregated features;

[0043] The standardized aggregated features are fused across modalities using a multi-head attention mechanism to obtain the fused feature matrix of the dynamic graph structure.

[0044] This application provides a feature fusion method. Based on a dynamic graph structure, it aggregates the multimodal feature vectors of each node within the modality based on the topological relationships between nodes, obtaining four corresponding aggregated features. Then, based on a multi-head attention mechanism, it further fuses the four aggregated features across modalities to obtain the final fused feature matrix. Although multimodal pedestrian recognition methods have been proposed to alleviate the limitations of single-modal methods, existing methods mostly employ simple feature concatenation or linear weighted fusion, lacking in-depth exploration of the correlation and complementarity between different modal features, resulting in limited fusion effects and difficulty in fully leveraging the advantages of multimodal data. This application, after standardization through intramodal aggregation, effectively integrates heterogeneous features through multi-head attention cross-modal fusion, fully fusing useful information provided by each modality and the interaction information between pedestrians. In this process, the attention mechanism focuses on key information, further improving the uniformity and discriminativeness of feature representation, and enhancing the accuracy of pedestrian recognition.

[0045] Furthermore, based on the topological relationships between nodes, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features, including:

[0046] The multimodal feature vectors of each node in the dynamic graph structure are traversed and updated to generate the updated feature vectors of each node. For any current node, several neighboring nodes of the current node are determined. Then, based on the comprehensive weight of the connection edges between the current node and each of the neighboring nodes, the multimodal feature vectors of each of the neighboring nodes are weighted and superimposed onto the current node to generate the updated feature vector of the current node.

[0047] The updated visual features, updated infrared features, updated depth features, and updated attribute features from the updated feature vectors of each node are respectively input into a preset softmax function, so that the softmax function performs feature concatenation on the updated feature vectors of each node according to the modality type to obtain four corresponding aggregated features.

[0048] In the process of intramodal feature aggregation, the embodiments of this application first update the features based on the weighted superposition of neighbor nodes, and then use softmax to concatenate the updated feature vectors intramodally to achieve efficient feature aggregation. While preserving the core information of each modality and avoiding information loss, the features of each node are also updated with weights based on the comprehensive weight of the connecting edges, which enhances the spatiotemporal representation capability of the nodes and improves the accuracy of subsequent pedestrian recognition.

[0049] In one possible implementation, the inference engine generates a corresponding pedestrian recognition result based on the fused feature matrix, including:

[0050] A pedestrian distribution map is generated based on the preset deployment locations of several visible light cameras. The pedestrian distribution map includes several sub-maps, each sub-map corresponding to one of the visible light cameras, and each node in the sub-map corresponding to a pedestrian.

[0051] Traverse the pedestrian distribution map, determine the time interval of the same pedestrian appearing in different cameras according to the fusion feature matrix, and then establish the connection edges between each sub-map and determine the edge weights to construct the pedestrian motion distribution map;

[0052] Based on the fused feature matrix and the pedestrian motion distribution map, pedestrian identification is performed on each pedestrian, and pedestrian identification results for each pedestrian are generated by reasoning.

[0053] For any current pedestrian, a probability diffusion is performed on the pedestrian movement distribution map based on a preset random walk algorithm and the fused feature matrix to predict the current pedestrian's movement path and generate a corresponding spatiotemporal trajectory prediction result.

[0054] Based on the temporal feature changes of the current pedestrian in the fusion feature matrix, several behavioral patterns of the current pedestrian on the corresponding movement path are predicted to obtain the prediction result of the current pedestrian's action state. The behavioral patterns include standing, walking, running, or gathering.

[0055] Based on feature similarity matching, the current pedestrian is matched with several pedestrian identities in a preset database to obtain the identity matching result of the current pedestrian.

[0056] This application provides a method for generating pedestrian recognition results. By constructing a pedestrian movement distribution map, it accurately represents the spatiotemporal correlation between various cameras and combines this with a fused feature matrix to infer pedestrian recognition results for each individual pedestrian. This application supports pedestrian feature matching across different cameras and different scenes. It improves the consistency of pedestrian features from different perspectives through graph network aggregation, and reduces interference from different cameras, perspectives, and lighting conditions by fusing multimodal data, ensuring the accuracy of pedestrian recognition. Simultaneously, it integrates random walk, temporal feature analysis, and feature matching to achieve unified pedestrian trajectory prediction, behavior analysis, and identity matching, effectively improving the global reasoning capability for pedestrian movement patterns and identities.

[0057] Secondly, correspondingly, embodiments of this application provide a pedestrian recognition system based on multimodal video data, including an acquisition module, a feature extraction module, a graph construction module, a graph attention fusion module, and an inference module;

[0058] The acquisition module is used to acquire a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data.

[0059] The feature extraction module is used to input the multimodal dataset into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian.

[0060] The graph construction module is used to construct a dynamic graph structure based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian.

[0061] The graph attention fusion module is used to input the dynamic graph structure into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node to obtain the fusion feature matrix of the dynamic graph structure;

[0062] The inference module is used to input the fused feature matrix into a preset inference engine, so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include spatiotemporal trajectory prediction results, action state prediction results, and identity matching results for each pedestrian.

[0063] Furthermore, the graph construction module constructs a dynamic graph structure based on each of the multimodal feature vectors, including:

[0064] Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian;

[0065] The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed.

[0066] Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge.

[0067] Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established;

[0068] The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix. Attached Figure Description

[0069] Figure 1 A flowchart illustrating a pedestrian recognition method based on multimodal video data provided in this application embodiment;

[0070] Figure 2 Another flowchart illustrating a pedestrian recognition method based on multimodal video data provided in this application embodiment;

[0071] Figure 3 A schematic diagram illustrating the feature extraction process in a pedestrian recognition method based on multimodal video data provided in this application embodiment;

[0072] Figure 4 This is a flowchart illustrating the construction of a dynamic graph structure in a pedestrian recognition method based on multimodal video data, as provided in an embodiment of this application.

[0073] Figure 5 A schematic diagram illustrating the graph-based feature fusion process in a pedestrian recognition method based on multimodal video data provided in this application embodiment;

[0074] Figure 6 A flowchart illustrating deep inference in a pedestrian recognition method based on multimodal video data provided in this application embodiment;

[0075] Figure 7 This is a schematic diagram of the structure of a pedestrian recognition system based on multimodal video data, provided in an embodiment of this application. Detailed Implementation

[0076] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0077] It should be noted that the step numbers in this document are only for the convenience of explaining the specific embodiments and are not intended to limit the order in which the steps are performed. In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0078] Example 1:

[0079] like Figure 1 As shown, Embodiment 1 provides a pedestrian recognition method based on multimodal video data, including steps S1-S5:

[0080] Step S1: Obtain a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data;

[0081] Step S2: Input the multimodal dataset into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian.

[0082] Step S3: Construct a dynamic graph structure based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian.

[0083] Step S4: Input the dynamic graph structure into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node, and obtains the fusion feature matrix of the dynamic graph structure;

[0084] Step S5: Input the fused feature matrix into a preset inference engine so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include the spatiotemporal trajectory prediction results, action state prediction results, and identity matching results of each pedestrian.

[0085] This application provides a pedestrian recognition method based on multimodal video data. It extracts features by fusing multimodal video data, constructs a dynamic graph structure based on the extracted multimodal feature vectors, and finally performs pedestrian recognition based on the dynamic graph structure and graph attention mechanism to generate corresponding pedestrian recognition results. Existing pedestrian recognition systems mostly rely on single-modal information such as visible light or infrared images, which are easily affected by complex environmental factors such as changes in illumination, occlusion, pose changes, camera perspective differences, and background interference, leading to a significant decrease in recognition accuracy and failing to meet the high stability and accuracy requirements of complex video surveillance environments. This application, however, combines multimodal video data to construct a dynamic graph structure, which can accurately model the spatiotemporal relationships between pedestrians, significantly improving the accuracy and robustness of pedestrian recognition in complex environments. Through the aggregation mechanism of the graph attention fusion network, the correlation between different modal features is automatically learned during the interaction process of each node, effectively overcoming the feature space difference problem, improving multimodal fusion performance, enhancing feature representation capabilities, and thus improving the accuracy of pedestrian recognition.

[0086] Furthermore, in step S1, obtaining the multimodal dataset includes:

[0087] Several visible light video streams within a preset time period are acquired using several visible light cameras, and the RGB frame sequence data is extracted from each of the visible light video streams.

[0088] Several infrared video streams within a preset time period are acquired using several infrared cameras, and the IR frame sequence data is extracted from each of the infrared video streams.

[0089] Several depth map sequences within a preset time period are acquired using several depth cameras, and the depth data is extracted from each of the depth map sequences.

[0090] Each of the visible light video streams is input into a preset YOLOv8 model so that the YOLOv8 model generates pedestrian attribute data for each pedestrian.

[0091] The multimodal dataset is constructed by combining the RGB frame sequence data, the IR frame sequence data, the depth data, and the pedestrian attribute data.

[0092] This application provides a method for acquiring a multimodal dataset. It involves deploying several visible light cameras, several infrared cameras, and several depth cameras to acquire RGB frame sequence data, IR frame sequence data, and depth data within a preset time period. Additionally, a YOLOv8 model is used to perform preliminary identification of the visible light video stream, obtaining pedestrian attribute data for each pedestrian, thus completing the construction of the multimodal dataset. In this application embodiment, multi-source data acquisition (visible light, infrared, and depth cameras) ensures data diversity and adapts to different lighting and occlusion scenarios. The attribute data generated by YOLOv8 supplements semantic information, enhancing the comprehensiveness of pedestrian feature description. Specifically, this application embodiment considers the reduced performance of single-image modality recognition when visible light imaging quality is poor, such as at night or in inclement weather. Therefore, IR frame sequence data and depth data are introduced to compensate for missing, interfered, or abnormal modalities by relying on information from other modalities, thereby improving the accuracy of pedestrian recognition.

[0093] In a preferred embodiment, such as Figure 2 As shown, the system acquires visible light video streams, infrared video streams, depth map sequences, and pedestrian attribute labels within a preset time period. These data are then processed sequentially by multiple modules to obtain the final pedestrian recognition result. The visible light video stream in the input layer provides pedestrian appearance feature data, typically provided by an RGB video stream. This data can be used for pedestrian recognition and pose estimation. The infrared video stream provides an image stream, which can be used to assist recognition in low-light or nighttime environments. This data is primarily used to capture the thermal radiation characteristics of pedestrians, enhancing recognition capabilities in complex environments. The depth map sequence mainly consists of depth data acquired through a depth camera, providing 3D structural information of the scene and helping to identify pedestrian distance and spatial distribution. The pedestrian attribute labels contain detailed pedestrian information and can supplement personalized pedestrian labels, providing information for further recognition.

[0094] In one possible implementation, in step S2, the feature extraction model uses neural networks of different structures to extract visual features, infrared features, depth features, and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses these features to output a multimodal feature vector for each pedestrian, including:

[0095] Several first feature vectors are extracted from the RGB frame sequence data in the multimodal dataset using a preset ResNet50 network to construct the visual features of each pedestrian.

[0096] Several second feature vectors are extracted from the IR frame sequence data in the multimodal dataset using a pre-set CNN convolutional network to construct the infrared features of each pedestrian.

[0097] Several third feature vectors are extracted from the depth data in the multimodal dataset using a pre-defined 3D-CNN convolutional network to construct the depth features of each pedestrian.

[0098] Several encoded attributes are extracted from the pedestrian attribute data in the multimodal dataset using a preset MLP encoding model to construct the attribute features of each pedestrian.

[0099] The visual features, infrared features, depth features, and attribute features of each pedestrian are sequentially subjected to feature concatenation, feature dimensionality reduction, and regularization operations to obtain a multimodal feature vector of each pedestrian with a preset dimension.

[0100] This application provides a feature extraction method that uses dedicated neural networks ResNet50, CNN, 3D-CNN, and MLP to extract different modal features based on the characteristics of each modal feature, ensuring the professionalism and independence of each modal feature. At the same time, the extracted features are aligned and fused to reduce data redundancy between different modal features and improve the accuracy and computational efficiency of subsequent pedestrian recognition.

[0101] In a preferred embodiment, such as Figure 3 As shown, the feature extraction model processes visible light video streams, infrared video streams, depth maps, and pedestrian attribute features respectively. It extracts features using convolutional neural networks (CNNs) and other deep learning methods, and further processes the features by aligning them through cross-modal contrastive loss. Specifically, the model includes the following steps:

[0102] (1) Preprocessing

[0103] The RGB frame is adjusted to a dimension of 3x224x224 and then normalized; the IR frame is adjusted to a dimension of 3x224x224 through dynamic range stretching and three-channel pseudo-color; the depth map is adjusted to a dimension of 1x224x224 through truncation and logarithmic compression; and the attributes are adjusted to 224xd (d represents the number of attributes) through one-hot encoding.

[0104] (2) Feature extraction

[0105] The RGB image processing branch uses a pre-trained ResNet50 deep convolutional neural network to extract a 1x2048-dimensional feature vector from the RGB three-channel image. The infrared image processing branch uses an untrained ResNet50 network structure to extract a 1x2048-dimensional feature vector from a single-channel infrared image. Training begins with randomly initialized weights. The depth map processing branch utilizes a lightweight dedicated convolutional network containing 7x7 convolutional layers, a Swish activation function, and residual blocks to extract a 1x2048-dimensional feature vector from the depth image. Attribute encoding requires first encoding categorical attributes. Discrete attributes such as gender are processed through embedding layers to map categorical data to a 1x1024-dimensional feature vector. Numerical attributes such as age are then processed through fully connected layers and mapped to a 1x1024-dimensional feature vector. Finally, feature fusion is used to concatenate these into a 1x2048-dimensional feature vector.

[0106] (3) Align intermediate stage data

[0107] The original features are concatenated as 2048+2048+2048+2048; then, a fully connected (FC) layer is used to reduce the dimensionality from 1x8192 to 1x2048; finally, the L2-regularized vector is set to 2048xd; finally, a comparative loss is calculated, and the cross-modal features are aligned according to the loss to output the final multimodal feature vector. The loss calculation formula is as follows:

[0108]

[0109] In one possible implementation, step S3, which involves constructing a dynamic graph structure based on each of the multimodal feature vectors, includes:

[0110] Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian;

[0111] The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed.

[0112] Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge.

[0113] Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established;

[0114] The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix.

[0115] Furthermore, the step of calculating the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge based on each of the multimodal feature vectors includes:

[0116] Based on each of the multimodal feature vectors, the cosine similarity between the two nodes corresponding to each of the connecting edges is calculated, and the cosine similarity includes visual feature cosine similarity and infrared feature cosine similarity.

[0117] Calculate the visual similarity weight of each connecting edge based on the cosine similarity of each edge;

[0118] Optical flow estimation is performed on the RGB frame sequence data, and the cosine value of the angle between the two optical flow vectors corresponding to the two nodes of each connection edge is calculated;

[0119] Calculate the motion consistency weight of each connecting edge based on the cosine value of each included angle;

[0120] Based on each of the multimodal feature vectors, calculate the attribute feature matching degree between the two nodes of each connection edge;

[0121] The attribute association weight of each connection edge is calculated based on the matching degree of each attribute feature.

[0122] This application provides a method for constructing a dynamic graph structure. It establishes corresponding node feature matrices and weighted adjacency matrices using multimodal feature vectors, thereby constructing the dynamic graph structure. Existing pedestrian recognition methods mostly process individual pedestrian samples independently, ignoring the potential spatial, temporal, or attribute-based relationships between pedestrians and failing to effectively utilize contextual information between pedestrians, resulting in insufficient discrimination capability in large-scale video surveillance scenarios. However, this application, in constructing the weighted adjacency matrix, determines the spatiotemporal adjacency relationship between nodes by determining the time interval between each pedestrian appearing in the same camera. Furthermore, it calculates the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge to capture the complex interactions and dynamic changes between pedestrians, enhancing the accuracy of edge connections. This means that in subsequent pedestrian recognition, this embodiment does not predict the individual movement trajectory and state of a pedestrian in isolation, but fully considers the mutual influence and interaction between pedestrians, optimizing pedestrian interaction modeling and improving the accuracy of pedestrian recognition.

[0123] In a preferred embodiment, such as Figure 4As shown, a dynamic graph structure is constructed through three steps: graph node construction, edge weight calculation, and dynamic updating, representing the temporal and spatial relationships between pedestrians. Each node represents a pedestrian, and each edge represents the spatiotemporal relationship between pedestrians. The construction of the dynamic graph enables the model to capture the complex interactions and dynamic changes between pedestrians, especially in multi-person scenarios, allowing for dynamic adjustment of the graph structure to reflect environmental changes. The specific process is as follows:

[0124] (1) Initialization of graph structure

[0125] Node definition: Each detected pedestrian instance is treated as a graph node (V), and the node features are composed of a 2048xd unified feature vector output by the multimodal feature extraction module.

[0126] Edge establishment rules: Initial edges (E) are established based on spatiotemporal proximity. A connection is established when two pedestrians appear in the same camera's field of view and the time interval is less than Δt.

[0127] (2) Dynamic calculation of edge weights

[0128] Visual similarity weight (W_visual): Calculate the cosine similarity of RGB / IR features, while introducing a deep feature difference penalty term.

[0129] Motion consistency weight (W_motion): Based on the cosine of the angle between the motion vectors estimated by optical flow, while also introducing a penalty term for the difference in velocity magnitude.

[0130] Attribute association weight (W_attribute): Jaccard similarity is used to calculate the matching degree of attributes such as gender / age between pedestrians, while co-occurrence statistics of carried items are introduced (such as association reinforcement when both are carrying backpacks).

[0131] (3) Dynamic update mechanism

[0132] Temporal sliding window: edge weights are recalculated every T frames (typically T=10).

[0133] Adaptive decay: Apply exponential decay to edges that have not been updated for a long time.

[0134] Cross-camera association: When the target leaves the camera's field of view (FOV), retain the node state for τ seconds (typically τ=5).

[0135] Use ReID feature matching to establish edges across cameras.

[0136] (4) Graph structure optimization

[0137] Sparsification: Preserve the top-k strong connections (k=5).

[0138] Conflict resolution: When multiple nodes with the same ID are detected, merge nodes with similar characteristics.

[0139] Memory management: Remove orphaned nodes that have exceeded their lifespan (typically 60 seconds).

[0140] (5) Output of dynamic graph structure

[0141] Real-time graph representation: Output weighted adjacency matrix A∈R^(N×N) and node feature matrix X∈R^(N×2048).

[0142] In one possible implementation, in step S4, the graph attention fusion network performs feature fusion based on a multi-head attention mechanism and the topological relationships between nodes to obtain a fusion feature matrix of the dynamic graph structure, including:

[0143] Based on the topological relationship between each node, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features;

[0144] Each of the aggregated features is normalized by layer convolution and projected onto a preset dimension to obtain the corresponding normalized aggregated features;

[0145] The standardized aggregated features are fused across modalities using a multi-head attention mechanism to obtain the fused feature matrix of the dynamic graph structure.

[0146] This application provides a feature fusion method. Based on a dynamic graph structure, it aggregates the multimodal feature vectors of each node within the modality based on the topological relationships between nodes, obtaining four corresponding aggregated features. Then, based on a multi-head attention mechanism, it further fuses the four aggregated features across modalities to obtain the final fused feature matrix. Although multimodal pedestrian recognition methods have been proposed to alleviate the limitations of single-modal methods, existing methods mostly employ simple feature concatenation or linear weighted fusion, lacking in-depth exploration of the correlation and complementarity between different modal features, resulting in limited fusion effects and difficulty in fully leveraging the advantages of multimodal data. This application, after standardization through intramodal aggregation, effectively integrates heterogeneous features through multi-head attention cross-modal fusion, fully fusing useful information provided by each modality and the interaction information between pedestrians. In this process, the attention mechanism focuses on key information, further improving the uniformity and discriminativeness of feature representation, and enhancing the accuracy of pedestrian recognition.

[0147] Furthermore, based on the topological relationships between nodes, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features, including:

[0148] The multimodal feature vectors of each node in the dynamic graph structure are traversed and updated to generate the updated feature vectors of each node. For any current node, several neighboring nodes of the current node are determined. Then, based on the comprehensive weight of the connection edges between the current node and each of the neighboring nodes, the multimodal feature vectors of each of the neighboring nodes are weighted and superimposed onto the current node to generate the updated feature vector of the current node.

[0149] The updated visual features, updated infrared features, updated depth features, and updated attribute features from the updated feature vectors of each node are respectively input into a preset softmax function, so that the softmax function performs feature concatenation on the updated feature vectors of each node according to the modality type to obtain four corresponding aggregated features.

[0150] In the process of intramodal feature aggregation, the embodiments of this application first update the features based on the weighted superposition of neighbor nodes, and then use softmax to concatenate the updated feature vectors intramodally to achieve efficient feature aggregation. While preserving the core information of each modality and avoiding information loss, the features of each node are also updated with weights based on the comprehensive weight of the connecting edges, which enhances the spatiotemporal representation capability of the nodes and improves the accuracy of subsequent pedestrian recognition.

[0151] In a preferred embodiment, such as Figure 5 As shown, after constructing the graph structure, a graph attention mechanism is used for weighted fusion of information. This mechanism dynamically adjusts the weights of information transmission based on the similarity and relationships between nodes, focusing on key pedestrian nodes and spatiotemporal relationships to improve the model's recognition and reasoning capabilities. This process automatically selects the most relevant features and relationships for fusion, achieving more accurate behavior analysis and identity matching. The specific process is as follows:

[0152] (1) Intramodal feature aggregation

[0153] The node identifiers for each of the four modes are as follows: The representation of each node is calculated as follows: Where Nu represents the neighbors of node u, a(u, v) is the attention score between nodes u and v, obtained through training the GATv2 model, σ is a non-linear activation function, and finally the features of each node are concatenated using softmax to obtain the corresponding modal feature vector.

[0154] (2) Hierarchical integration and cross-modal interaction

[0155] The four-modal features are normalized by convolutional layers, unifying the projection to 256-d. Each 256-1 vector is then linearly projected into three vectors: Q (query identifier), K (key identifier), and V (value identifier). Finally, a four-head attention mechanism is used for cross-modal association. Here, dk is used as a scaling factor, and finally the fused feature fg is obtained after passing through the laynorm layer.

[0156] In one possible implementation, in step S5, the inference engine generates a corresponding pedestrian recognition result based on the fused feature matrix, including:

[0157] A pedestrian distribution map is generated based on the preset deployment locations of several visible light cameras. The pedestrian distribution map includes several sub-maps, each sub-map corresponding to one of the visible light cameras, and each node in the sub-map corresponding to a pedestrian.

[0158] Traverse the pedestrian distribution map, determine the time interval of the same pedestrian appearing in different cameras according to the fusion feature matrix, and then establish the connection edges between each sub-map and determine the edge weights to construct the pedestrian motion distribution map;

[0159] Based on the fused feature matrix and the pedestrian motion distribution map, pedestrian identification is performed on each pedestrian, and pedestrian identification results for each pedestrian are generated by reasoning.

[0160] For any current pedestrian, a probability diffusion is performed on the pedestrian movement distribution map based on a preset random walk algorithm and the fused feature matrix to predict the current pedestrian's movement path and generate a corresponding spatiotemporal trajectory prediction result.

[0161] Based on the temporal feature changes of the current pedestrian in the fusion feature matrix, several behavioral patterns of the current pedestrian on the corresponding movement path are predicted to obtain the prediction result of the current pedestrian's action state. The behavioral patterns include standing, walking, running, or gathering.

[0162] Based on feature similarity matching, the current pedestrian is matched with several pedestrian identities in a preset database to obtain the identity matching result of the current pedestrian.

[0163] This application provides a method for generating pedestrian recognition results. By constructing a pedestrian movement distribution map, it accurately represents the spatiotemporal correlation between various cameras and combines this with a fused feature matrix to infer pedestrian recognition results for each individual pedestrian. This application supports pedestrian feature matching across different cameras and different scenes. It improves the consistency of pedestrian features from different perspectives through graph network aggregation, and reduces interference from different cameras, perspectives, and lighting conditions by fusing multimodal data, ensuring the accuracy of pedestrian recognition. Simultaneously, it integrates random walk, temporal feature analysis, and feature matching to achieve unified pedestrian trajectory prediction, behavior analysis, and identity matching, effectively improving the global reasoning capability for pedestrian movement patterns and identities.

[0164] In a preferred embodiment, such as Figure 6 As shown, the spatiotemporal relationship reasoning engine STRE further utilizes graph structures and attention mechanisms for deep reasoning to infer pedestrians' spatiotemporal trajectories, behavioral models, and identity matching. By modeling spatiotemporal relationships, STRE can accurately infer pedestrians' movement paths, behavior types (such as walking, standing, etc.), and their interactions with pedestrians or the environment in complex scenes. This module integrates features from different modalities in spatiotemporal reasoning, improving the reliability and accuracy of the reasoning. The specific process is as follows:

[0165] (1) Spatiotemporal consistency constraint stage

[0166] Time difference calculation: Record the timestamps of the target appearing on different cameras and calculate the reasonable transfer time range.

[0167] Consistency score: Verify the credibility of cross-camera association based on factors such as movement speed and path rationality.

[0168] (2) Multi-camera association stage

[0169] Establishment of edges across cameras: Generate a pedestrian distribution map based on the preset deployment locations of several visible light cameras, and establish connecting edges between submaps of different cameras when the spatiotemporal constraints are met.

[0170] Association strength calculation: Combine ReID feature similarity and spatiotemporal matching to determine edge weights and construct a pedestrian movement distribution map.

[0171] (3) Reasoning optimization stage

[0172] Random walk algorithm: Based on the fused feature matrix, a probability diffusion is performed on the pedestrian movement distribution map to simulate the possible movement paths of the target.

[0173] Path scoring: The feasibility of the generated candidate paths is evaluated (avoiding obstacle areas, etc.), and the final spatiotemporal trajectory prediction result is determined.

[0174] Other identification results: Based on the fusion feature matrix, several behavioral patterns of each pedestrian on the corresponding movement path are inferred and predicted, and each pedestrian is matched with the pedestrian identity in the preset database to obtain the action state prediction result and identity matching result.

[0175] (4) Identity dissemination stage

[0176] Modal dynamic balancing: Automatically adjusts the contribution weight of different modalities (such as RGB / infrared) in identity matching.

[0177] Conflict resolution: When identity assignment conflicts occur, the propagation path with the highest confidence is selected.

[0178] (5) Feature caching mechanism

[0179] Short-term memory cache: Retains feature data from the last 5 minutes for fast retrieval.

[0180] Long-term feature database: Stores confirmed identity features in the database to support cross-time period queries.

[0181] In a preferred embodiment, the final pedestrian identification result includes pedestrian ID, spatiotemporal trajectory, behavior analysis, and identity matching:

[0182] 1. Pedestrian ID: Based on pedestrian characteristics and behavioral patterns, the STRE module assigns a unique ID to each pedestrian to complete identity recognition.

[0183] 2. Spatiotemporal Trajectory: Records the dynamic movement trajectory of pedestrians in a scene, which can effectively predict the future location and behavior of pedestrians.

[0184] 3. Behavior Analysis: Analyzing pedestrian behavior to identify their activity patterns (such as walking, standing, turning, etc.) to provide data support for subsequent behavior prediction and monitoring.

[0185] 4. Identity Matching: Based on multimodal features and spatiotemporal relationship reasoning, pedestrians are matched with the existing identity database to ensure multiple identity verifications in complex scenarios.

[0186] based on Figure 2 The steps shown are illustrated below, and the complete input / output flow is as follows:

[0187] 1. Input layer

[0188] Input: RGB frame sequence, single-channel heatmap sequence, depth value matrix sequence, pedestrian attribute labels.

[0189] Dimensions: [B,T,3,H,W], [B,T,1,H,W], [B,T,1,H,W] (batch size, time step, channel, height, width), [B,N,A] (batch size, number of pedestrians, attribute dimension).

[0190] 2. Multimodal Feature Extractor (MFE)

[0191] Processing steps:

[0192] Visible light branch: ResNet50 extracts spatial features → output [B,N,2048].

[0193] Infrared branch: Custom CNN extracts thermal features → Output [B,N,2048].

[0194] Deep branch: 3D-CNN extracts spatiotemporal features → output [B,N,512].

[0195] Attribute branch: MLP encoded attribute → output [B,N,128].

[0196] Output: Unified multimodal feature vector [B,N,1024].

[0197] 3. Dynamic Graph Structure Building Module (DGC)

[0198] Input: Multimodal features [B,N,1024] + spatiotemporal coordinates.

[0199] Processing steps:

[0200] Node creation: Create a graph node for each pedestrian.

[0201] Edge establishment: Connecting spatiotemporally adjacent nodes of the same camera.

[0202] Cross-camera association: Node connections that satisfy transfer constraints.

[0203] Edge weight calculation: feature similarity + spatiotemporal rationality.

[0204] Output: A dynamic graph data structure containing a node feature matrix: [N, 1024], an adjacency matrix: [N, N], and an edge weight matrix: [N, N].

[0205] 4. Graph Attention Fusion Network (GAFN)

[0206] Input: Dynamic graph structure (nodes + edges).

[0207] Processing steps:

[0208] Intramodal aggregation: GAT processing of each modal subgraph.

[0209] Cross-modal interaction: multi-head attention mechanism.

[0210] Residual fusion: Preserves original features.

[0211] Dynamic weight adjustment: based on ambient lighting conditions.

[0212] Output: fused feature matrix [N, 1024].

[0213] 5. Spatiotemporal Relationship Reasoning Engine (STRE)

[0214] Input: fused feature matrix [N, 1024].

[0215] Processing steps:

[0216] Trajectory reasoning: Random walk generates candidate paths.

[0217] Behavioral analysis: LSTM processes temporal feature changes.

[0218] Identity matching: Feature similarity calculation.

[0219] Conflict resolution: Confidence priority principle.

[0220] Output: Pedestrian ID: UUID list [N], Spatiotemporal trajectory: [N,P,4] (path points, coordinates + time), Behavior analysis: [N,C] (probabilities of C types of behavior), Identity matching: [M,2] (M matching pairs).

[0221] 6. Output layer

[0222] Pedestrian ID output: Format: ["ID_001","ID_002",...], Update rule: New ID is assigned for each new detection.

[0223] Spatiotemporal trajectory output: Format: {ID:[(x1,y1,t1),(x2,y2,t2),...]}, Length: retain the 100 most recent location points.

[0224] Behavioral analysis output: Category: ["Standing", "Walking", "Running", "Gathering"], Format: {ID:{"Behavior Type":Confidence}}.

[0225] Identity matching output: Format: [("cam1_ID5","cam2_ID7"),...], with cross-camera transfer time.

[0226] Furthermore, the pedestrian recognition method provided in this application embodiment has alternative solutions that can be further improved, including: using VIT instead of CNN to extract visible light image features, since VIT is strong at capturing long distances and improves recognition accuracy in complex scenes, but has a large computational cost; trying to use a multi-branch fusion network to replace the self-built infrared sub-network to save model parameters and improve cross-modal feature alignment capability; considering the use of TabTransformer or Feature-wiseLinear Modulation (FiLM) mechanism; introducing GRaphSAGE instead of GCN, as GraphSAGE can handle large-scale data, GAF dynamically adjusts neighbor weights, and DGCNN is suitable for point cloud data.

[0227] Example 2:

[0228] like Figure 7 As shown, Embodiment 2 provides a pedestrian recognition system based on multimodal video data, including an acquisition module 10, a feature extraction module 20, a graph construction module 30, a graph attention fusion module 40, and an inference module 50;

[0229] The acquisition module 10 is used to acquire a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data.

[0230] The feature extraction module 20 is used to input the multimodal dataset into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian.

[0231] The graph construction module 30 is used to construct a dynamic graph structure based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian.

[0232] The graph attention fusion module 40 is used to input the dynamic graph structure into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node to obtain the fusion feature matrix of the dynamic graph structure;

[0233] The inference module 50 is used to input the fused feature matrix into a preset inference engine, so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include spatiotemporal trajectory prediction results, action state prediction results, and identity matching results for each pedestrian.

[0234] Furthermore, the acquisition module 10 acquires the multimodal dataset, including:

[0235] Several visible light video streams within a preset time period are acquired using several visible light cameras, and the RGB frame sequence data is extracted from each of the visible light video streams.

[0236] Several infrared video streams within a preset time period are acquired using several infrared cameras, and the IR frame sequence data is extracted from each of the infrared video streams.

[0237] Several depth map sequences within a preset time period are acquired using several depth cameras, and the depth data is extracted from each of the depth map sequences.

[0238] Each of the visible light video streams is input into a preset YOLOv8 model so that the YOLOv8 model generates pedestrian attribute data for each pedestrian.

[0239] The multimodal dataset is constructed by combining the RGB frame sequence data, the IR frame sequence data, the depth data, and the pedestrian attribute data.

[0240] In one possible implementation, the feature extraction module 20 uses neural networks with different structures to extract visual features, infrared features, depth features, and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses these features to output a multimodal feature vector for each pedestrian, including:

[0241] Several first feature vectors are extracted from the RGB frame sequence data in the multimodal dataset using a preset ResNet50 network to construct the visual features of each pedestrian.

[0242] Several second feature vectors are extracted from the IR frame sequence data in the multimodal dataset using a pre-set CNN convolutional network to construct the infrared features of each pedestrian.

[0243] Several third feature vectors are extracted from the depth data in the multimodal dataset using a pre-defined 3D-CNN convolutional network to construct the depth features of each pedestrian.

[0244] Several encoded attributes are extracted from the pedestrian attribute data in the multimodal dataset using a preset MLP encoding model to construct the attribute features of each pedestrian.

[0245] The visual features, infrared features, depth features, and attribute features of each pedestrian are sequentially subjected to feature concatenation, feature dimensionality reduction, and regularization operations to obtain a multimodal feature vector of each pedestrian with a preset dimension.

[0246] In one possible implementation, the graph construction module 30 constructs a dynamic graph structure based on each of the multimodal feature vectors, including:

[0247] Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian;

[0248] The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed.

[0249] Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge.

[0250] Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established;

[0251] The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix.

[0252] Furthermore, the step of calculating the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge based on each of the multimodal feature vectors includes:

[0253] Based on each of the multimodal feature vectors, the cosine similarity between the two nodes corresponding to each of the connecting edges is calculated, and the cosine similarity includes visual feature cosine similarity and infrared feature cosine similarity.

[0254] Calculate the visual similarity weight of each connecting edge based on the cosine similarity of each edge;

[0255] Optical flow estimation is performed on the RGB frame sequence data, and the cosine value of the angle between the two optical flow vectors corresponding to the two nodes of each connection edge is calculated;

[0256] Calculate the motion consistency weight of each connecting edge based on the cosine value of each included angle;

[0257] Based on each of the multimodal feature vectors, calculate the attribute feature matching degree between the two nodes of each connection edge;

[0258] The attribute association weight of each connection edge is calculated based on the matching degree of each attribute feature.

[0259] In one possible implementation, the graph attention fusion module 40 performs feature fusion based on a multi-head attention mechanism and the topological relationships between nodes to obtain a fusion feature matrix of the dynamic graph structure, including:

[0260] Based on the topological relationship between each node, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features;

[0261] Each of the aggregated features is normalized by layer convolution and projected onto a preset dimension to obtain the corresponding normalized aggregated features;

[0262] The standardized aggregated features are fused across modalities using a multi-head attention mechanism to obtain the fused feature matrix of the dynamic graph structure.

[0263] Furthermore, based on the topological relationships between nodes, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features, including:

[0264] The multimodal feature vectors of each node in the dynamic graph structure are traversed and updated to generate the updated feature vectors of each node. For any current node, several neighboring nodes of the current node are determined. Then, based on the comprehensive weight of the connection edges between the current node and each of the neighboring nodes, the multimodal feature vectors of each of the neighboring nodes are weighted and superimposed onto the current node to generate the updated feature vector of the current node.

[0265] The updated visual features, updated infrared features, updated depth features, and updated attribute features from the updated feature vectors of each node are respectively input into a preset softmax function, so that the softmax function performs feature concatenation on the updated feature vectors of each node according to the modality type to obtain four corresponding aggregated features.

[0266] In one possible implementation, the inference module 50 generates a corresponding pedestrian recognition result based on the fused feature matrix, including:

[0267] A pedestrian distribution map is generated based on the preset deployment locations of several visible light cameras. The pedestrian distribution map includes several sub-maps, each sub-map corresponding to one of the visible light cameras, and each node in the sub-map corresponding to a pedestrian.

[0268] Traverse the pedestrian distribution map, determine the time interval of the same pedestrian appearing in different cameras according to the fusion feature matrix, and then establish the connection edges between each sub-map and determine the edge weights to construct the pedestrian motion distribution map;

[0269] Based on the fused feature matrix and the pedestrian motion distribution map, pedestrian identification is performed on each pedestrian, and pedestrian identification results for each pedestrian are generated by reasoning.

[0270] For any current pedestrian, a probability diffusion is performed on the pedestrian movement distribution map based on a preset random walk algorithm and the fused feature matrix to predict the current pedestrian's movement path and generate a corresponding spatiotemporal trajectory prediction result.

[0271] Based on the temporal feature changes of the current pedestrian in the fusion feature matrix, several behavioral patterns of the current pedestrian on the corresponding movement path are predicted to obtain the prediction result of the current pedestrian's action state. The behavioral patterns include standing, walking, running, or gathering.

[0272] Based on feature similarity matching, the current pedestrian is matched with several pedestrian identities in a preset database to obtain the identity matching result of the current pedestrian.

[0273] This application provides a pedestrian recognition system based on multimodal video data. It extracts features by fusing multimodal video data, constructs a dynamic graph structure based on the extracted multimodal feature vectors, and finally performs pedestrian recognition based on the dynamic graph structure and graph attention mechanism, generating corresponding pedestrian recognition results. Existing pedestrian recognition systems often rely on single-modal information such as visible light or infrared images, which are easily affected by complex environmental factors such as changes in lighting, occlusion, pose changes, camera perspective differences, and background interference, leading to a significant decrease in recognition accuracy and failing to meet the high stability and accuracy requirements of complex video surveillance environments. This application, however, combines multimodal video data to construct a dynamic graph structure, which can accurately model the spatiotemporal relationships between pedestrians, significantly improving the accuracy and robustness of pedestrian recognition in complex environments. Through the aggregation mechanism of the graph attention fusion network, the correlation between different modal features is automatically learned during the interaction process of each node, effectively overcoming the feature space difference problem, improving multimodal fusion performance, enhancing feature representation capabilities, and thus improving the accuracy of pedestrian recognition.

[0274] For a more detailed explanation of the working principle and procedures of this embodiment, please refer to the relevant description in Embodiment 1.

[0275] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of this application. It should be understood that the above descriptions are merely specific embodiments of this application and are not intended to limit the scope of protection of this application. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application for those skilled in the art.

Claims

1. A pedestrian recognition method based on multimodal video data, characterized in that, include: Obtain a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data; The multimodal dataset is input into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian. A dynamic graph structure is constructed based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian. The dynamic graph structure is input into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node to obtain the fusion feature matrix of the dynamic graph structure. The fused feature matrix is ​​input into a preset inference engine so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include spatiotemporal trajectory prediction results, action state prediction results, and identity matching results for each pedestrian.

2. The pedestrian recognition method based on multimodal video data as described in claim 1, characterized in that, The acquisition of the multimodal dataset includes: Several visible light video streams within a preset time period are acquired using several visible light cameras, and the RGB frame sequence data is extracted from each of the visible light video streams. Several infrared video streams within a preset time period are acquired using several infrared cameras, and the IR frame sequence data is extracted from each of the infrared video streams. Several depth map sequences within a preset time period are acquired using several depth cameras, and the depth data is extracted from each of the depth map sequences. Each of the visible light video streams is input into a preset YOLOv8 model so that the YOLOv8 model generates pedestrian attribute data for each pedestrian. The multimodal dataset is constructed by combining the RGB frame sequence data, the IR frame sequence data, the depth data, and the pedestrian attribute data.

3. The pedestrian recognition method based on multimodal video data as described in claim 1, characterized in that, The feature extraction model uses neural networks with different structures to extract visual features, infrared features, depth features, and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses these features to output a multimodal feature vector for each pedestrian, including: Several first feature vectors are extracted from the RGB frame sequence data in the multimodal dataset using a preset ResNet50 network to construct the visual features of each pedestrian. Several second feature vectors are extracted from the IR frame sequence data in the multimodal dataset using a pre-set CNN convolutional network to construct the infrared features of each pedestrian. Several third feature vectors are extracted from the depth data in the multimodal dataset using a pre-defined 3D-CNN convolutional network to construct the depth features of each pedestrian. Several encoded attributes are extracted from the pedestrian attribute data in the multimodal dataset using a preset MLP encoding model to construct the attribute features of each pedestrian. The visual features, infrared features, depth features, and attribute features of each pedestrian are sequentially subjected to feature concatenation, feature dimensionality reduction, and regularization operations to obtain a multimodal feature vector of each pedestrian with a preset dimension.

4. The pedestrian recognition method based on multimodal video data as described in claim 1, characterized in that, The construction of the dynamic graph structure based on each of the multimodal feature vectors includes: Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian; The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed. Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge. Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established; The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix.

5. The pedestrian recognition method based on multimodal video data as described in claim 4, characterized in that, The step of calculating the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge based on each of the multimodal feature vectors includes: Based on each of the multimodal feature vectors, the cosine similarity between the two nodes corresponding to each of the connecting edges is calculated, and the cosine similarity includes visual feature cosine similarity and infrared feature cosine similarity. Calculate the visual similarity weight of each connecting edge based on the cosine similarity of each edge; Optical flow estimation is performed on the RGB frame sequence data, and the cosine value of the angle between the two optical flow vectors corresponding to the two nodes of each connection edge is calculated; Calculate the motion consistency weight of each connecting edge based on the cosine value of each included angle; Based on each of the multimodal feature vectors, calculate the attribute feature matching degree between the two nodes of each connection edge; The attribute association weight of each connection edge is calculated based on the matching degree of each attribute feature.

6. The pedestrian recognition method based on multimodal video data as described in claim 1, characterized in that, The graph attention fusion network performs feature fusion based on a multi-head attention mechanism and the topological relationships between nodes to obtain a fusion feature matrix of the dynamic graph structure, including: Based on the topological relationship between each node, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features; Each of the aggregated features is normalized by layer convolution and projected onto a preset dimension to obtain the corresponding normalized aggregated features; The standardized aggregated features are fused across modalities using a multi-head attention mechanism to obtain the fused feature matrix of the dynamic graph structure.

7. The pedestrian recognition method based on multimodal video data as described in claim 6, characterized in that, Based on the topological relationships between nodes, the visual features, infrared features, depth features, and attribute features in each of the multimodal feature vectors are respectively subjected to intramodal feature aggregation to obtain four corresponding aggregated features, including: The multimodal feature vectors of each node in the dynamic graph structure are traversed and updated to generate the updated feature vectors of each node. For any current node, several neighboring nodes of the current node are determined. Then, based on the comprehensive weight of the connection edges between the current node and each of the neighboring nodes, the multimodal feature vectors of each of the neighboring nodes are weighted and superimposed onto the current node to generate the updated feature vector of the current node. The updated visual features, updated infrared features, updated depth features, and updated attribute features from the updated feature vectors of each node are respectively input into a preset softmax function, so that the softmax function performs feature concatenation on the updated feature vectors of each node according to the modality type to obtain four corresponding aggregated features.

8. The pedestrian recognition method based on multimodal video data as described in claim 1, characterized in that, The inference engine generates corresponding pedestrian recognition results based on the fused feature matrix, including: A pedestrian distribution map is generated based on the preset deployment locations of several visible light cameras. The pedestrian distribution map includes several sub-maps, each sub-map corresponding to one of the visible light cameras, and each node in the sub-map corresponding to a pedestrian. Traverse the pedestrian distribution map, determine the time interval of the same pedestrian appearing in different cameras according to the fusion feature matrix, and then establish the connection edges between each sub-map and determine the edge weights to construct the pedestrian motion distribution map; Based on the fused feature matrix and the pedestrian motion distribution map, pedestrian identification is performed on each pedestrian, and pedestrian identification results for each pedestrian are generated by reasoning. For any current pedestrian, a probability diffusion is performed on the pedestrian movement distribution map based on a preset random walk algorithm and the fused feature matrix to predict the current pedestrian's movement path and generate a corresponding spatiotemporal trajectory prediction result. Based on the temporal feature changes of the current pedestrian in the fusion feature matrix, several behavioral patterns of the current pedestrian on the corresponding movement path are predicted to obtain the prediction result of the current pedestrian's action state. The behavioral patterns include standing, walking, running, or gathering. Based on feature similarity matching, the current pedestrian is matched with several pedestrian identities in a preset database to obtain the identity matching result of the current pedestrian.

9. A pedestrian recognition system based on multimodal video data, characterized in that, It includes an acquisition module, a feature extraction module, a graph construction module, a graph attention fusion module, and an inference module; The acquisition module is used to acquire a multimodal dataset, which includes RGB frame sequence data, IR frame sequence data, depth data, and pedestrian attribute data. The feature extraction module is used to input the multimodal dataset into a preset feature extraction model, so that the feature extraction model uses neural networks with different structures to extract the visual features, infrared features, depth features and attribute features of each pedestrian from the multimodal dataset, and aligns and fuses the features to output the multimodal feature vector of each pedestrian. The graph construction module is used to construct a dynamic graph structure based on each of the multimodal feature vectors, wherein each node in the dynamic graph structure corresponds to each pedestrian, the node features of each node correspond to the multimodal feature vectors of each pedestrian, and each edge in the dynamic graph structure corresponds to the spatiotemporal relationship between each pedestrian. The graph attention fusion module is used to input the dynamic graph structure into a preset graph attention fusion network, so that the graph attention fusion network performs feature fusion based on the multi-head attention mechanism and the topological relationship between each node to obtain the fusion feature matrix of the dynamic graph structure; The inference module is used to input the fused feature matrix into a preset inference engine, so that the inference engine generates corresponding pedestrian recognition results based on the fused feature matrix. The pedestrian recognition results include spatiotemporal trajectory prediction results, action state prediction results, and identity matching results for each pedestrian.

10. A pedestrian recognition system based on multimodal video data as described in claim 9, characterized in that, The graph construction module constructs a dynamic graph structure based on each of the multimodal feature vectors, including: Based on the various multimodal feature vectors, establish a node feature matrix corresponding to each pedestrian; The time interval between the appearance of each pedestrian in the same camera is determined based on the multimodal feature vectors of each pedestrian, the spatiotemporal proximity relationship between each node is determined, and then the connection edges between each node are constructed. Based on each of the multimodal feature vectors, the visual similarity weight, motion consistency weight, and attribute association weight of each connecting edge are calculated, thereby determining the comprehensive weight of each connecting edge. Based on each of the connecting edges and the corresponding comprehensive weights, a corresponding weighted adjacency matrix is ​​established; The dynamic graph structure is obtained by combining the node feature matrix and the weighted adjacency matrix.