Group pedestrian re-identification method based on geographic spatial-temporal feature map attention network

By constructing a graph model based on geographical spatiotemporal features and integrating geographical location, time and behavioral information, the shortcomings in identifying group pedestrian behavior patterns and interactive relationships in complex traffic scenarios in the prior art are solved, and higher recognition accuracy and robustness are achieved.

CN120088815APending Publication Date: 2025-06-03XINJIANG UNIVERSITY +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510160014.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing pedestrian re-identification methods have shortcomings in identifying group pedestrian behavior patterns and interactive relationships in complex geographical space-time environments. Especially in traffic scenarios with complex geographical locations and dynamic changes, it is difficult to accurately identify the identity and behavior of group pedestrians.

Method used

A group pedestrian re-identification method based on the geo-spatial-temporal feature map attention network is adopted. By integrating precise geographical location, time and geographical spatial-temporal behavior information, a graph model between group pedestrians is constructed to comprehensively capture their interactive relationships and behavior patterns.

Benefits of technology

It improves the recognition accuracy and robustness of group pedestrians in complex traffic environments, and can more accurately capture and handle complex geographical and spatio-temporal behavior patterns and interactive relationships.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120088815A_ABST
    Figure CN120088815A_ABST
Patent Text Reader

Abstract

The invention discloses a group pedestrian re-identification method based on a geographic spatial-temporal feature map attention network, and belongs to the field of pedestrian re-identification. The method comprises the following steps: S1, collecting group image data; s2, obtaining a group image with a label; s3, constructing a BRT undirected graph, and extracting a time feature, a space feature and a speed feature; s4, extracting visual features, and integrating the visual features into a node feature matrix; s5, combining the time features, the space features and the speed features to form an edge feature matrix, and constructing a group pedestrian geographic space-time relation graph; s6, analyzing the geographic space-time relation graph of the group pedestrians to generate a prediction matrix; and S7, optimizing the accuracy of the prediction matrix. According to the method, a graph model among the group pedestrians is constructed by integrating accurate geographic position, time and geographic time-space behavior information, and the interaction relationship and behavior mode of the group pedestrians are comprehensively captured, so that the recognition precision and robustness of the group pedestrians in a complex traffic environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pedestrian re-identification, and in particular to a group pedestrian re-identification method based on a geographical spatio-temporal feature map attention network. Background Art

[0002] As an important technical means in the field of video surveillance, pedestrian re-identification solves the recognition difficulties caused by problems such as perspective changes and lighting differences in traditional surveillance by continuously tracking target individuals in cross-camera video data.

[0003] Currently, most pedestrian re-identification methods mainly rely on visual features, and some pedestrian re-identification methods are based on the fusion of spatio-temporal features and visual features. They have the following defects

[0004] 1. The combination of visual features and geographical location is not accurate enough: Currently, most pedestrian re-identification methods mainly rely on visual features and use simple spatio-temporal constraints (such as camera ID and timestamp) to assist in recognition. Although this method effectively reduces visual blur, it performs poorly in complex and dynamically changing traffic scenarios with respect to geographical location.

[0005] 2. Unable to handle complex geographical spatio-temporal behavior patterns: Current methods mostly use simple spatio-temporal constraints to assist visual matching and are unable to handle the geographical spatio-temporal behavior changes of group pedestrian behavior in complex traffic scenarios.

[0006] 3. Unable to capture the interaction relationships between group pedestrians: Pedestrians in a group, especially in complex scenarios, often have strong interdependencies, such as always acting together or maintaining a certain structure. Existing methods cannot well identify such dependencies, especially when the positions and relationships of pedestrians in the group are constantly changing, and existing recognition models will miss these dynamic interactions. Summary of the Invention

[0007] The purpose of the present invention is to provide a group pedestrian re-identification method based on a geographical spatio-temporal feature map attention network. This method constructs a graph model between group pedestrians by integrating accurate geographical location, time, and geographical spatio-temporal behavior information, comprehensively captures their interaction relationships and behavior patterns, thereby improving the recognition accuracy and robustness of group pedestrians in complex traffic environments.

[0008] To achieve the above purpose, the present invention provides a group pedestrian re-identification method based on a geographical spatio-temporal feature map attention network, including the following steps

[0009] S1, collect group image data, and then use the group image data to construct an ST-BRT dataset D, and the ST-BRT dataset D contains a training set D t and a test set D e ;

[0010] S2, assign a unified identity label to the group images in the training set D to obtain labeled group images; t in

[0011] S3, use the ST-BRT dataset D to construct a BRT undirected graph, and extract temporal features, spatial features, and speed features from the labeled group images;

[0012] S4, use the Vision Transformer model to extract visual features from the labeled group images and integrate them into a node feature matrix;

[0013] S5, combine the temporal features, spatial features, and speed features to form an edge feature matrix, and comprehensively use the edge feature matrix and the node feature matrix to construct a group pedestrian geo-spatiotemporal relationship graph;

[0014] S6, use the graph attention network model to analyze the group pedestrian geo-spatiotemporal relationship graph and generate a prediction matrix of the group global feature representation;

[0015] S7, calculate the cross-entropy loss between the prediction matrix and the true label matrix in the test set D e to optimize the accuracy of the prediction matrix.

[0016] Preferably, the expression of the ST-BRT dataset D in S1 is

[0017] D = D t ∪D e ;

[0018]

[0019] where |D| represents the number of group images in the dataset, |x i | is the number of internal group members in the i-th group image, x i,j represents the j-th member in the i-th group image, b i,j represents the bounding box coordinates of the j-th member in the image x i , represents the identity identification of the j-th member in the image x i , represents the group identity of the i-th group.

[0020] Preferably, in S2, the labeled group image is expressed as

[0021]

[0022] where x i represents the i-th group image, b iIndicates the bounding box annotation for each member. Indicates the group identity of the i-th group. Indicates the identity identifier for each member.

[0023] Preferably, in S3, the extracted temporal feature is

[0024] T = {T 12 , T 13 , …, T ij};

[0025] The spatial feature is

[0026] D = {D 12 , D 13 , …, D ij};

[0027] The velocity feature is

[0028] V = {V 12 , V 13 , …, V ij};

[0029] Where i and j represent the j-th member in the group image i.

[0030] Preferably, in S4, the process of extracting visual features is

[0031] S41, Cut the group pedestrian images in the labeled group image into a number of pixel blocks of fixed size;

[0032] S42, Project each pixel block through the linear projection function ψ into a feature vector R of dimension D D , The formula is as follows

[0033]

[0034] Where Represents the n-th pixel block of the group image i, Is the feature representation of this pixel block;

[0035] S43, Introduce a first-level token denoted as t f , And interact t f With each pixel block feature to obtain the overall appearance feature of the group pedestrian, the formula is as follows

[0036]

[0037] Where S i,j Represents the serialized feature of the group pedestrian in the group image i, E p Represents the position information encoding of each image block;

[0038] S44. Input the serialized features of the group images into the Vision Transformer model for optimization. The formula is as follows:

[0039]

[0040] Among them, represents the global appearance feature of the group pedestrian of the j-th member of the group image i;

[0041] S45. Integrate the global features of all group pedestrians into the node feature matrix H. The formula is as follows:

[0042]

[0043] Among them, N represents a total of N groups, and each group image i has M i members;

[0044] S46. Since the number of members in each group is different, use the all-zero vector 0 ∈ R D to complete the feature matrix, and then integrate the group pedestrian features. The formula is as follows:

[0045]

[0046] Among them, h i represents the node feature vector corresponding to the i-th group, and Flatten means splicing the feature vectors corresponding to multiple groups into one vector;

[0047] S47. Obtain the integrated node feature matrix H node ,

[0048]

[0049] Preferably, in S5, the fusion process of the time feature, space feature, and speed feature is as follows

[0050] S51. Determine the time feature T between group pedestrian i and group pedestrian j ij

[0051] T ij = |t i - t j |;

[0052] Among them, t i represents the time of group i, and t j represents the time of group j;

[0053] Determine the space feature D between group pedestrian i and group pedestrian j ij

[0054] Dij = D[i][j];

[0055] Among them, D is the distance matrix between stations, and D[i][j] represents the BRT distance from station to station;

[0056] Determine the speed characteristics V of group pedestrian i and group pedestrian j ij

[0057]

[0058] S52. Integrate the time, space, and speed characteristics to calculate the spatio-temporal characteristics F of each edge ij

[0059]

[0060] S53. Calculate the attention weight A using the time and space characteristics ij , and the formula is as follows

[0061]

[0062] Among them, Q = F ij W Q , Q represents the feature F ij The query vector obtained after linear transformation, K = F ij W K , K represents the feature F ij The key vector obtained after linear transformation, W Q and W K Are learnable projection matrices respectively, and d k Represents the dimension of the key vector;

[0063] S54. Use the speed characteristic V ij To perform feature integration, and the formula is as follows

[0064] E ij = A ij .V;

[0065] Among them, V = F ij W V , V represents the value vector obtained after the feature F ij After linear transformation W V ;

[0066] S55. Incorporate the integrated features into an edge feature matrix E, and the formula is as follows

[0067]

[0068] Preferably, in S6, the analysis process of the graph attention network for the group pedestrian geo-spatiotemporal relationship graph is as follows

[0069] S61. Perform a linear transformation on the node feature matrix, and the formula is as follows

[0070]

[0071] Wherein, W is a learnable weight matrix, F' is the transformed feature dimension, and H' is the transformed node feature matrix;

[0072] S62. Calculate the attention score e between any two nodes i and j through the combination of node features and edge features ij , and the formula is as follows

[0073]

[0074] a is a learnable weight vector, h i ′ and h j ′ are the features of nodes j and j respectively, and E ij is the edge feature between nodes i and j;

[0075] S63. Normalize the attention scores of the neighbor nodes of each node i, and the formula is as follows

[0076]

[0077] Where N i represents the neighbor set of node i, and α ij is the normalized attention weight;

[0078] S64. Use the normalized attention weights to perform weighted summation on the features of neighbor nodes, aggregate and update the features of each node, and the formula is as follows

[0079]

[0080] Wherein, σ is the activation function, and N i represents the neighbor set of node i;

[0081] Use the multi-head attention mechanism to obtain the final node features, and the formula is as follows

[0082]

[0083] Where K is the number of attention heads, is the feature of node i calculated by the k-th attention head.

[0084] Preferably, in S7, the cross-entropy loss function is used to measure the difference between the class probabilities predicted by the model and the true classes. The formula of the cross-entropy loss function is as follows

[0085]

[0086] where y i,c represents the true label, represents the probability that the i-th sample of the prediction model belongs to class c, C represents the total number of classes, and N represents the total number of samples.

[0087] Therefore, the group pedestrian re-identification method based on the geographical spatio-temporal feature graph attention network with the above structure has the following advantages

[0088] (1) The graph attention network is introduced, and weights are assigned to different nodes through an adaptive learning mechanism. The activities of group pedestrians in the urban geographical space scenario, especially bus travel, present graph structure characteristics, which highly match the expression mode of the graph model, making it an effective method for realizing the group pedestrian re-identification task.

[0089] (2) By combining the geographical spatio-temporal behavior pattern with the graph model, the dynamic migration relationship of group pedestrians in the geographical space-time can be accurately modeled. The robustness of group pedestrian re-identification is improved by dynamically adjusting the number of nodes and edge weights.

[0090] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. Description of the Drawings

[0091] Figure 1 is a schematic diagram of the group pedestrian re-identification method based on the geographical spatio-temporal feature graph attention network of the present invention;

[0092] Figure 2 is a flowchart of the group pedestrian re-identification method based on the geographical spatio-temporal feature graph attention network of the present invention. Detailed Embodiments

[0093] Embodiment

[0094] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0095] Accordingly, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0096] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0097] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is customarily placed during use. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0098] In the description of the present invention, it should also be noted that unless otherwise clearly specified and defined, the terms "set", "install", "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0099] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0100] As Figure 1 , Figure 2 shown, a group pedestrian re-identification method based on a geographical spatio-temporal feature map attention network according to the present invention includes the following steps

[0101] S1. Collect group image data, and then use the group image data to construct an ST-BRT data set D. The ST-BRT data set D includes a training set D t and a test set D e ; The expression of the ST-BRT data set D in S1 is

[0102] D = D t ∪ D e ;

[0103]

[0104] Among them, |D| represents the number of group images in the dataset, |X i | is the number of internal group members in the i-th group image, and x i,j represents the j-th member in the i-th group image, and b i,j represents the bounding box coordinates of the j-th member in the image x i . represents the identity identifier of the j-th member in the image x i . represents the group identity of the i-th group.

[0105] S2. Assign a unified identity label to the group images in the training set D to obtain labeled group images;

[0106] In S2, the labeled group images are represented as

[0107]

[0108] where x i represents the i-th group image, b i represents the bounding box annotation of each member, represents the group identity of the i-th group, represents the identity identifier of each member.

[0109] S3. Use the ST-BRT dataset D to construct a BRT undirected graph, and extract temporal features, spatial features, and velocity features from the labeled group images;

[0110] In S3, the extracted temporal features are

[0111] T = {T 12 , T 13 , …, T ij};

[0112] Each group image is associated with a time t i , which represents the specific time when the image was taken. The time information provides an important constraint for the recognition process and can effectively exclude false matches that occur at different time periods. The time difference between group image i and group image j is calculated, and the formula is as follows

[0113] d t (t i , t j ) = |t i - t j |;

[0114] Each group image also contains geographical coordinates p i = (lat i , lng i) to provide information about the shooting location. Geographical location information is particularly suitable for large-scale surveillance scenarios and helps to determine the movement paths and location associations of group pedestrians under different cameras.

[0115] The spatial feature is

[0116] D = {D 12 , D 13 , …, D ij};

[0117] Each group image also has a corresponding speed information, which can be deduced from the position information of multiple frames of images. The speed information can further enhance the spatio-temporal constraints because it not only reflects the positions of group pedestrians but also describes their movement trends.

[0118] The speed feature is

[0119] V = {V 12 , V 13 , …, V ij};

[0120] where i and j represent the j-th member in the i-th group image.

[0121] To perform group pedestrian re-identification, the most important thing is to obtain a deep model (F·; θ) that contains learnable parameters and can show high inter-group differences in group pedestrian features.

[0122] Specifically, the model (F·; θ) accurately classifies the group structure by inputting the group image x i and the constraint information (t i , p i , v i ) and maximizes the average probability of correct classification for all group images.

[0123]

[0124] S4. Use the Vision Transformer model to extract visual features from the labeled group images and integrate them into a node feature matrix;

[0125] In S4, the process of extracting visual features is

[0126] In the graph model, each node represents the appearance feature of a group pedestrian. Therefore, use the Vision Transformer model to extract the visual features of group pedestrians from the image data and generate the global appearance representation of each group member.

[0127] S41. Cut the group pedestrian images in the labeled group images into several pixel blocks of fixed size;

[0128] S42. Project each pixel block into a feature vector R of dimension D through the linear projection function ψ D , as shown in the following formula

[0129]

[0130] where represents the nth pixel block of the group image i, is the feature representation of this pixel block;

[0131] S43. Introduce a first-level token denoted as t f , and interact t f with each pixel block feature to obtain the overall appearance feature of the group pedestrians, as shown in the following formula

[0132]

[0133] where S i,j represents the serialized feature of the group pedestrians in the group image i, and E p represents the position information encoding of each image block;

[0134] Introducing a learnable first-level token aims to generate the global appearance feature of each group of pedestrians. The cls-token acts as a representation of the global feature in the Vision Transformer model. By interacting with each local pixel block feature, the overall appearance feature of the group pedestrians is learned.

[0135] S44. Input the serialized feature of the group image into the Vision Transformer model for optimization. After introducing the cls-token, the serialized feature of the group pedestrians is input into the Vision Transformer model. The multi-layer self-attention mechanism of the Vision Transformer model interacts the cls-token with the local features, learns and aggregates the local information, so as to obtain the global appearance feature of the group pedestrians. After multiple layers of processing, the cls-token will accumulate the feature information of all local pixel blocks to generate the optimized group appearance representation, as shown in the following formula

[0136]

[0137] where represents the global appearance feature of the group pedestrians of the jth member of the group image i;

[0138] S45. Integrate the global features of all group pedestrians into a node feature matrix H. To incorporate the global appearance features of group pedestrians into the calculation of the graph attention network, the global features of all group pedestrians are integrated into a node feature matrix, as shown in the following formula

[0139]

[0140] where N represents a total of N groups, and each group image i has M i members;

[0141] S46. Since the number of members M in the group i may be different, it is necessary to pad the feature matrix. Here, the all-zero vector 0 ∈ R D is used to pad the feature matrix, and then the features of group pedestrians are integrated, as shown in the following formula

[0142]

[0143] where h i represents the node feature vector corresponding to the i-th group, and Flatten means concatenating the feature vectors corresponding to multiple groups into one vector;

[0144] S47. Obtain the integrated node feature matrix H node ,

[0145]

[0146] S5. Combine the time feature, space feature, and speed feature to form an edge feature matrix, and construct a geographical spatio-temporal relationship graph of group pedestrians by integrating the edge feature matrix and the node feature matrix;

[0147] In S5, the fusion process of the time feature, space feature, and speed feature is

[0148] S51. Determine the time feature T ij

[0149] T ij = |t i - t j |;

[0150] where t i represents the time of group i, and t j represents the time of group j;

[0151] The space feature S ij measures the geographical distance between two groups. In this embodiment, the data in the ST-BRT dataset is used to determine the geographical distance between group pedestrians and determine the space feature D between group pedestrian i and group pedestrian jij

[0152] D ij = D[i][j];

[0153] Among them, D is the distance matrix between stations, and D[i][j] represents the BRT distance between stations;

[0154] The speed feature is a combination of time and space features, used to capture the dynamic movement patterns of group pedestrians, and determine the speed feature V of group pedestrian i and group pedestrian j ij

[0155]

[0156] S52. Fuse the time, space, and speed features to calculate the spatio-temporal feature F of each edge ij , Define the spatio-temporal feature of each edge, and the formula is as follows

[0157]

[0158] S53. The time feature and the space feature can stably describe the global relationship between nodes (such as the frequency of migration and geographical distribution), so use the time and space features to calculate the attention weight A ij , The formula is as follows

[0159]

[0160] Among them, Q = F ij W Q , Q represents the feature F ij The query vector obtained by linear transformation of, K = F ij W K , K represents the feature F ij The key vector obtained by linear transformation of, W Q and W K are learnable projection matrices respectively, and d k represents the dimension of the key vector, used to normalize the attention score and prevent the gradient from vanishing due to excessive numerical values;

[0161] S54. The speed feature is essentially a local combination of time and space, which can dynamically reflect the immediate behavior characteristics between nodes (such as whether the moving speed at a certain moment is abnormal), so use the speed feature V ij to perform feature fusion, and the formula is as follows

[0162] E ij = A ij ·V;

[0163] Among them, V = F ij WV where \(V\) represents the feature \(F\) ij after a linear transformation \(W\) V to obtain the value vector;

[0164] S55. To incorporate the comprehensive features of time, space, and speed into the calculation of the graph attention network, the comprehensive features are integrated into an edge feature matrix, and the fused features are incorporated into an integrated edge feature matrix \(E\), as shown in the following formula

[0165]

[0166] S6. Use the graph attention network model to analyze the group pedestrian geo-spatiotemporal relationship graph and generate a prediction matrix of the group global feature representation;

[0167] The core idea of Graph Attention Networks (GAT) is to dynamically calculate the interaction weights between nodes through the self-attention mechanism and use these weights to aggregate the features of neighbor nodes to generate a new node feature representation. Here, taking the node feature matrix \(H\) node and the edge feature matrix \(E\) as inputs.

[0168] In S6, the analysis process of the graph attention network for the group pedestrian geo-spatiotemporal relationship graph is

[0169] S61. Perform a linear transformation on the node feature matrix, as shown in the following formula

[0170]

[0171] where \(W\) is a learnable weight matrix, \(F'\) is the transformed feature dimension, and \(H'\) is the transformed node feature matrix;

[0172] S62. Calculate the attention score \(e\) for any two nodes \(i\) and \(j\) through the combination of node features and edge features ij , as shown in the following formula

[0173]

[0174] \(a\) is a learnable weight vector, \(h'\) i and \(h'\) j are the features of nodes \(i\) and \(i\) respectively, and \(E\) ij is the edge feature between nodes \(i\) and \(j\);

[0175] S63. To make the attention scores comparable among different neighbors, normalize the attention scores of the neighbor nodes of each node \(i\), as shown in the following formula

[0176]

[0177] where $N$ i represents the neighbor set of node $i$, and $\alpha$ ij is the normalized attention weight;

[0178] S64, using the normalized attention weights to weight-sum the features of neighbor nodes, aggregating and updating the features of each node, the formula is as follows

[0179]

[0180] where $\sigma$ is the activation function, and $N$ i represents the neighbor set of node $i$;

[0181] Through the aggregation operation, the features of each node not only contain its own information but also incorporate the information of neighbor nodes.

[0182] To enhance the model's expressive power, GAT usually adopts the multi-head attention mechanism. Specifically, each attention head independently calculates a set of attention weights and updates the features, and the final node features are obtained by concatenating the multi-head outputs, the concatenated node feature matrix. The final node features are obtained using the multi-head attention mechanism, and the formula is as follows

[0183]

[0184] where $K$ is the number of attention heads, is the feature of node $i$ calculated by the $k$-th attention head.

[0185] S7, calculating the cross-entropy loss between the prediction matrix and the true label matrix in the test set $D$ e to optimize the accuracy of the prediction matrix.

[0186] In S7, the group pedestrian recognition is transformed from a similarity-based retrieval problem into a multi-classification task, that is, according to the multi-modal features of each group of pedestrians, assigning a class label to it. To optimize the model to accurately classify each group of pedestrians, the cross-entropy loss function is used to measure the difference between the class probabilities predicted by the model and the true classes. The cross-entropy loss function formula is as follows

[0187]

[0188] where $y$ i,c represents the true label, represents the probability that the $i$-th sample of the prediction model belongs to class $C$, $C$ represents the total number of classes, and $N$ represents the total number of samples.

[0189] Therefore, the present invention adopts the above-mentioned group pedestrian re-identification method based on the geographical spatio-temporal feature map attention network. This method constructs a graph model among group pedestrians by integrating accurate geographical location, time, and geographical spatio-temporal behavior information, comprehensively capturing their interaction relationships and behavior patterns, thereby improving the recognition accuracy and robustness of group pedestrians in complex traffic environments.

[0190] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A group person re-identification method based on geographic spatiotemporal feature graph attention network, characterized by: The following steps are included S1, collect group image data, and then use the group image data to construct the ST-BRT dataset D, which contains a training set D t and a test set D e ; S2, is the training set D t The group images in are given a unified identity label to obtain a labeled group image; S3, construct a BRT undirected graph using the ST-BRT dataset D, and extract temporal features, spatial features, and speed features from labeled group images; S4, uses the Vision Transformer model to extract visual features from labeled group images and integrates them into a node feature matrix; S5, combining the time feature, space feature and speed feature to form an edge feature matrix, and constructing a group pedestrian geographic spatiotemporal relationship graph by integrating the edge feature matrix and the node feature matrix; S6, using the graph attention network model to parse the spatial-temporal relationship graph of the group of pedestrians and generate a prediction matrix representing the global features of the group; S7, the prediction matrix and the test set D e The cross entropy loss of the true label matrix in is calculated to optimize the accuracy of the prediction matrix.

2. According to claim 1, a group pedestrian re-identification method based on geographic spatiotemporal feature graph attention network is characterized by: S1 The expression of ST-BRT dataset D is: D=D t ∪D e ; Where |D| represents the number of group images in the dataset, |x i | is the number of internal group members in the i-th group image, x i,j represents the jth member in the jth group image, b i,j Denotes image x i The bounding box coordinates of the jth member in , Represents image x i The identity of the jth member in represents the group identity of the i-th group.

3. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 2, characterized in that: In S2, the labeled group images are represented as Among them, x i represents the i-th group image, b i represents the bounding box annotation of each member, represents the group identity of the i-th group, Indicates the identity of each member.

4. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 3, characterized in that: In S3, the extracted The time feature is T={T 12 ,T 13 ,…,T ij }; The spatial characteristics are D={D 12 ,D 13 ,…,D ij }; The speed characteristic is V={V 12 ,V 13 ,…,V ij }; Where i and j represent the jth member in group image i.

5. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 4, characterized in that: In S4, the process of extracting visual features is: S41, dividing the group pedestrian image in the labeled group image into a plurality of pixel blocks of fixed size; S42, project each pixel block into a feature vector R with dimension D through a linear projection function ψ. D , the formula is as follows in, represents the nth pixel block of group image i, is the feature representation of the pixel block; S43, introduce the first-level token denoted as t f , t f By interacting with each pixel block feature, the overall appearance feature of the group of pedestrians is obtained. The formula is as follows Among them, S i,j represents the sequential features of the group pedestrians in group image i, E p Indicates the position information encoding of each image block; S44, the serialized features of the group image are input into the Vision Transformer model for optimization. The formula is as follows in, Represents the global appearance features of the group pedestrians of the jth member of group image i; S45, integrate the global features of all group pedestrians into the node feature matrix H, the formula is as follows Among them, N means there are N groups in total, and each group image i has M i members; S46, because the number of members in each group is different, use the all-zero vector 0∈R D The feature matrix is ​​completed, and then the group pedestrian features are integrated. The formula is as follows Among them, h i represents the node feature vector corresponding to the i-th group, and Flatten means concatenating the feature vectors corresponding to multiple groups into one vector; S47, get the integrated node feature matrix H node , 6. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 5, characterized in that: In S5, the fusion process of the time feature, the space feature and the speed feature is S51, and the time feature T between the pedestrian group i and the pedestrian group j is determined. ij T ij =|t i -t j |; Among them, t i represents the time of group i, t j represents the time of group j; Determine the spatial feature D between group pedestrians i and group pedestrians j ij D ij =D[i][j]; Where D is a distance matrix between stations, and D[i][j] represents the BRT distance between stations; Determine the velocity characteristics V of group pedestrians i and group pedestrian j ij S52, integrate the time, space, and speed features to calculate the geographic spatiotemporal features F of each edge ij S53, use temporal and spatial features to calculate the attention weight A ij , the formula is as follows Where Q = F ij W Q , Q represents the feature F ij The query vector obtained by linear transformation, K = F ij W K , K represents feature F ij The key vector obtained by linear transformation, W Q and W K are respectively the learnable projection matrices, d k represents the dimension of the key vector; S54, use speed feature V ij To perform feature fusion, the formula is as follows E ij =A ij ·V; Where V = F ij W V , V represents feature F ij After linear transformation W V The value vector obtained after S55, the fused features are integrated into an edge feature matrix E, the formula is as follows 7. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 6, characterized in that: In S6, the graph attention network analyzes the spatial-temporal relationship graph of group pedestrians as follows: S61, linear transformation of the node feature matrix is ​​performed, the formula is as follows in, W is the learnable weight matrix, F' is the transformed feature dimension, and H' is the transformed node feature matrix; S62, calculate the attention score e of any two nodes i and j by combining node features and edge features ij , the formula is as follows a is the learnable weight vector, h' i and h' j are the features of nodes i and j respectively, E ij is the edge feature between nodes i and j; S63, normalize the attention scores of neighboring nodes of each node i, the formula is as follows Where N i represents the neighbor set of node i, α ij is the normalized attention weight; S64, use the normalized attention weights to weight the features of neighboring nodes, aggregate and update the features of each node. The formula is as follows Among them, σ is the activation function, N i represents the neighbor set of node i; Use the multi-head attention mechanism to get the final node features. The formula is as follows Where K is the number of attention heads, is the feature of node i computed by the kth attention head.

8. The method for group person re-identification based on geographic spatiotemporal feature graph attention network according to claim 7, characterized in that: In S7, the cross entropy loss function is used to measure the difference between the category probability predicted by the model and the actual category. The cross entropy loss function formula is as follows Among them, y i,c represents the true label, It represents the probability that sample i of the prediction model belongs to category c, C represents the total number of categories, and N represents the total number of samples.