A pedestrian intention prediction method for complex traffic scenarios
Through the feature fusion of local flow module, global flow module and cross-attention module, the problem of insufficient feature fusion of pedestrian intention prediction in complex traffic scenarios is solved, and the prediction accuracy and model expression ability are improved.
Patent Information
- Application Number
- CN202510544521.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing pedestrian intention prediction methods have problems such as insufficient local and global features integration, insufficient cross-current information interaction, and insufficient multi-scale feature capture in complex traffic scenarios, resulting in insufficient prediction accuracy.
Local flow module, global flow module, cross-attention module and feature fusion module are adopted to extract and fuse local and global features through convolutional neural networks, long and short-term memory networks, multi-head attention mechanisms and adaptive multi-scale convolutional blocks to enhance feature interaction and multi-scale expression capabilities.
It significantly improves the accuracy of pedestrian intention prediction and the expressive ability of the model, reduces information loss, and enhances the generalization ability and robustness of the model.
Smart Images

Figure CN120071262B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of pedestrian intention prediction, and specifically relates to a pedestrian intention prediction method for complex traffic scenes. Background Art
[0002] Pedestrian intention prediction is a key link in autonomous driving technology. By identifying and predicting pedestrians' behaviors such as whether they are crossing the road in advance, it can effectively avoid potential collision risks and improve traffic safety.
[0003] In recent years, deep learning-based methods combined with computer vision technology have provided new possibilities for pedestrian intention prediction. By using a large amount of pedestrian behavior data to train the model, researchers can extract effective features to improve prediction accuracy. Therefore, the research on pedestrian intention prediction based on deep learning has important theoretical and practical significance.
[0004] Chinese patent application number CN118781575A discloses a pedestrian intention prediction method based on a multimodal fusion strategy of a Transformer structure. This prediction method emphasizes the use of strong coupling and cross-attention mechanisms between different modalities. Although this prediction method theoretically helps to enhance the correlation between the modalities, this fusion strategy may face greater challenges in practical applications, especially when the quality of multimodal information is unstable, which may weaken the generalization ability and robustness of the model.
[0005] Chinese patent application number CN117152209A discloses a pedestrian trajectory prediction and intention estimation method based on dual-stream LSTM. This estimation method combines relative displacement technology to simulate the dynamic relationship between pedestrians and cars, captures the interaction between pedestrians and cars, and relies on pre-defined boundary lines. However, this estimation method may lead to inaccurate boundary line delineation or error accumulation caused by dynamic environmental changes. The information fusion of trajectory stream and intention stream depends on the interaction between the two, and this interaction process may be complex and not easy to optimize.
[0006] In summary, the following technical problems exist in the prior art: (1) Existing pedestrian intention prediction methods often process local features and global features independently, and fail to fully integrate local and global information, resulting in limited ability to accurately predict pedestrian behavior; (2) Traditional prediction methods find it difficult to effectively interact information between local feature streams and global feature streams, resulting in weak interdependence between feature streams, which affects the performance of the model. (3) Existing prediction methods lack multi-scale semantic information in the feature fusion process and cannot fully capture features at different levels, which affects the expressiveness and generalization capabilities of the model. (4) Traditional prediction methods often face the problem of insufficient accuracy when dealing with pedestrian intention prediction in complex traffic environments. Summary of the Invention
[0007] To solve the above technical problems, the present application provides a pedestrian intention prediction method for complex traffic scenarios. This prediction method effectively improves the accuracy of pedestrian intention prediction by involving local feature flow, global feature flow, cross-attention module, and multi-scale feature fusion module, has a wide range of application prospects, and solves problems such as insufficient fusion of local and global features, insufficient cross-flow information interaction, and insufficient capture of multi-scale features in the prior art.
[0008] To achieve the above object, the present application is implemented through the following technical solutions:
[0009] The present application is a pedestrian intention prediction method for complex traffic scenarios. This pedestrian intention prediction method is implemented through a pedestrian intention prediction model. The pedestrian intention prediction model includes a local flow module, a global flow module, a cross-attention module, and a feature fusion module. The local flow module extracts pedestrian bounding box data from the pedestrian crossing video to generate local flow features. There are N global flow modules in total, which extract and capture pedestrian bounding box data, pedestrian center points, and self-vehicle speed from the pedestrian crossing video as global flow features to realize the modeling of environmental relationships. The cross-attention module realizes the interaction between global process features and local flow features as cross-attention flow features. The feature fusion module fuses cross-attention flow features with local flow features and global flow features. Specifically, the pedestrian intention prediction method specifically includes the following steps:
[0010] Step 1: Extract pedestrian bounding box data, pedestrian center points, and self-vehicle speed from the pedestrian crossing video, use the pedestrian bounding box data as local flow features, use the pedestrian bounding box data, pedestrian center points, and self-vehicle speed as global flow features, input the local flow features into the local flow module, and input the global flow features into the global flow module;
[0011] Step 2: The local flow module extracts local flow features through a convolutional neural network and captures time series dependence through a long short-term memory network (LSTM);
[0012] Step 3: The global flow module extracts global flow features through a multi-head attention mechanism layer and an adaptive multi-scale convolutional block to obtain information at different scales;
[0013] Step 4: The local flow module and the global flow module perform feature interaction through the cross-attention module to generate cross-attention flow features;
[0014] Step 5: The local flow features, global flow features, and cross-attention flow features are fused through the feature fusion module;
[0015] Step 6: Map the fused features to the prediction result through a fully connected layer to complete the pedestrian intention prediction.
[0016] A further improvement of this application lies in that the local flow module in the pedestrian intention prediction model includes:
[0017] A local feature encoding layer, including a convolutional neural network, embeds the extracted pedestrian bounding box data into a high-dimensional space through one-dimensional convolution and ReLU activation function of the convolutional neural network, and uses residual connection to retain the pedestrian bounding box data to generate local flow features;
[0018] A self-attention mechanism layer performs self-attention calculation on the extracted pedestrian bounding box data to capture the global dependence between local flow features;
[0019] Two long short-term memory network time series modeling layers, including long short-term memory network (LSTM), are used to further extract the features of the pedestrian bounding box data, capture the time series dependence, and splice the output of the local feature encoding layer and the output of the self-attention mechanism layer to generate the feature representation of the final decoder;
[0020] There are N global flow modules in total, and the global flow module includes:
[0021] A feature mapping layer: maps the pedestrian bounding box data, pedestrian center point, and ego vehicle speed data to a high-dimensional space through an encoder to generate global flow features;
[0022] A position encoding layer: adds the position information of the pedestrian bounding box data, ego vehicle speed, and pedestrian center point to the input features;
[0023] A multi-head attention mechanism layer: calculates the pedestrian bounding box data, pedestrian center point, and ego vehicle speed data passing through the position encoding layer through the multi-head self-attention mechanism layer to capture the global dependence between the pedestrian bounding box data, pedestrian center point, and ego vehicle speed;
[0024] An adaptive multi-scale convolutional block: extracts the local flow features of the time series, and combines adaptive pooling and weighted fusion to generate unified features;
[0025] A feed-forward neural network layer: maps the global flow features through two fully connected layers and a non-linear activation function to enhance the non-linear expression ability of the global flow features and improve the feature extraction ability of the pedestrian intention prediction model;
[0026] The cross-attention module includes
[0027] A global flow attention fusion layer combines multi-layer attention representations through weighted summation to generate the attention output of the global flow, and inputs the generated attention output of the global flow into the cross-attention module for global flow feature and local flow feature interaction;
[0028] A local flow self-attention layer generates an attention output of local flow features through the local flow self-attention layer;
[0029] A cross-attention interaction layer interacts the attention output of global flow features and the attention output of local flow features through the cross-attention interaction layer to generate cross-attention flow features, fuse global flow features and local flow features, and enhance the expression ability of the pedestrian intention prediction model.
[0030] A feature fusion module is used to fuse cross-attention flow features with local flow features and global flow features.
[0031] Specifically, the feature fusion module extracts local flow features using multiple convolutional encoders and enhances the non-linear expression ability through the ReLU activation function; compresses the output of the convolutional kernels of each feature fusion module into a one-dimensional feature vector through global max pooling; calculates the selection weights of global flow, local flow, and cross-attention flow through a fully connected layer, normalizes the weight values using the softmax operation, and selects the features with the maximum weights; and generates a mask to retain the selected features according to the selected weights to achieve feature fusion.
[0032] A further improvement of this application is that in step 1, pedestrian bounding box data , the ego-vehicle speed and the pedestrian center point are extracted from the dataset, and then the pedestrian bounding box data, the ego-vehicle speed, and the pedestrian center point are concatenated together as the global flow features , and only the pedestrian bounding box data is used as the local flow features :
[0033]
[0034]
[0035] Among them, is the local flow feature, represents the global flow feature, is the pedestrian bounding box data, is the ego-vehicle speed, is the pedestrian center point, is the number of samples in a batch, that is, the number of samples input into the pedestrian intention prediction model at the same time, is the length of the time series, that is, the number of frames of data included in each sample.
[0036] A further improvement of this application is that step 2 specifically includes the following steps:
[0037] Step 2.1, local flow feature encoding: Through a series of one-dimensional convolutions Embed the local flow features into a high-dimensional space. The ReLU activation function introduces non-linearity, and the original information is retained through residual connections to improve the training efficiency and extract the local flow bounding box features , specifically:
[0038]
[0039] Among them, represents the input local flow features The local flow bounding box features obtained after being processed by the encoder, is the residual connection data, is the number of convolution times.
[0040] Step 2.2, Self-attention mechanism: Perform self-attention (SelfAttention) calculation on the local flow bounding box features extracted in Step 2.1 to obtain the local flow enhanced features :
[0041] ;
[0042] Step 2.3, Long short-term memory network time series modeling: The local flow enhanced features and the pedestrian motion features after converting the original pedestrian bounding box data are respectively subjected to time series modeling through different long short-term memory network time series modeling layers, and the outputs of the two long short-term memory networks are concatenated in the feature dimension to generate the final decoder features , where the method for extracting the pedestrian motion features includes the following steps:
[0043] Step 2.3.1. To extract the pedestrian motion features, the center coordinates of the pedestrian bounding box at each time step ( ), the width of the pedestrian bounding box and the height of the pedestrian bounding box are used as the position features. The center coordinates of the pedestrian bounding box are defined as follows:
[0044]
[0045]
[0046] Among them, is the upper left corner of the pedestrian bounding box at the th time step coordinate, is the lower right corner of the pedestrian bounding box at the th time step coordinate, is the The coordinates of the upper left corner of the bounding box for a time step are the coordinates of the lower right corner of the bounding box for the th time step;
[0047] The width of the pedestrian bounding box and the height of the pedestrian bounding box are defined as follows:
[0048]
[0049] Wherein, represents the width of the pedestrian bounding box for the th time step, represents the height of the pedestrian bounding box for the th time step;
[0050] Step 2.3.2. For each pair of adjacent time steps time step , the velocity feature is defined as the change in position and the change in size :
[0051]
[0052]
[0053]
[0054]
[0055] Wherein, is the horizontal coordinate of the center point of the current time step , is the horizontal coordinate of the center point of the previous time step , is the absolute difference between the horizontal coordinates of the center points of the current time step and the previous time step , i.e., the horizontal displacement amplitude, is the vertical coordinate of the center point of the current time step , is the vertical coordinate of the center point of the previous time step , is the absolute difference between the vertical coordinates of the center points of the current time step and the previous time step , i.e., the vertical displacement amplitude, is the width of the bounding box of the current time step , is the width of the bounding box of the previous time step The width of the bounding box, is the current time step and the previous time step The absolute difference between the widths of the bounding boxes is the width change amplitude, is the current time step The height of the bounding box, is the previous time step The height of the bounding box, is the current time step and the previous time step The absolute difference between the heights of the bounding boxes is the height change amplitude;
[0056] Step 2.3.3. The pedestrian motion feature at each time step is , and the final decoder feature output by the long short-term memory network for temporal modeling is:
[0057]
[0058] Among them, is the first temporal modeling layer, is the second temporal modeling layer.
[0059] A further improvement of this application is that: Step 3 specifically includes the following steps:
[0060] Step 3.1. Global feature mapping: Map the pedestrian bounding box data, pedestrian center point, and ego vehicle speed data to a high-dimensional space through an encoder composed of a fully connected layer:
[0061]
[0062] Among them, is the global flow feature, represents the global feature of the bounding box obtained after being processed by the encoder;
[0063] Step 3.2. Position encoding: The position encoding layer adds position information to the input features by introducing a learnable smooth position encoding matrix to enhance the model's perception ability of the element order in the sequence. Assume the position encoding is , then for each position , the smoothed encoding is expressed as:
[0064]
[0065] Among them, is the position encoding of position , is the position The position encoding, is the position encoding, is the position corresponding smoothing factor, is the position corresponding smoothing factor, is the position corresponding smoothing factor;
[0066] Step 3.3, Multi-Head Attention Mechanism: Calculate the global features of the bounding boxes with position encoding through the multi-head self-attention mechanism to obtain :
[0067]
[0068] Among them, represents the result after calculation by the multi-head attention mechanism layer in the th layer, is the number of layers;
[0069] Step 3.4, Adaptive Multi-Scale Convolution Block: Extract the local flow features of the time series through the convolution kernels of the adaptive multi-scale convolution block, and generate unified features by combining adaptive pooling and weighted fusion:
[0070]
[0071]
[0072] Among them, is the output of each convolution operation, is the convolution kernel of each convolution operation, is the bias term, is the unified feature extracted from the input data of the adaptive multi-scale convolution block under the processing of different convolution kernels;
[0073] Step 3.5, Feed-Forward Neural Network: After being processed by the adaptive multi-scale convolution block to obtain unified features, map the global flow features through two fully connected networks and a non-linear activation function:
[0074]
[0075] Among them, is the weight matrix of the intermediate layer of the feed-forward neural network between the input layer and the output layer of the feed-forward neural network, is the weight matrix from the intermediate layer of the feed-forward neural network to the output layer of the feed-forward neural network, is the output feature of the feedforward network layer.
[0076] A further improvement of this application is that step 4 specifically includes the following steps:
[0077] Step 4.1: Through the global flow attention fusion layer, generate the attention output of the global flow, and input the generated attention output of the global flow into the cross-attention module for feature interaction. The specific formula is as follows:
[0078]
[0079] Among them, is the final multi-layer attention feature after weighted summation, is the weight of the layer, the layer after being calculated by the multi-head attention mechanism layer;
[0080] Step 4.2: Through the local flow self-attention layer, generate the local flow enhanced feature ;
[0081] Step 4.3: The cross-attention module uses the multi-head attention mechanism layer to combine the local flow enhanced feature and the multi-layer attention feature of the weighted sum of the entire flow , and generate the cross-attention flow feature, expressed as:
[0082]
[0083] Among them, is the final multi-layer attention feature after weighted summation, is the local flow enhanced feature, is the dimension of the key.
[0084] A further improvement of this application is that step 5 specifically includes the following steps:
[0085] Step 5.1: The feature fusion module uses multiple convolutional encoders to extract local flow features. The specific formula for each convolutional kernel to process is as follows:
[0086]
[0087] Among them, is a one-dimensional feature vector, is the input feature of each convolutional kernel.
[0088] Step 5.2: The process of feature selection is that each flow calculates the selection weight through the fully connected layer, and then normalizes the weight value through the softmax operation to select the feature with the maximum weight:
[0089]
[0090] Among them, is the fully connected layer of the th stream, are the local stream, the global stream, and the cross-attention stream;
[0091] According to the selected weights, top-K features are selected, and the selected features are retained using a mask. The specific formula for each stream is as follows:
[0092]
[0093] Among them, is the mask selected through the top-K index.
[0094] A further improvement of this application lies in that: Step 6 is specifically: mapping the fused features to the final prediction result through a fully connected layer, and using the cross-entropy loss function to complete the training of the pedestrian intention prediction model, specifically:
[0095]
[0096] Among them, represents the fused features, is the final prediction result.
[0097] The beneficial effects of this application are:
[0098] The two-stream architecture proposed in this application, namely the global stream and the local stream, can extract global and local features respectively and use the cross-attention module for feature fusion, which can capture complex patterns in the time series more comprehensively. Compared with traditional models that only rely on a single feature extraction method, this application realizes the complementarity of global and local features through the two-stream architecture, effectively reducing information loss or deviation and significantly improving the expression ability of the model.
[0099] The cross-attention module proposed in this application can dynamically combine the features of the Transformer stream and the LSTM stream, enhance the interaction between the two streams, and improve the effect of feature fusion. Compared with most models that only use the attention mechanism in a single stream, this application realizes the dynamic fusion of the features of the two streams through the cross-attention mechanism, and can better utilize the complementary information of the two streams.
[0100] The multi-scale feature fusion module proposed in this application can extract multi-scale features by combining different convolutional kernel sizes, thereby enhancing the expression ability for complex patterns. Compared with single-scale convolutional operations, multi-scale convolution can capture the diversity of data more comprehensively, significantly improving the generalization ability and robustness of the model.
[0101] By introducing learnable positional encoding and a smoothing mechanism, this application can more flexibly capture sequential information and local patterns in time series. Compared with fixed positional encoding such as sine / cosine encoding in traditional models, the design of this application can better adapt to and process complex time series data. BRIEF DESCRIPTION OF THE DRAWINGS
[0102] Figure 1 is a flowchart of this application.
[0103] Figure 2 is a schematic diagram of the model of this application.
[0104] Figure 3 is a comparison chart of PIE results between the method of this application and the prior art method.
[0105] Figure 4 is a comparison chart of JAAD results between the method of this application and the prior art method. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0106] The following will disclose the embodiments of the present invention with reference to the drawings. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit this application. That is to say, in some embodiments of this application, these practical details are not necessary.
[0107] As Figure 1 shown, this application is a pedestrian intention prediction method for complex traffic scenarios. The pedestrian intention prediction method is implemented through a pedestrian intention prediction model. The pedestrian intention prediction model includes a local flow module, a global flow module, a cross-attention module, and a feature fusion module. The local flow module extracts pedestrian bounding box data from a pedestrian crossing video to generate local flow features. There are N global flow modules in total, which extract and capture pedestrian bounding box data, pedestrian center points, and the speed of the host vehicle from the pedestrian crossing video to generate global flow features, realizing the modeling of environmental relationships. The cross-attention module realizes the interaction between the global flow features and the local flow features to generate cross-attention flow features. The feature fusion module fuses the cross-attention flow features with the local flow features and the global flow features.
[0108] As Figure 2 shown, the local flow module of this application includes:
[0109] A local feature encoding layer, including a convolutional neural network. The extracted pedestrian bounding box data is embedded into a high-dimensional space through one-dimensional convolution and ReLU activation function of the convolutional neural network, and the pedestrian bounding box data is retained using residual connections to generate local flow features; the local flow features are enhanced through the local flow module and the modeling of the temporal changes of pedestrians is realized.
[0110] A self-attention mechanism layer performs self-attention calculation on the extracted pedestrian bounding box data to enhance the representation of the pedestrian bounding box data and capture the global dependencies between local flow features.
[0111] Two long short-term memory network temporal modeling layers, including long short-term memory networks (LSTM), are used to further extract the features of the human bounding box data, capture the time series dependencies, and concatenate the outputs of the local feature encoding layer and the self-attention mechanism layer to generate the final feature representation for the decoder.
[0112] The global flow module extracts and captures the pedestrian bounding box data, the pedestrian center point, and the ego vehicle speed from the pedestrian crossing video, enhances the global flow features, and realizes the modeling of the environmental relationship. There are N such global flow modules. The global flow module includes:
[0113] A feature mapping layer: Maps the data of the pedestrian bounding box data, the pedestrian center point, and the ego vehicle speed to a high-dimensional space through an encoder composed of fully connected layers to generate global flow features.
[0114] A position encoding layer: Adds the position information of the pedestrian bounding box data, the ego vehicle speed, and the pedestrian center point to the input features by introducing a learnable smooth position encoding matrix, enhancing the pedestrian intention prediction model's perception ability of the element order in the sequence.
[0115] A multi-head attention mechanism layer: Calculates the pedestrian bounding box data, the pedestrian center point, and the ego vehicle speed data passing through the position encoding layer through a multi-head self-attention mechanism layer to capture the global dependencies between the pedestrian bounding box data, the pedestrian center point, and the ego vehicle speed.
[0116] An adaptive multi-scale convolution block: Extracts the local flow features of the time series through the convolution kernels in the adaptive multi-scale convolution block, and combines adaptive pooling and weighted fusion to generate unified features, enhancing the pedestrian intention prediction model's ability to capture different time scale patterns.
[0117] A feed-forward neural network layer: Maps the global flow features through two fully connected layers and a non-linear activation function to enhance the non-linear expression ability of the global flow features and improve the feature extraction ability of the pedestrian intention prediction model.
[0118] The interaction between the global flow features and the local flow features is realized through a cross-attention module to enhance the expression ability of the pedestrian intention prediction model and retain the key feature information. It includes:
[0119] A global flow attention fusion layer combines the multi-layer attention representations in a weighted summation manner to generate the attention output of the global flow, and inputs the generated attention output of the global flow into the cross-attention module for the interaction between the global flow features and the local flow features.
[0120] A local flow self-attention layer generates an attention output of local flow features through the local flow self-attention layer;
[0121] A cross-attention interaction layer interacts the attention output of global flow features and the attention output of local flow features through the cross-attention interaction layer to generate cross-attention flow features, fusing global flow features and local flow features and enhancing the expressive ability of the pedestrian intention prediction model.
[0122] A feature fusion module is used to fuse cross-attention flow features with local flow features and global flow features.
[0123] Specifically, the feature fusion module is as follows: Multiple convolutional encoders are used to extract local flow features, and the ReLU activation function is used to enhance the non-linear expression ability; The output of the convolution kernel of each feature fusion module is compressed into a one-dimensional feature vector through global max pooling operation; The selection weights of each flow, namely global flow, local flow, and cross-attention flow, are calculated through a fully connected layer, and the weight values are normalized using the softmax operation to select the features of each flow with the maximum weight; According to the selected weights, a mask is generated through top-K indexing to retain the selected features, realizing feature fusion.
[0124] Specifically, as Figure 1 shown, the pedestrian intention prediction method specifically includes the following steps:
[0125] Step 1: Extract pedestrian bounding box data, pedestrian center points, and ego vehicle speed from the pedestrian crossing video annotation file, and use the pedestrian bounding box data as local flow features. By fusing fine-grained and macroscopic information, the accuracy and robustness of model prediction are improved. The pedestrian bounding box data, pedestrian center points, and ego vehicle speed are used as global flow features. The local flow features are input into the local flow module, and the global flow features are input into the global flow module.
[0126] Extract pedestrian bounding box data from the dataset ego vehicle speed and pedestrian center points , then concatenate the pedestrian bounding box data, ego vehicle speed, and pedestrian center points together as global flow features , and only use the pedestrian bounding box data as local flow features , the global flow features reflect the relative position and motion state between the pedestrian and the ego vehicle, while the local flow features only focus on the position information of the pedestrian:
[0127]
[0128]
[0129] Among them, is the number of samples in a batch, that is, the number of samples input into the pedestrian intention prediction model at the same time. is the length of the time series, that is, how many frames of data each sample contains.
[0130] Step 2: The local flow module extracts local flow features through a convolutional neural network and captures time series dependencies through a long short-term memory network (LSTM). The specific steps are as follows:
[0131] Step 2.1: Local flow feature encoding: Through a series of one-dimensional convolutions embed the local flow features into a high-dimensional space, introduce non-linearity with the ReLU activation function, and retain the original information through residual connections to improve training efficiency and extract local flow bounding box features Specifically:
[0132]
[0133] Among them, represents the input local flow features and the local flow bounding box features obtained after being processed by the encoder, is the residual connection data, is the number of convolutions
[0134] Step 2.2: Self-attention mechanism: Perform self-attention (SelfAttention) calculation on the local flow bounding box features extracted in Step 2.1 to obtain local flow enhanced features :
[0135] ;
[0136] Step 2.3: Long short-term memory network time series modeling: Pass the local flow enhanced features and the pedestrian motion features after converting the original pedestrian bounding box data through different long short-term memory network time series modeling layers for time series modeling respectively, and concatenate the outputs of the two long short-term memory networks in the feature dimension to generate the final decoder features . Among them, the extraction method of the pedestrian motion features includes the following steps:
[0137] Step 2.3.1: To extract pedestrian motion features, the center coordinates of the pedestrian bounding box at each time step ( ), the width of the pedestrian bounding box and the height of the pedestrian bounding box are used as position features. The center coordinates of the pedestrian bounding box are defined as follows:
[0138]
[0139]
[0140] Among them, is the upper left coordinate of the pedestrian bounding box at the th time step, is the lower right coordinate of the pedestrian bounding box at the th time step, is the upper left coordinate of the bounding box at the th time step, is the lower right coordinate of the bounding box at the th time step;
[0141] The width of the pedestrian bounding box and the height of the pedestrian bounding box are defined as follows:
[0142]
[0143] Among them, represents the width of the pedestrian bounding box at the th time step, represents the height
[0144] Step 2.3.2. Calculate the differences between the positions and dimensions between adjacent time steps, and use the differences as velocity features. For each pair of adjacent time steps time step , the velocity feature is the difference in the position feature, and the velocity feature is defined as the position change and the dimension change :
[0145]
[0146]
[0147]
[0148]
[0149] Among them, is the horizontal coordinate of the center point at the current time step , is the horizontal coordinate of the center point at the previous time step , is the current time step and the previous time step The absolute difference in the horizontal coordinate of the center point is the horizontal displacement amplitude, is the current time step The vertical coordinate of the center point, is the previous time step The vertical coordinate of the center point, is the current time step and the previous time step The absolute difference in the vertical coordinates of the center points is the vertical displacement amplitude, is the current time step The width of the bounding box, is the previous time step The width of the bounding box, is the current time step and the previous time step The absolute difference in the widths of the bounding boxes is the width change amplitude, is the current time step The height of the bounding box, is the previous time step The height of the bounding box, is the current time step and the previous time step The absolute difference in the heights of the bounding boxes is the height change amplitude;
[0150] Step 2.3.3. The pedestrian motion feature at each time step is , and the final decoder feature output by the long short-term memory network for temporal modeling is:
[0151]
[0152] Among them, is the first temporal modeling layer, is the second temporal modeling layer.
[0153] Step 3. The global flow module extracts global flow features through the multi-head attention mechanism layer and the adaptive multi-scale convolution block to obtain information at different scales. The specific steps are as follows:
[0154] Step 3.1. Global feature mapping: Map the pedestrian bounding box data, pedestrian center point, and ego vehicle speed data to a high-dimensional space through an encoder composed of a fully connected layer:
[0155]
[0156] Among them, is the global flow feature, Represents the global bounding box features obtained after being processed by the encoder;
[0157] Step 3.2, Positional Encoding: The positional encoding layer adds positional information to the input features by introducing a learnable smooth positional encoding matrix to enhance the model's perception of the element order in the sequence. Assume the positional encoding is , then for each position , the smoothed encoding is expressed as:
[0158]
[0159] where, is the positional encoding of position , is the positional encoding of position , is the positional encoding of position , is the positional encoding of position , is the smoothing factor corresponding to position , is the smoothing factor corresponding to position ;
[0160] Step 3.3, Multi-Head Attention Mechanism: Calculate the globally encoded bounding box features that have undergone positional encoding through the multi-head attention mechanism (multi-head self-attention) to obtain :
[0161]
[0162] where, represents the result after being calculated by the multi-head attention mechanism layer in the th layer, is the number of layers;
[0163] Step 3.4, Adaptive Multi-Scale Convolutional Block: Extract the local flow features of the time series through the convolutional kernels of the adaptive multi-scale convolutional block, and combine adaptive pooling and weighted fusion to generate unified features: to enhance the pedestrian intention prediction model's ability to capture different time-scale patterns.
[0164]
[0165]
[0166] where, is the output of each convolutional operation, is the convolutional kernel of each convolutional operation, is the bias term, is the unified feature extracted from the input data of the adaptive multi-scale convolution block under the processing of different convolution kernels;
[0167] Step 3.5, Feed-forward neural network: After being processed by the adaptive multi-scale convolution block to obtain the unified feature, the global flow feature is mapped through two layers of fully connected networks and a non-linear activation function to enhance the non-linear expression ability of the global flow feature and improve the feature extraction ability of the pedestrian intention prediction model.
[0168]
[0169] Among them, is the weight matrix of the intermediate layer of the feed-forward neural network between the input layer and the output layer of the feed-forward neural network, is the weight matrix from the intermediate layer of the feed-forward neural network to the output layer of the feed-forward neural network, is the output feature of the feed-forward network layer.
[0170] Step 4, The local flow module and the global flow module perform feature interaction through the cross-attention module to generate cross-attention flow features. Specifically, it includes the following steps:
[0171] Step 4.1, Through the global flow attention fusion layer, generate the attention output of the global flow, and input the generated attention output of the global flow into the cross-attention module for feature interaction. The specific formula is as follows:
[0172]
[0173] Among them, is the final multi-layer attention feature after weighted summation, is the weight of the th layer, the th layer after being calculated by the multi-head attention mechanism layer;
[0174] Step 4.2, Through the local flow self-attention layer, generate the local flow enhanced feature ;
[0175] Step 4.3, The cross-attention module uses the multi-head attention mechanism layer to and the multi-layer attention feature after weighted summation of the entire flow , generate the cross-attention flow feature, expressed as:
[0176]
[0177] Among them, is the final multi-layer attention feature after weighted summation, is the local flow enhancement feature, is the dimension of the key.
[0178] Step 5: The local flow feature, the global flow feature, and the cross-attention flow feature are fused through a feature fusion module; specifically, it includes the following steps:
[0179] Step 5.1: The feature fusion module uses multiple convolutional encoders to extract the local flow feature, and uses the ReLU activation function to enhance the non-linear expression ability. The output of each convolutional kernel is compressed into a one-dimensional feature vector through the global max pooling operation. The specific formula processed by each convolutional kernel is as follows:
[0180]
[0181] where, is the one-dimensional feature vector, is the input feature of each convolutional kernel.
[0182] In the process of step 5.2 for feature selection, each flow calculates the selection weight through a fully connected layer, and then normalizes the weight value through the softmax operation to select the feature with the maximum weight:
[0183]
[0184] where, is the fully connected layer of the th flow, are the local flow, the global flow, and the cross-attention flow;
[0185] According to the selected weights, top-K features are selected. In the embodiment of the present application, the top 90% important features are selected, and the selected features are retained with a mask (mask). The specific formula for each flow is as follows:
[0186]
[0187] where, is the mask selected through the top-K index.
[0188] Step 6: Map the fused feature to the prediction result through a fully connected layer to complete the pedestrian intention prediction. Specifically: Map the fused feature to the final prediction result through a fully connected layer, and use the cross-entropy loss function to complete the training of the pedestrian intention prediction model. Specifically:
[0189]
[0190] where, represents the fused feature, is the final prediction result.
[0191] To better illustrate the technical effects of this application, the following experiments are disclosed in this application:
[0192] 1. Experimental setup
[0193] Datasets: The performance of this method is evaluated on two benchmark datasets PIE and JAAD that are widely used in the field of pedestrian intention prediction.
[0194] Evaluation metrics: Accuracy, AUC (Area Under the Curve), and F1-score are selected as the main evaluation metrics. Accuracy reflects the proportion of correctly predicted samples by the model and is applicable to the case where the class distribution of the dataset is balanced. AUC measures the classification ability of the model at different decision thresholds through the area under the curve, and the F1-score can comprehensively consider the performance of the model on positive and negative samples, especially applicable to tasks with class imbalance or different misclassification costs.
[0195] 2. Experimental results
[0196] This application is compared with the state-of-the-art methods in recent years in Table 1. The performance of these benchmark methods is obtained from their respective original papers. This application reports the accuracy, AUC, and F1 on two widely used public datasets PIE and JAAD. Table 1 shows the frame-level AUC (%) comparison on 2 benchmark datasets. The best performance on each dataset is shown in bold.
[0197] Table 1
[0198]
[0199] As can be seen from Table 1 and Figure 3 , Figure 4 it can be seen that compared with other methods, this application basically achieves the best performance on the two datasets, which proves the effectiveness of this application.
[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit the technical solutions. Those of ordinary skill in the art should understand that any modifications or equivalent replacements made to the technical solutions of the present invention without departing from the purpose and scope of the present technical solution should be covered within the scope of the claims of the present invention.
[0201] The above are only the embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A pedestrian intention prediction method for complex traffic scenes, characterized by: The pedestrian intention prediction method is implemented by a pedestrian intention prediction model, which includes a local flow module, a global flow module, a cross-attention module and a feature fusion module. The specific pedestrian intention prediction method includes the following steps: Step 1: extract pedestrian bounding box data, pedestrian center point and vehicle speed from the pedestrian crossing video, use the pedestrian bounding box data as local flow features, use the pedestrian bounding box data, pedestrian center point and vehicle speed as global flow features, input the local flow features into the local flow module, and input the global flow features into the global flow module; Step 2: The local flow module extracts local flow features through a convolutional neural network and captures time series dependencies through a long short-term memory network; Step 3: The global flow module extracts global flow features through a multi-head attention mechanism layer and an adaptive multi-scale convolution block to obtain information at different scales. Step 4: The local stream module and the global stream module interact with each other through the cross-attention module to generate cross-attention stream features; Step 5: local flow features, global flow features and cross-attention flow features are fused through the feature fusion module; Step 6: Map the features fused in step 5 to the prediction results through the fully connected layer to complete the pedestrian intention prediction, where: The local flow module includes: a local feature encoding layer, including a convolutional neural network, embedding the extracted pedestrian bounding box data into a high-dimensional space through a one-dimensional convolution and a ReLU activation function of the convolutional neural network, and retaining the pedestrian bounding box data using a residual connection to generate a local flow feature; A self-attention mechanism layer that performs self-attention calculations on the extracted pedestrian bounding box data to capture the global dependencies between local flow features; Two LSTM temporal modeling layers, including LSTM, are used to further extract the features of the person bounding box data, capture the temporal series dependencies, and concatenate the output of the local feature encoding layer with the output of the self-attention mechanism layer to generate the feature representation of the final decoder; There are N global flow modules in total, and the global flow modules include: A feature mapping layer: The encoder maps the pedestrian bounding box data, pedestrian center point and vehicle speed data to a high-dimensional space to generate global flow features; A position encoding layer: adds pedestrian bounding box data, vehicle speed, and pedestrian center point location information to the global stream attention module input features; A multi-head attention mechanism layer: The multi-head self-attention mechanism layer calculates the pedestrian bounding box data, pedestrian center point and vehicle speed data after the position encoding layer to capture the global dependency between pedestrian bounding box data, pedestrian center point and vehicle speed; An adaptive multi-scale convolutional block: extracts local flow features of the time series and combines adaptive pooling and weighted fusion to generate unified features; A feedforward neural network layer: maps the global flow features through two fully connected layers and nonlinear activation functions, enhances the nonlinear expression ability of the global flow features, and improves the feature extraction ability of the pedestrian intention prediction model; The cross-attention module includes A global stream attention fusion layer combines the multi-layer attention representations by weighted summation to generate a global stream attention output, and inputs the generated global stream attention output into the cross-attention module to interact the global stream features with the local stream features; A local stream self-attention layer, which generates the attention output of the local stream features through the local stream self-attention layer; A cross-attention interaction layer, through which the attention output of the global stream feature interacts with the attention output of the local stream feature to generate cross-attention stream features, fuse the global stream features and the local stream features, and improve the expression ability of the pedestrian intention prediction model; The feature fusion module is used to fuse the cross-attention stream features with the local stream features and the global stream features. The feature fusion module is specifically as follows: multiple convolutional encoders are used to extract local flow features, and the nonlinear expression ability is enhanced through the ReLU activation function; the output of the convolution kernel of each feature fusion module is compressed into a one-dimensional feature vector through the global maximum pooling operation; the selection weights of the global flow, local flow and cross-attention flow are calculated through the fully connected layer, and the weight values are normalized using the softmax operation to select the features with the largest weights. According to the selected weights, masks are generated through the top-K index to retain the selected features and realize feature fusion.
2. The method for predicting pedestrian intention in complex traffic scenarios according to claim 1, characterized in that: In step 1, the pedestrian bounding box data X is extracted from the dataset bbox ∈R B×T×4 , vehicle speed X speed ∈R B×T×1 and pedestrian center point X center ∈R B×T×2 , then the pedestrian bounding box data, the vehicle speed and the pedestrian center point are spliced together as the global flow feature X2, and only the pedestrian bounding box data is used as the local flow feature X1. The global flow feature reflects the relative position and motion state between the pedestrian and the vehicle, while the local flow feature only focuses on the position information of the pedestrian: X1=X bbox X2=X bbox +X speed +X center Among them, X1 is the local flow feature, X2 is the global flow feature, and X bbox is the pedestrian bounding box data, X speed is the vehicle speed, X center is the center point of the pedestrian, B is the number of samples in a batch, that is, the number of samples simultaneously input into the pedestrian intention prediction model, and T is the length of the time series, that is, how many frames of data each sample contains.
3. The method for predicting pedestrian intention in complex traffic scenarios according to claim 2 is characterized by: Step 2 specifically includes the following steps: Step 2.1, local flow feature encoding: The local flow feature X1 is embedded into the high-dimensional space through a series of one-dimensional convolution Conv1D, the ReLU activation function introduces nonlinearity, and the original information is retained through the residual connection to extract the local flow bounding box feature X′1, specifically: X′1=ReLU(Conv1D N (X1))+X shortcut Among them, X′1 represents the local flow bounding box feature obtained after the input local flow feature X1 is processed by the encoder, and X shortcut is the residual connection data, N is the number of convolutions; Step 2.2, self-attention mechanism: Perform self-attention calculation on the local flow bounding box feature X′1 extracted in step 2.1 to obtain the local flow enhanced feature R1: R1=SA(X′1); Step 2.3, Long Short-Term Memory Network Time Series Modeling: The local flow enhancement feature R1 is combined with the pedestrian motion feature X converted from the original pedestrian bounding box data pv Time series modeling is performed through different long short-term memory network time series modeling layers respectively, and the outputs of the two layers of long short-term memory networks are concatenated in the feature dimension to generate the final decoder feature D1.
4. The method for predicting pedestrian intention in complex traffic scenarios according to claim 3 is characterized by: In step 2.3, the pedestrian motion feature X pv The extraction method comprises the following steps: Step 2.3.1: Set the center coordinates (x c ,y c ), pedestrian bounding box width w and pedestrian bounding box height h are used as position features, and the center coordinates of the pedestrian bounding box are defined as follows: in, is the x coordinate of the upper left corner of the pedestrian bounding box at the tth time step, is the x-coordinate of the lower right corner of the pedestrian bounding box at the t-th time step, is the y coordinate of the upper left corner of the bounding box at the tth time step, is the y coordinate of the lower right corner of the bounding box at the tth time step; The pedestrian bounding box width w and pedestrian bounding box height h are defined as follows: Among them, w (t) represents the width of the pedestrian bounding box at the tth time step, h (t) represents the height of the pedestrian bounding box at the tth time step; Step 2.3.2: For each pair of adjacent time steps t and time step t-1, define the velocity feature as the position change Δx c , Δy c and dimensional changes Δw, Δh: Δw (t) =|in (t) -In (t-1) | Δh (t) =|h (t) -h (t-1) | in, is the horizontal coordinate of the center point at the current time step t, is the horizontal coordinate of the center point at the previous time step t-1, is the absolute difference between the horizontal coordinates of the center point at the current time step t and the previous time step t-1, that is, the lateral displacement amplitude. is the vertical coordinate of the center point at the current time step t, is the vertical coordinate of the center point at the previous time step t-1, is the absolute difference between the vertical coordinates of the center point at the current time step t and the previous time step t-1, i.e., the longitudinal displacement amplitude, w (t) is the bounding box width at the current time step t, w (t-1) is the bounding box width at the previous time step t-1, Δw (t) is the absolute difference between the width of the bounding box at the current time step t and the previous time step t-1, i.e., the width change, h (t) is the bounding box height at the current time step t, h (t-1) is the bounding box height at the previous time step t-1, Δh (t) is the absolute difference between the bounding box height at the current time step t and the previous time step t-1, i.e., the height change amplitude; Step 2.3.3: The pedestrian motion feature at each time step is The final decoder feature D1 output by the long short-term memory network temporal modeling layer for temporal modeling is: D1=LSTM1(R1)+LSTM2(X pv ) Among them, LSTM1 is the first time series modeling layer, and LSTM2 is the second time series modeling layer.
5. The method for predicting pedestrian intention in complex traffic scenarios according to claim 1, characterized in that: The step 3 specifically includes the following steps: Step 3.1, global feature mapping: The pedestrian bounding box data, pedestrian center point and vehicle speed data are mapped to a high-dimensional space through an encoder composed of a fully connected layer: X′2=Linear(X2) Among them, X2 is the global flow feature, and X′2 represents the global feature of the bounding box obtained after being processed by the encoder; Step 3.2, position coding: The position coding layer introduces a smooth position coding matrix, assuming that the position coding is P = {P1, P2, ..., P T }, then for each position i, the smoothed code P smooth (i) is expressed as: Among them, P i is the position code of position i, P i-1 is the position code of position i-1, P i+1 is the position code of position i+1, α i is the smoothing factor corresponding to position i, α i-1 is the smoothing factor corresponding to position i-1, α i+1 is the smoothing factor corresponding to position i+1; Step 3.3, multi-head attention mechanism: The position-encoded bounding box global feature X′2 is calculated through the multi-head attention mechanism in, It represents the result of the calculation of the lth layer after the multi-head attention mechanism layer, where l is the number of layers; Step 3.4, Adaptive multi-scale convolution block: The convolution kernel of the adaptive multi-scale convolution block is used to extract the local flow features of the time series, and combined with adaptive pooling and weighted fusion to generate unified features: F=[f1,f2,...,f N ] Among them, f i For the output of each convolution operation, W i For each convolution operation, b i is the bias term, and F is the unified feature extracted from the input data of the adaptive multi-scale convolution block after being processed by different convolution kernels; Step 3.5, feedforward neural network: After being processed by the adaptive multi-scale convolution block, unified features are obtained, and the global flow features are mapped through a two-layer fully connected network and a nonlinear activation function: D2 = W2·Dropout(ReLU(W1·F)) Among them, W1 is the weight matrix of the middle layer of the feedforward neural network between the input layer and the output layer of the feedforward neural network, W2 is the weight matrix from the middle layer of the feedforward neural network to the output layer of the feedforward neural network, and D2 is the output feature of the feedforward network layer.
6. The method for predicting pedestrian intention in complex traffic scenarios according to claim 1, characterized in that: Step 4 specifically includes the following steps: Step 4.1: Generate the attention output of the global stream through the global stream attention fusion layer, and input the generated global stream attention output into the cross-attention module for feature interaction. The specific formula is as follows: Among them, R2 is the final weighted sum of multi-layer attention features, w l is the weight of the lth layer, The result of layer l after calculation by the multi-head attention mechanism layer; Step 4.2: Generate local flow enhanced feature R1 through the local flow self-attention layer; Step 4.3, the cross-attention module generates the cross-attention flow feature by using the multi-head attention mechanism layer to add the weighted sum of the local flow enhancement feature R1 and the multi-layer attention feature R2 of all flows, which is expressed as: Among them, d k is the dimension of the key.
7. The method for predicting pedestrian intention in complex traffic scenarios according to claim 1, characterized in that: The step 5 specifically includes the following steps: Step 5.1: The feature fusion module uses multiple convolutional encoders to extract local flow features. The specific formula for each convolution kernel processing is as follows: F po =MaxPool.ReLU(Conv1D(X)) / Among them, F po is a one-dimensional feature vector, X is the input feature of each convolution kernel; Step 5.2: The feature selection process is as follows: each stream q is selected by calculating the weight through the fully connected layer, and then the weight value is normalized through the softmax operation to select the feature with the maximum weight: w q =Softmax(FC n (q)) Among them, FC n is the fully connected layer of the nth stream, q is the local stream, global stream and cross-attention stream; According to the selected weights, the top-K features are selected and the selected features are retained with masks. The specific formula for each stream q is as follows: q selected =q×mask q Among them, mask q is the mask selected by top-K indexes.
8. The method for predicting pedestrian intention in complex traffic scenarios according to claim 1, characterized in that: The step 6 is specifically as follows: mapping the fused features to the final prediction result through a fully connected layer, and using the cross entropy loss function to complete the training of the pedestrian intention prediction model, specifically as follows: Among them, X fused represents the fused features, is the final prediction result.
Citation Information
Patent Citations
Pedestrian trajectory prediction and intention estimation method based on double-flow LSTM
CN117152209A
Pedestrian intention prediction method based on Transform multi-modal fusion strategy
CN118781575A
Pedestrian crossing intention prediction method based on multi-self-attention mechanism fused with multi-source information
CN117765568A
Multi-feature fusion algorithm for pedestrian intention recognition
CN118608906A