A pedestrian crossing intention prediction method and system using fusion enhancement

Through the action prediction framework of the Transformer-GRU model, the visual and non-visual information are combined for feature fusion, which solves the problem of insufficient utilization of modal information in pedestrian crossing prediction and realizes efficient prediction of pedestrian crossing intention.

CN120524306BActive Publication Date: 2025-09-23NORTHWESTERN POLYTECHNICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511033757.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-09-23
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

In existing pedestrian crossing prediction technologies, single-modal data fails to fully consider the impact of traffic scene information, multi-modal data lacks effective feature fusion, resulting in poor prediction results, and recurrent neural networks are prone to overfitting in the time dimension.

Method used

The action prediction framework adopts the Transformer-GRU model, which fuses visual and non-visual information and exploits the complementarity of visual and non-visual information to perform long-term sequence modeling and flexible iterative decoding. Feature fusion methods include splicing, learnable weighted sum, multi-layer perceptron fusion, attention fusion, and modal cross-fusion.

Benefits of technology

The accuracy of pedestrian crossing intention prediction is significantly improved, the differences and complementarities between modalities are fully utilized, and the modeling ability and prediction effect of the model on pedestrian crossing intention are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524306B_ABST
    Figure CN120524306B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for predicting pedestrian crossing intention using fusion enhancement. The method includes: inputting visual information and non-visual information, extracting visual features from the visual information using a visual feature extractor; splicing non-visual information in the feature dimension to form non-visual features; integrating visual features and non-visual features to obtain fused features; adding position encoding to the fused features to obtain position features, inputting the position features into a Transformer model, which outputs encoded features; inputting the encoded features into a GRU model, which outputs predicted target features, and inputting the predicted target features into a classifier to obtain pedestrian crossing intention. The present invention fully utilizes the complementarity of visual and non-visual information to strengthen the interaction between modalities; and adopts a Transformer-GRU motion prediction framework in the encoding and decoding process of the fused features, performing long-term sequence modeling and flexible iterative decoding, resulting in significant prediction results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent traffic situation awareness, and specifically relates to a method and system for predicting pedestrian crossing intention using fusion enhancement. Background Art

[0002] Traffic control systems are a crucial component of traffic management, regulating and managing road traffic flows through signal control and traffic monitoring. While ensuring driving safety, these systems also provide a safe travel environment for pedestrians by optimizing signal timing and installing pedestrian crossing facilities.

[0003] While traffic control systems can ensure pedestrian safety, according to the World Health Organization, more than half of the people who die in traffic accidents each year are vulnerable road users. As the most exposed and vulnerable participants in road traffic, protecting pedestrian safety is urgent.

[0004] Currently, within the scope of traffic control systems, the key to improving pedestrian safety lies in avoiding conflicts between pedestrians and vehicles. However, how to effectively avoid conflicts between pedestrians and vehicles has become a pressing issue in urban traffic control. Existing solutions include raising awareness of traffic rules, optimizing road design, and improving traffic facilities. With the rapid development of advanced driving technology and autonomous driving technology, in order to further meet the safety of pedestrians guided by traffic control systems, vehicles on the road use on-board sensors such as vision, lidar, and millimeter-wave radar to perceive the road environment, other vehicles, pedestrians, and other targets around the vehicle, and then predict pedestrian behavior, allowing the driver to plan and make decisions in advance, effectively avoiding conflicts between pedestrians and vehicles.

[0005] Considering that pedestrian-vehicle conflicts in the context of traffic management often occur at crosswalks, accurately predicting whether pedestrians will cross the crosswalk in front of the vehicle can not only avoid pedestrian-vehicle conflicts, but also improve road driving efficiency, thereby reducing pedestrian delays and traffic congestion, and assisting driving vehicles in finding a balance between safety and efficiency.

[0006] In recent years, efforts to improve the accuracy of pedestrian crossing prediction have focused on using single-modality data, such as image frame sequences, single-frame static images, and human poses. Some studies have enhanced prediction by adding traffic scene data from different modalities (e.g., pedestrian detection boxes, grid images, and semantic segmentation information). Others have refined network architectures (e.g., various attention mechanisms, scene relationship modeling, and visual information extraction) to improve prediction performance.

[0007] In terms of predicting pedestrians crossing crosswalks, although the above studies have solved many problems and improved the prediction effect, the following problems still exist in the pedestrian crossing prediction task: (1) Using single-modal data to predict pedestrian crossings does not take into account the impact of other information such as traffic scenes on pedestrian crossing intentions; (2) Some studies add multimodal data to improve the model prediction effect, but do not consider the differences and complementarities between modalities, and lack effective and efficient feature fusion methods; (3) The input features of pedestrian crossing prediction are mainly in the form of time series, but existing studies directly use recurrent neural networks to extract information from various modal data and use a large number of linear layers in the time dimension, which easily leads to overfitting and poor prediction effect. Summary of the Invention

[0008] The purpose of the present invention is to solve the technical problems existing in the above-mentioned existing pedestrian crossing prediction technology, and to provide a pedestrian crossing intention prediction method and system using fusion enhancement, which fully utilizes the complementarity of visual information and non-visual information and strengthens the interaction between modalities; the action prediction framework of the Transformer-GRU model is used in the encoding and decoding process of the fusion features to perform long-term sequence modeling and flexible iterative decoding, significantly improving the prediction effect.

[0009] To achieve the above objectives, the technical solutions provided by the present invention are:

[0010] In one aspect, the present invention provides a method for predicting pedestrian crossing intention using fusion enhancement, comprising:

[0011] Step 1: input visual information and non-visual information, extract visual features from the visual information through a visual feature extractor; and splice the non-visual information in a feature dimension to form non-visual features;

[0012] Step 2: Integrate the visual feature sequence and the non-visual feature sequence through the visual-non-visual fusion enhancement method to obtain a fused feature sequence;

[0013] Step 3: Add position coding to the fusion feature sequence to obtain a position feature sequence, input the position feature sequence into the Transformer model, the Transformer model obtains an output feature sequence, and averages the output feature sequence to obtain a global coding feature;

[0014] Step 4: Input the global encoding features into the GRU model, the GRU model outputs the predicted target features, and the predicted target features are input into the classifier to obtain the pedestrian crossing intention.

[0015] Based on one aspect, in a preferred embodiment of the present invention, in step 1:

[0016] The visual information includes: a series of image frames, the visual features represent the spatial information of a series of image frames; wherein the pedestrian is included The original image frame sequence is extracted by a visual feature extractor to extract the visual features;

[0017] The expression for extracting the visual features is:

[0018]

[0019] Where, represents the visual features extracted by the visual feature extractor, represents the visual feature extractor, represents the image input to the visual feature extractor, where , Indicates the time corresponding to the last frame of the image frame sequence;

[0020] Based on one aspect, in a preferred embodiment of the present invention, in step 1:

[0021] The non-visual information includes: the self-vehicle speed information sequence corresponding to the image frame where the pedestrian is located , pedestrian bounding box coordinate information sequence , pedestrian position trajectory information sequence , the non-visual features represent a series of traffic object attribute information; wherein:

[0022] Ego vehicle speed information sequence Expressed as: , where represents the ego vehicle speed at different moments read from the vehicle’s ego vehicle system;

[0023] Pedestrian bounding box coordinate information sequence Expressed as: , where Represents the coordinate information of the pedestrian bounding box at different times, Indicates time Target pedestrians The bounding box coordinates of , Indicates time The horizontal coordinate of the upper left corner of the pedestrian bounding box, Indicates time The vertical coordinate of the upper left corner of the pedestrian bounding box, Indicates time The horizontal coordinate of the lower right corner of the pedestrian bounding box, Indicates time The vertical coordinate of the lower right corner of the pedestrian bounding box;

[0024] Pedestrian position trajectory information sequence Expressed as: , where Indicates the pedestrian location trajectory information at different times, , Indicates time The horizontal coordinate of the pedestrian center point is Indicates time The vertical coordinate of the pedestrian center point;

[0025] The non-visual information is spliced ​​according to the channel dimension and processed by one-dimensional convolution to obtain the non-visual features. The non-visual features are expressed as: , where Represents non-visual features, Represents a splicing operation, represents the dimension of splicing, represents the channel dimension, Represents one-dimensional convolution Parameters.

[0026] Based on one aspect, in a preferred embodiment of the present invention, the visual-non-visual fusion enhancement method in step 2 is selected from any one of splicing, learnable weighted sum, multi-layer perceptron fusion, attention fusion, and modality cross fusion; wherein:

[0027] Splicing: directly splicing the visual features and the non-visual features in the channel dimension;

[0028] A learnable weighted sum is performed to perform a weighted sum of the visual features and the non-visual features using a learnable weight parameter;

[0029] Multi-layer perceptron fusion: visual features and non-visual features are spliced ​​along the channel dimension and input into the multi-layer perceptron for fusion;

[0030] Attention fusion, using the attention mechanism to perform attention weighted sum on the visual features and the non-visual features;

[0031] Modality cross-fusion,a visual cross-fusion method is adopted to capture the inherent temporal correlation,and the cross-correlation between visual and non-visual features,respectively.

[0032] Based on one aspect, in a preferred embodiment of the present invention, the integration process in step 2 is expressed as:

[0033]

[0034] Where, represents the fusion feature, represents the visual-non-visual fusion enhancement method, Represents non-visual features, Represents visual features, represents the pedestrian's id; 、 、 The expressions are:

[0035]

[0036] Where, Indicates The fusion characteristics of the moment, express The feature matrix of dimension;

[0037]

[0038] Where, Indicates non-visual features of the moment; express The feature matrix of dimension;

[0039]

[0040] Where, Indicates the visual characteristics of the moment; express The feature matrix of dimension.

[0041] Based on one aspect, in a preferred embodiment of the present invention, in step 3, the position code is represented as:

[0042] ,

[0043] Where, represents the definition of position encoding, Indicates the sequence position, represents the dimension position of the position encoding vector, Indicates the dimension of the feature after fusion;

[0044] In step 3, the Transformer model includes Transformer blocks, each Transformer block includes a multi-head self-attention mechanism layer, a residual connection layer with layer normalization, and a feedforward neural network layer. The propagation process of each Transformer block is expressed as:

[0045] ,

[0046] Where, Indicates the The normalized output of the Transformer block after the multi-head self-attention mechanism layer and residual connection, Indicates the - the output of 1 Transformer block, Indicates the The output of the first Transformer block is normalized after the feedforward neural network FFN and residual connection. The output of the Transformer block Take the average value to obtain the global encoding feature .

[0047] Based on one aspect, in a preferred embodiment of the present invention, the step four is specifically as follows:

[0048] Globally encoded features Input the GRU model as the initial hidden state of the GRU model ;

[0049] For any prediction moment in the prediction time step T , the last time step in the fusion feature Characteristics of the moment , as the input of the GRU model;

[0050] GRU model according to the characteristics and the hidden state at the current time step , calculate the hidden state of the next time step , whose expression is ,in , where Indicates the moment of prediction;

[0051] In the GRU model, the number of iterations is set to be consistent with the prediction time step T, and all iterations are completed, and the hidden state of the last iteration is As the predicted target feature, it is output by the GRU model. is the hidden state at time T;

[0052] The predicted target features are input into the classifier to obtain a pedestrian crossing intention score, and whether the pedestrian will perform a crossing action at the predicted time node is determined based on the pedestrian crossing intention score, thereby obtaining the pedestrian crossing intention.

[0053] In another aspect, the present invention provides a pedestrian crossing intention prediction system using fusion enhancement, comprising:

[0054] A visual feature extraction module, wherein the visual feature extraction module extracts a visual feature sequence through a visual feature extractor;

[0055] A non-visual splicing module, wherein the non-visual splicing module splices non-visual information in a feature dimension to form a non-visual feature sequence;

[0056] a fusion enhancement module, wherein the fusion enhancement module integrates the visual feature sequence and the non-visual feature sequence to obtain a fused feature sequence;

[0057] An encoding module, wherein the encoding module adopts a Transformer model for encoding, inputs the position feature sequence after adding the position encoding to the fusion feature sequence into the Transformer model, the Transformer model obtains an output feature sequence, and averages the output feature sequence to obtain a global encoding feature;

[0058] A decoding module, wherein the decoding module uses a GRU model to capture the dependency between current observations and future actions, inputs the global encoding features into the GRU model, and outputs predicted target features.

[0059] The advantages of the present invention are:

[0060] 1. This invention uses a feature fusion method to efficiently integrate multimodal information between visual and non-visual modes, fully leveraging the differences between different modal information and interactively integrating these differences, which is very beneficial for learning the Transformer-GRU motion prediction framework. In addition, feature fusion helps the encoder better understand and utilize the correlation between temporal and spatial context, improving the Transformer-GRU motion prediction framework's ability to model pedestrian crossing intention data.

[0061] 2. This invention constructs an encoder branch based on the Transformer model, which uses a self-attention mechanism to capture long-term dependencies between different modalities and generate fused features. This effectively captures global contextual information, enhancing the Transformer model's ability to understand categorical features. Furthermore, this invention uses a decoder branch based on the GRU model, combining the advantages of parallel and autoregressive models to flexibly iterate decoding for different prediction times.

[0062] 3. The present invention fully utilizes the complementarity of visual and non-visual information to strengthen the interaction between modalities; the Transformer-GRU action prediction framework is used in the encoding and decoding process of the fused features to perform long-term sequence modeling and flexible iterative decoding, with significant prediction effects.

[0063] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned by practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments with reference to the accompanying drawings, in which:

[0065] Figure 1 : A flow chart of a method for predicting pedestrian crossing intention using fusion enhancement provided by the present invention;

[0066] Figure 2 : The overall framework diagram of the pedestrian crossing intention prediction method provided by the present invention;

[0067] Figure 3 : Schematic diagram of the visual-non-visual fusion enhancement method provided by the present invention;

[0068] Figure 4 : The present invention provides a functional block diagram of a pedestrian crossing intention prediction system using fusion enhancement. DETAILED DESCRIPTION

[0069] The following describes in detail embodiments of the present invention. The embodiments are illustrative and intended to explain the present invention, but are not to be construed as limiting the present invention.

[0070] See also Figure 1 and Figure 2 , an embodiment of the present invention provides a pedestrian crossing intention prediction method using fusion enhancement, comprising the following steps:

[0071] Step 1: Input visual information and non-visual information, extract visual features from the visual information through a visual feature extractor; and splice the non-visual information in the feature dimension to form non-visual features.

[0072] In the above step 1 of the embodiment of the present invention, the visual information includes: a series of image frames, and the visual features represent the spatial information of the series of image frames; wherein the pedestrian is included The original image frame sequence The visual features are extracted by the visual feature extractor; wherein, the expression of visual feature extraction is:

[0073]

[0074] Where, represents the visual features extracted by the visual feature extractor, where represents a set of real numbers, that is, the extracted visual features belong to the real number range, represents the channel dimension of visual features, represents the visual feature extractor, represents the image input to the visual feature extractor, where , Indicates the time corresponding to the last frame in the image frame sequence. Indicates that the dimensions of the original input image are [C, H, W], where C represents the number of channels, usually 3, H represents the height of the image, and W represents the width of the image.

[0075] In practical applications, it is preferred that the embodiment of the present invention resizes the image to 224×224, that is, both H and W are 224. After passing through the visual feature extractor, the dimensions of the image features are mapped to the visual features extracted by the visual feature extractor. ,in , represents the channel dimension of the visual features, and , the visual feature sequence is represented as , Represents visual features at different moments. The speed information of the self-vehicle, the coordinate information of the pedestrian bounding box, and the position trajectory information of the pedestrian in the visual scene are spliced ​​in the channel dimension, and a non-visual feature sequence is obtained through one-dimensional convolution.

[0076] ,in, The channel dimension representing non-visual features, Represent non-visual features at different moments.

[0077] In the above step 1 of the embodiment of the present invention, the non-visual information includes: the self-vehicle speed information sequence corresponding to the image frame where the pedestrian is located , pedestrian bounding box coordinate information sequence , pedestrian position trajectory information sequence , non-visual features represent a series of traffic object attribute information; among them:

[0078] Ego vehicle speed information sequence Expressed as: , where represents the speed of the ego vehicle at different times;

[0079] Pedestrian bounding box coordinate information sequence Expressed as: , where Represents the coordinate information of the pedestrian bounding box at different times, Indicates time Target pedestrians The bounding box coordinates of , Indicates time The horizontal coordinate of the upper left corner of the pedestrian bounding box, Indicates time The vertical coordinate of the upper left corner of the pedestrian bounding box, Indicates time The horizontal coordinate of the lower right corner of the pedestrian bounding box, Indicates time The vertical coordinate of the lower right corner of the pedestrian bounding box;

[0080] Pedestrian position trajectory information sequence Expressed as: , where Indicates the pedestrian location trajectory information at different times, , Indicates time The horizontal coordinate of the pedestrian center point is Indicates time The vertical coordinate of the pedestrian center point;

[0081] The non-visual information is spliced ​​according to the channel dimension and processed by one-dimensional convolution to obtain the non-visual features. The non-visual features are expressed as: , where Represents non-visual features, Represents a splicing operation, represents the dimension of splicing, represents the channel dimension, Represents one-dimensional convolution Parameters.

[0082] The embodiment of the present invention obtains the above series of image frames and uses the visual feature extractor After that, visual features are obtained. Non-visual features are obtained by splicing the speed information of the ego vehicle, the coordinate information of the pedestrian's bounding box, and the position and trajectory information of the pedestrian in the visual scene in the channel dimension, making full use of the complementarity between different modalities.

[0083] It should be noted that the visual feature extractor of the embodiment of the present invention , a series of powerful feature extraction networks such as TSN, I3D, irCSN152, ViViT, etc. can be used.

[0084] Step 2: Integrate the visual feature sequence and the non-visual feature sequence through the visual-non-visual fusion enhancement method to obtain a fused feature sequence. Figure 3 In the above step 2 of the embodiment of the present invention, the visual-non-visual fusion enhancement method selects any one of the fusion enhancement methods such as splicing, learnable weighted sum, multi-layer perceptron fusion, attention fusion, and modality cross fusion; wherein:

[0085] Splicing: Visual features and non-visual features are directly spliced ​​in the channel dimension; for example, assuming the shape of visual features is [batch_size, sequence, F_dim], and the shape of non-visual features is [batch_size, sequence, Z_dim]; when splicing in the feature dimension, splicing is performed along the second dimension, that is, along the feature dimension. The shape of the spliced ​​features is: [batch_size, sequence, F_dim+Z_dim].

[0086] Learnable weighted sum: The visual features and non-visual features are weighted summed by learnable weight parameters; when the dimensions of the visual features and non-visual features are inconsistent, a fully connected layer is used to map the non-visual features to the visual features. The fully connected layer is a neural network layer that maps the input features to the output features by performing a linear transformation on the input; when the dimensions of the visual features and non-visual features are inconsistent, a fully connected layer is used to map the non-visual features to the same dimension as the visual features. Specifically, for the input features and the weight matrix , fully connected layer The calculation formula is: , where Represents the bias vector. For the visual feature sequence and non-visual feature sequences , using a fully connected layer to transform non-visual features Dimensions Mapping to visual features Same dimensions , expressed as: mapping non-visual features ,in, represents the weight of the fully connected layer, , represents the bias of the fully connected layer, ; Multiply the mapped non-visual features and visual features by their respective weights and sum them to complete the fusion. In practical applications, learnable weights can be used to allow the model to learn the most suitable weights for different modalities during training. The specific process is: the fused features after splicing ,in: represents the learnable weight parameter one, Represents the learnable weight parameter two.

[0087] Multi-layer perceptron fusion: Visual features and non-visual features are spliced ​​along the channel dimension and input into the multi-layer perceptron for fusion. Based on the splicing of the two features, the embodiment of the present invention inputs the spliced ​​features into a multi-layer perceptron (MLP) for fusion. The multi-layer perceptron (MLP) is continuously optimized during the training process. Specifically, the visual feature sequence and non-visual feature sequences Splicing is performed on the feature dimension to obtain a spliced ​​feature sequence , , where the visual feature sequence , non-visual feature sequence , splicing feature sequence The multilayer perceptron (MLP) consists of a linear layer and an activation function. The linear layer transforms the two input dimensions Mapping to output dimensions The specific process is as follows:

[0088] Splicing feature sequences Input into the linear layer, apply activation function (such as ReLU) to introduce nonlinearity: fusion feature sequence ,in: represents the linear layer weight, Representing the bias vector, using the multi-layer perceptron (MLP) structure, the feature representation is continuously optimized during the training process, which can effectively fuse the feature information of different modalities and improve the prediction ability.

[0089] Attention fusion: Use the attention mechanism to perform attention weighting on visual features and non-visual features. In attention fusion, the embodiment of the present invention uses the mapped visual feature results as the query , and use the non-visual feature sequence as the key Sum Perform attention calculation. The detailed steps include:

[0090] Input visual feature sequence and non-visual feature sequences , mapping visual feature sequences to queries , non-visual feature sequences are mapped as keys Sum :Query ,key ,value . Use the scaled dot product attention mechanism to calculate the attention output , where express The transpose of represents the dot product, Represents the scaled dot product. By calculating the dot product , and scaled dot product results ,application The function gets the attention weight and adds the attention weight to the value Multiply and get the attention output .

[0091] Modal cross-fusion: A visual cross-fusion method is used to capture the inherent temporal correlation within visual features and non-visual features, as well as the cross-correlation between visual features and non-visual features. The detailed steps include:

[0092] Input visual feature sequence and non-visual feature sequences . For visual features Initialize a cross mark and with Splicing to get ; For non-visual features Initialize a and with Splicing to get . At the same time, initialize a shared cross mark .Will After passing through a multi-layer perceptron (MLP), ,Will After passing through a multi-layer perceptron (MLP), ;Will 、 and Splicing to obtain fusion markers , and then The encoding is fused through the multi-layer perceptron (MLP) to obtain the fusion mark, and the fusion mark is extracted from the fusion mark. ;Will Respectively and Splice and obtain the visual features after cross fusion and non-visual features ; then for and Each layer of multi-layer perceptron (MLP) is used to obtain enhanced visual features through residual connections. and non-visual features , and finally the visual features and non-visual features The two are added together to obtain the fusion feature.

[0093] In the above step 2 of the embodiment of the present invention, the integration process is expressed as follows:

[0094]

[0095] Where, represents the fusion feature sequence, represents the visual-non-visual fusion enhancement method, represents a sequence of visual features, Indicates the pedestrian's id, represents a sequence of non-visual features, Represents the dimension of the feature after fusion. 、 、 The expressions are:

[0096]

[0097] Where, Indicates the fusion characteristics of the moment;

[0098]

[0099] Where, Indicates non-visual features of the moment;

[0100]

[0101] Where, Indicates The visual characteristics of the moment.

[0102] The present invention uses a visual-non-visual fusion enhancement method to effectively and efficiently fuse visual and non-visual information to obtain spatial information between modalities. This method encodes spatial information from the feature dimension and reconstructs the spatiotemporal features, leveraging the complementarity between different modalities.

[0103] Step 3: The fusion feature sequence obtained in step 2 Add positional encoding , obtain the position feature sequence, input the position feature sequence into the Transformer model, the Transformer model obtains the output feature sequence, and takes the average of the output feature sequence to obtain the global encoding feature. Since the Transformer model is not sensitive to the sequence, and the fusion feature obtained in step 2 is ordered, the embodiment of the present invention first adds the position encoding feature to the fusion feature sequence. (Positional Encoding) obtains a position feature sequence. In order to process variable-length input sequences, the embodiment of the present invention adopts two-dimensional fixed position encoding.

[0104] In step 3 of the embodiment of the present invention, the position code Expressed as:

[0105] ,

[0106] Where, Represents the fusion feature sequence The sequence position of a time step in is an integer, starting from 0, and the maximum value is the sequence length minus 1. If the input time series length is 16, then The value range of is 0, 1, ..., 15, representing the “number of time steps”; represents the dimension index position in the position encoding vector, Represents the fusion feature sequence Dimensions, if The value of is 512, then The value range is 0, 1, 2, ..., 255, and the position encoding formula is based on and , corresponding to even and odd dimensions respectively. Position feature sequence after adding position encoding Expressed as: , where represents the fusion feature sequence, Indicates the added position code. By adding position coding Get the position feature sequence.

[0107] In the position coding of the embodiment of the present invention, The position of the token in the sequence (counting starts from 0, the maximum value is "the input time series length - 1"), is the dimension index of the position encoding vector (from 0 to -1), 10000 is a hyperparameter that can be set to any value. Through the above position encoding method, each position will get a unique encoding vector, and the position encoding is added to the fusion feature sequence It is then input into the Transformer model for subsequent encoding operations.

[0108] In step 3 of the embodiment of the present invention, the Transformer model includes Transformer blocks, each of which includes a multi-head self-attention mechanism layer, residual connections with layer normalization, and a feed-forward neural network. The multi-head self-attention mechanism introduces learnable parameters that are continuously optimized during training, allowing different heads to learn different patterns. The propagation process of each Transformer block is expressed as:

[0109] ,

[0110] Where, Indicates the Intermediate variables in a Transformer block, represents the output of the m-1th Transformer block, Indicates the The output of the Transformer block is The output of the Transformer block Take the average value to obtain the global encoding feature ;in, Represents the processing of the multi-head self-attention mechanism layer, represents the processing process of the feedforward neural network, Represents the residual connection processing process with layer normalization.

[0111] Since the Transformer model of the embodiment of the present invention does not change the dimension of the input feature sequence, the output and input sequences maintain the same shape. In order to obtain the aggregated features of the fused feature sequence, the embodiment of the present invention takes the average value of the Transformer model output as the global encoding feature. .

[0112] When encoding the fused multimodal features, the embodiments of the present invention utilize the powerful temporal modeling capabilities of the Transformer model as an encoder to fully understand the observed content and fully capture the dependency between current observations and future actions, capturing the feature differences of different targets starting from global information.

[0113] Step 4: Input the global encoding features obtained in step 3 into the GRU model. The GRU model outputs the predicted target features, which are then input into the classifier to obtain the pedestrian's crossing intention.

[0114] Step 4 of the embodiment of the present invention is specifically: Input the GRU model as the initial hidden state of the GRU model ; For each prediction moment in the prediction time step T, the last time step in the fusion feature Characteristics of the moment As the input of the GRU model; the GRU model will decode according to the features each time and the current forecast time The hidden state , calculate the next prediction time The hidden state , whose expression is ,in , where Indicates the moment of prediction; the number of iterations of the GRU model is consistent with the prediction time step T. After all iterations are completed, the hidden state of the last iteration is As the prediction target feature, it is output by the GRU model. for The hidden state at the moment; the predicted target features are input into the classifier to obtain the pedestrian crossing intention score, and the pedestrian crossing intention score is used to determine whether the pedestrian will cross at the predicted time node.

[0115] See also Figure 4 An embodiment of the present invention also provides a pedestrian crossing intention prediction system using fusion enhancement, comprising a visual feature extraction module, a non-visual splicing module, a fusion enhancement module, an encoding module, and a decoding module. The visual feature extraction module extracts a visual feature sequence using a visual feature extractor. Specifically, the visual feature extraction module inputs a sequence of image frames and obtains a visual feature sequence using the feature extractor. The non-visual splicing module splices non-visual information in the feature dimension to form non-visual features. Specifically, the non-visual splicing module splices the ego vehicle's speed information, the pedestrian's bounding box coordinate information, and the pedestrian's position trajectory information in the channel dimension to obtain a non-visual feature sequence. The fusion enhancement module integrates the visual feature sequence and the non-visual feature sequence to obtain a fused feature sequence. Specifically, the fusion enhancement module effectively and efficiently fuses the visual and non-visual information, fully leveraging the complementarity between different modalities to obtain spatial information between the modalities. The encoding module uses a Transformer model for encoding. The position feature sequence after adding the position encoding to the fused feature sequence is input into the Transformer model. The Transformer model generates an output feature sequence, which is then averaged to obtain a global encoding feature. Specifically, the encoding module leverages the powerful temporal modeling capabilities of the Transformer model as an encoder to fully understand the observed content and fully capture the dependencies between the current observation and future actions. The decoding module utilizes the GRU model to capture the dependencies between the current observation and future actions, inputs the global encoded features into the GRU model, and outputs the predicted target features. Specifically, the decoding module predicts the future based on the encoder output, using the GRU's flexible iteration function as a decoder to solve predictions at different times.

[0116] The embodiment of the present invention utilizes the flexible iterative prediction time of the GRU model to make predictions within the corresponding prediction time for different prediction times. Compared to the Transformer model, the embodiment of the present invention is based entirely on the GRU model of the self-attention mechanism. The GRU model is both competitive and easy to understand and implement when computing resources are limited or the sequence length is moderate. In addition, the GRU model in the embodiment of the present invention effectively controls the flow of information through a gating mechanism, enabling the GRU model to capture long-range dependencies.

[0117] In addition, the embodiment of the present invention uses the encoder based on the Transformer model to perform long-term sequence modeling, and the decoder based on the GRU model to perform flexible iterative decoding. It combines the advantages of parallel and autoregressive models and is tested on the public datasets JAAD and PIE. From the experimental results, it can be seen that the Transformer-GRU-based action prediction framework disclosed in the embodiment of the present invention is very effective, verifying the effectiveness of the Transformer-GRU action prediction framework in the embodiment of the present invention.

[0118] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present invention, and these modifications or replacements should all be included in the scope of protection of the present invention.

Claims

1. A pedestrian crossing intention prediction method using fusion enhancement, characterized in that: include: Step 1: input visual information and non-visual information, and extract visual features from the visual information through a visual feature extractor; splicing the non-visual information in the feature dimension to form a non-visual feature; Step 2: Integrate the visual feature sequence and the non-visual feature sequence through the visual-non-visual fusion enhancement method to obtain a fused feature sequence; Step 3: Add position coding to the fusion feature sequence to obtain a position feature sequence, input the position feature sequence into the Transformer model, the Transformer model obtains an output feature sequence, and averages the output feature sequence to obtain a global coding feature; Step 4: Input the global encoding feature into the GRU model, the GRU model outputs the predicted target feature, and the predicted target feature is input into the classifier to obtain the pedestrian crossing intention; Specifically: The visual-non-visual fusion enhancement method in step 2 is selected from any one of splicing, learnable weighted sum, multi-layer perceptron fusion, attention fusion, and modality cross fusion; wherein: Splicing: directly splicing the visual features and the non-visual features in the channel dimension; Learnable weighted sum, which uses learnable weight parameters to perform weighted summation of the visual features and the non-visual features; Multi-layer perceptron fusion, which concatenates the visual features and the non-visual features along the channel dimension and inputs them into the multi-layer perceptron for fusion; Attention fusion, which uses the attention mechanism to perform attention weighted summation of the visual features and the non-visual features; Modality cross-fusion, which uses visual cross-fusion methods to capture the inherent temporal correlation within visual features and non-visual features, as well as the cross-correlation between visual features and non-visual features; In step 3, the position code is expressed as: Where PE represents the definition of position encoding, pos represents the sequence position, i represents the dimension position of the position encoding vector, and d x Indicates the dimension of the feature after fusion; In step 3, the Transformer model includes m Transformer blocks, each of which includes a multi-head self-attention mechanism layer, a residual connection layer with layer normalization, and a feedforward neural network layer. The propagation process of each Transformer block is expressed as: X′ (m) =LN(MHSA(X (m-1) )+X (m-1) ),X (m) =LN(MHSA(X (m-1 ))+X (m-1) ) Where X′ (m) represents the normalized output of the m-th Transformer block after the multi-head self-attention mechanism layer and residual connection, X (m-1) represents the output of the m-1th Transformer block, X (m) Represents the normalized output of the mth Transformer block after the feedforward neural network FFN and residual connection, and the output X of the mth Transformer block (m) Take the average value to obtain the global encoding feature The step 4 is specifically as follows: Globally encoded features Input GRU model as the initial hidden state h of the GRU model o ; For any prediction time t in the prediction time step T, the last time step t in the fusion feature o Characteristics of the moment As input to the GRU model; The GRU model calculates the hidden state h at the next time step based on the features and the hidden state h at the current time step t , and the expression for calculating the hidden state h at the next time step is t+1 where t < T, and in the formula, t represents the predicted time ​ In the GRU model, the number of iterations is set to be consistent with the prediction time step T, and all iterations are completed, and the hidden state h of the last iteration is T As the predicted target feature, h is output by the GRU model. T is the hidden state at time T; The predicted target features are input into the classifier to obtain a pedestrian crossing intention score, and whether the pedestrian will perform a crossing action at the predicted time node is determined based on the pedestrian crossing intention score, thereby obtaining the pedestrian crossing intention.

2. The method for predicting pedestrian crossing intention using fusion enhancement according to claim 1, characterized in that: In the step 1: The visual information includes: a series of image frames, and the visual features represent spatial information of the series of image frames; wherein the visual features are extracted from the original image frame sequence containing pedestrian i by a visual feature extractor; The expression for extracting the visual features is: Where, f i t represents the visual features extracted by the visual feature extractor, represents the visual feature extractor, V i t represents the image input to the visual feature extractor, where t = 1, 2, ..., t o , t o Indicates the time corresponding to the last frame in the image frame sequence.

3. The method for predicting pedestrian crossing intention using fusion enhancement according to claim 2, characterized in that: In the step 1: The non-visual information includes: the self-vehicle speed information sequence S corresponding to the image frame where the pedestrian is located i , pedestrian bounding box coordinate information sequence B i , pedestrian position trajectory information sequence L i , the non-visual features represent a series of traffic object attribute information; wherein: Ego vehicle speed information sequence S i Expressed as: Where, represents the speed of the ego vehicle at different times; Pedestrian bounding box coordinate information sequence B i Expressed as: Where, Represents the coordinate information of the pedestrian bounding box at different times, represents the bounding box coordinates of the target pedestrian i at time t, Represents the horizontal coordinate of the upper left corner of the pedestrian bounding box at time t, Represents the ordinate of the upper left corner of the pedestrian bounding box at time t, Represents the horizontal coordinate of the lower right corner of the pedestrian bounding box at time t, Represents the ordinate of the lower right corner of the pedestrian bounding box at time t; Pedestrian position trajectory information sequence L i Expressed as: Where, Indicates the pedestrian location trajectory information at different times, represents the horizontal coordinate of the pedestrian center point at time t, represents the vertical coordinate of the pedestrian center point at time t; The non-visual information is spliced ​​according to the channel dimension and processed by one-dimensional convolution to obtain the non-visual feature, which is expressed as: Z i =Conv1d(concat(S i , B i , L i , dim=channel); W), where Z i Represents non-visual features, concat represents the concatenation operation, dim represents the concatenation dimension, channel represents the channel dimension, and W represents the parameters of the one-dimensional convolution Conv1d.

4. The method for predicting pedestrian crossing intention using fusion enhancement according to claim 1, characterized in that: The integration process in step 2 is expressed as: X i =θ(Z i ,F i ) Where, X i represents the fusion feature, θ represents the visual-non-visual fusion enhancement method, Z i represents non-visual features, F i represents visual features, i represents the pedestrian’s id; where X i , Z i 、F i The expressions are: Where, Indicates that at t o The fusion characteristics of the moment, Indicates t o ×d x The feature matrix of dimension; Where, Indicates that at t o non-visual features of the moment; Indicates t o ×d s The feature matrix of dimension; Where, Indicates that at t o the visual characteristics of the moment; Indicates t o ×d v The feature matrix of dimension.

5. A pedestrian crossing intention prediction system using fusion enhancement, based on the pedestrian crossing intention prediction method using fusion enhancement according to any one of claims 1 to 4, characterized in that: The pedestrian crossing intention prediction system includes: A visual feature extraction module, wherein the visual feature extraction module extracts a visual feature sequence through a visual feature extractor; a non-visual splicing module, which splices non-visual information in a feature dimension to form a non-visual feature sequence; and a fusion enhancement module, which integrates the visual feature sequence and the non-visual feature sequence to obtain a fused feature sequence. An encoding module, wherein the encoding module adopts a Transformer model for encoding, inputs the position feature sequence after adding the position encoding to the fusion feature sequence into the Transformer model, the Transformer model obtains an output feature sequence, and averages the output feature sequence to obtain a global encoding feature; A decoding module, wherein the decoding module uses a GRU model to capture the dependency between current observations and future actions, inputs the global encoding features into the GRU model, and outputs predicted target features.

Citation Information

Patent Citations

  • Pedestrian intention multi-task identification and trajectory prediction method under view angle of intelligent automobile

    CN114120439A

  • Pedestrian crossing intention recognition method based on multi-source information fusion

    CN117173663A