A Visual Spatiotemporal Pedestrian Crossing Intention Prediction Method for Autonomous Driving
By constructing a spatial pyramid and a time pyramid model, combined with a dynamic space-time attention mechanism, the problem of accurate identification of pedestrians' intentions in autonomous driving is solved, and the safety of autonomous driving and feature discrimination capabilities in complex environments are improved.
Patent Information
- Application Number
- CN202510525077.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The prior art is difficult to accurately identify pedestrians' intention to pass through in complex traffic environments, resulting in insufficient safety in interaction between autonomous vehicles and pedestrians.
A spatial pyramid model of pedestrian pose and scene and a time pyramid model of pedestrian motion trajectory and velocity is constructed. Combined with the dual-path attention module to process spatial and temporal features, the probability of pedestrian travel intention is output through a classification network that can be separated and convolutionally separated in depth.
It improves the feature discrimination ability of autonomous driving in complex scenarios and the accuracy of pedestrian travel intention prediction, and improves the generalization ability and prediction efficiency of the model.
Smart Images

Figure CN120047924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly relates to a visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving. Background Art
[0002] Autonomous vehicles can eliminate or reduce factors such as fatigue, distraction, and drunk driving in human driving, thus significantly improving traffic safety. However, autonomous vehicles have difficulties in the safe and smooth interaction with pedestrians crossing the street, which has become one of the important obstacles restricting the implementation of high-level autonomous vehicles on urban streets. Therefore, accurately and real-time predicting the future movement intention of pedestrians is the basic guarantee for solving the conflict between humans and vehicles and improving road driving safety.
[0003] The prediction of pedestrian crossing intention is realized based on pedestrian detection and tracking. Generally, scholars studying pedestrian crossing intention prediction default that pedestrians have been detected and tracked. Currently, the prediction of pedestrian crossing intention is generally achieved through the following methods, for example:
[0004] Based on video images, relevant information about pedestrian crossing intention is obtained by extracting MCHOG features, and an SVM classifier is combined to determine whether pedestrians on the street cross the road;
[0005] 18 human key point skeleton points are located through a pose estimation and positioning algorithm, and a 396-dimensional feature vector is calculated by calculating the mathematical relationship between the above points. The 396-dimensional feature vector obtained from the image is input into the SVM, and the pedestrian motion state is obtained through the classifier;
[0006] The pedestrian frame data is cut into two images to capture the head and leg information respectively, feature extraction is performed using AlexNet, and finally the classifier is used to predict the action intention of the pedestrian. This network proves that the combination accuracy based on CNN and machine learning classifiers far exceeds other traditional methods;
[0007] With the development of recurrent neural networks, some research scholars have proposed the ConvLSTM network. The ConvLSTM network takes images as input, preprocesses them with a pre-trained CNN, and inputs the extracted features into a convolutional LSTM. The last hidden state is input into a fully connected layer for prediction.
[0008] The above research methods are proposed based on spatio-temporal points, human geometric features, or motion information. The information source is single. Once the driving environment is complex or lacks information, the prediction effect is poor. Therefore, the current intention recognition algorithm is difficult to accurately recognize the pedestrian crossing intention by autonomous vehicles in a complex traffic environment. Summary of the Invention
[0009] To meet the need for accurate recognition of pedestrians' crossing intentions in autonomous vehicles, the present invention describes a visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving, which has the advantages of high computational efficiency and accurate intention recognition. The specific technical solutions are as follows:
[0010] A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving includes the following steps:
[0011] S1. Extract the time features of pedestrian movement in the video captured by the on-vehicle camera, input them into a bidirectional LSTM network, and output a time context encoding vector;
[0012] S2. Perform max pooling on the pedestrian's pose features and average pooling on the scene semantic features to generate a multi-scale spatial feature map, and construct a spatial pyramid model of the pedestrian pose and the scene;
[0013] S3. Perform downsampling on the time context encoding vector in the time dimension, extract motion features with different time granularities, generate a multi-granularity time encoding vector, and construct a time pyramid model of the pedestrian movement trajectory and speed;
[0014] S4. Weight each level feature of the spatial pyramid model and the time pyramid model respectively; splice the weighted spatial features and time features to form a spliced spatio-temporal feature vector; sum the spliced feature vectors by level and output a fused spatio-temporal feature vector;
[0015] S5. Establish a classification network of depthwise separable convolution, input the fused spatio-temporal feature vector, and output the pedestrian crossing intention probability value P∈[0,1].
[0016] Furthermore, the time features include long-term slow-varying features and / or short-term sudden-change features. The long-term slow-varying features include the instantaneous speed, acceleration, and motion direction angle of the pedestrian during movement, and the short-term sudden-change features include sudden turning and sudden acceleration of the pedestrian during movement.
[0017] Furthermore, the time features are extracted by the following method:
[0018] S101. Set a sliding time window with a length of 5 frames;
[0019] S102. Extract continuous frame data with a quantity of ≧20 frames from the video captured by the on-vehicle camera;
[0020] S103. Process each frame image in the continuous frame data and real-time locate the pedestrian bounding box;
[0021] S104. Perform Kalman filter tracking on the center point of the pedestrian bounding box to generate a smooth pedestrian movement trajectory sequence;
[0022] S105. Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, acceleration, and motion direction angle of the pedestrian's motion trajectory, and establish long-term slowly varying features.
[0023] S106. Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, motion direction angle of the pedestrian's motion trajectory, as well as the angle change and speed change between adjacent trajectory points, and use the captured feature changes as short-term mutation features for subsequent pedestrian crossing intention prediction; set an angle change threshold according to the normal walking angle of the pedestrian, and if the normal walking angle change of the pedestrian exceeds the set threshold, it is determined that the pedestrian suddenly turns; set an angle change threshold according to the normal walking speed of the pedestrian, and if the normal walking speed change of the pedestrian exceeds the set threshold, it is determined that the pedestrian suddenly accelerates.
[0024] Furthermore, the pose features of the pedestrian include the coordinates of the pedestrian's key points, the limb angle vector, and the head orientation angle vector; the scene semantic features include at least one of zebra markings, sidewalk markings, motor vehicle lane markings, and traffic light states.
[0025] Furthermore, perform adaptive cropping and normalization processing on the pedestrian area within the bounding box through the YOLOv5s-Tiny model to generate a sequence of ROI images with a fixed size; perform semantic segmentation on the ROI image sequence through the MobileNetV3 Small network to extract the semantic features of the pedestrian; extract the coordinates of the pedestrian's key points through the Lite-HRNet network.
[0026] Furthermore, establish a spatial attention model and a temporal attention model to process spatial features and temporal features respectively, and output spatial attention weights and temporal attention weights; add spatial weights to different levels of the spatial pyramid model and add temporal weights to different levels of the temporal pyramid model, splice the spatial features and temporal features, and output a fused and enhanced spatio-temporal feature vector.
[0027] Furthermore, the input of the spatial attention model is a multi-scale spatial feature map, and the output is the spatial attention weight. The formula of the spatial attention model is:
[0028]
[0029] In the formula, is the spatial attention weight; is the sigmoid function; and are learning parameters; is the global average pooling; is the spatial feature;
[0030] The input of the temporal attention model is a multi-granularity temporal encoding vector, and the output is the temporal attention weight. The formula of the temporal attention model is:
[0031]
[0032] Wherein, is the time attention weight; is the sigmoid function; is the learning parameter; is the historical hidden state; is the current frame feature.
[0033] Furthermore, the concatenation of spatial and temporal features includes the following steps:
[0034] The weighted features of each level of the spatial pyramid are concatenated together along the channel dimension to form a fused spatial feature vector, and the expression is:
[0035]
[0036] Wherein, is the fused spatial feature vector; is the spatial attention weight of the nth level; is the spatial feature of the nth level; n is the number of levels of the spatial pyramid;
[0037] The weighted features of each level of the temporal pyramid are concatenated together along the channel dimension to form a fused temporal feature vector, and the expression is:
[0038]
[0039] Wherein, is the fused temporal feature vector; is the temporal attention weight of the mth level; is the temporal feature of the mth level; m is the number of levels of the temporal pyramid;
[0040] The fused spatial feature vector and the fused temporal feature vector are concatenated together along the channel dimension to form a fused spatio-temporal feature vector, and the expression is:
[0041]
[0042] Wherein, is the fused spatio-temporal feature vector; is the fused spatial feature vector; is the fused temporal feature vector.
[0043] Furthermore, establishing a classification network of depthwise separable convolution includes the following steps:
[0044] a. Depth convolution: Independently apply depth convolution operations to each channel of the input data. A separate 3×3 convolutional kernel is used for convolution on each input channel. The number of output channels generated by depth convolution is the same as the number of input channels, and each channel contains the convolution result of that channel;
[0045] b. Pointwise convolution: Apply pointwise convolution operations to the feature maps generated by depth convolution. The pointwise convolution uses a 1×1 convolutional kernel and performs convolution on all channels at each position, linearly combining each pixel point of the feature map to mix the feature of each channel generated by depth convolution;
[0046] c. After pointwise convolution, apply the ReLU function as the non - linear activation function;
[0047] d. Perform average pooling.
[0048] Furthermore, set a sliding time window with a length of 5 frames, perform weighted average of the prediction probabilities of 5 consecutive frames using the sliding window, and calculate the confidence through the following formula:
[0049]
[0050] In the formula, is the confidence; is the sliding window size; is the predicted probability value of the i - th frame within the sliding window; is the average probability within the window.
[0051] Based on the above technical solutions, the present invention has the following beneficial effects:
[0052] 1. By jointly modeling the spatial pyramid model of pedestrian pose and scene and the temporal pyramid model of pedestrian motion trajectory and speed, the problem of spatio - temporal feature fragmentation in the prior art is solved.
[0053] 2. Establish a dynamic spatio - temporal attention mechanism. The spatial and temporal features are processed separately through a dual - path attention module to output spatial attention weights and temporal attention weights, and a fusion strategy of spatial - temporal attention weights is adopted, which improves the feature discrimination ability of autonomous driving in complex scenarios.
[0054] 3. Through the spatio - temporal context fusion model and the dynamic spatio - temporal attention mechanism, compared with the traditional LSTM baseline model, the prediction accuracy of the pedestrian crossing intention prediction method recorded in the present invention has been significantly improved. For example, the mean average precision (mAP) of the lightweight YOLOv5s - Tiny model in object detection has increased by 2.26%, which is comparable to the YOLOv8s model, but the number of parameters is only 53% of the YOLOv8s model, as shown in the following table:
[0055] Parameter comparison data table of the lightweight YOLOv5s-Tiny model and the YOLOv8s model
[0056] Model Mean Average Precision (mAP) Number of Parameters (M) FLOPs (B) Yolov5s-Tiny Improved by 2.26% compared to the baseline model 4.2 8.7 Yolov8s 44.9 11.2 28.6 Description of the Drawings
[0057] Figure 1 : Schematic diagram of the overall process of the method described in the present invention;
[0058] Figure 2 : Schematic diagram of the structure of the spatio-temporal feature pyramid and the attention mechanism. Detailed Implementation Modes
[0059] It should be noted that:
[0060] 1. Certain terms are used in the description and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. The description and claims of this specification do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction. Unless otherwise defined, the technical terms or scientific terms used in this disclosure should have the ordinary meaning understood by those of ordinary skill in the art within the field to which this disclosure belongs.
[0061] 2. In this embodiment, a monocular camera is used for video and / or image shooting. On-vehicle monocular cameras have been widely used in the field of autonomous driving technology. Through a monocular camera, both images can be taken and continuous video streams of the driving environment can be collected. Each static image of the taken images and / or videos is defined as an original image. This embodiment does not make improvements to the structure, principle, operation mode, etc. of the monocular camera, and will not be elaborated further.
[0062] As shown in Attach Figure 1 and Attach Figure 2 This embodiment describes a visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving, including the following steps.
[0063] S1. Extract the temporal features of pedestrian movement from the video captured by the on-vehicle monocular camera, and input the temporal features into a bidirectional LSTM network to output a temporal context encoding vector.
[0064] The temporal features include long-term slow-changing features and / or short-term sudden-changing features. The long-term slow-changing features refer to the instantaneous speed, acceleration, and movement direction angle of the pedestrian during movement, and the short-term sudden-changing features refer to the sudden turning and sudden acceleration of the pedestrian during movement.
[0065] The extraction of the temporal features of pedestrian movement preferably adopts the following method or steps, specifically:
[0066] S101. Set a sliding time window
[0067] Set a sliding time window with a length of 5 frames to capture the short-term dynamic features of pedestrian movement.
[0068] S102, Extract consecutive frame data
[0069] Extract consecutive frame data from the video captured by the on-vehicle monocular camera, and the number of consecutive frame data ≥ 20 frames.
[0070] S103, Locate the pedestrian bounding box
[0071] Use the lightweight object detection model YOLOv5s-Tiny to process each frame in the consecutive frame data and locate the pedestrian bounding box in real time.
[0072] S104, Kalman filter tracking
[0073] Perform Kalman filter tracking on the center point of the pedestrian bounding box to generate a smooth sequence of pedestrian movement trajectories.
[0074] S105, Extract the long-term slow-varying features of pedestrian movement
[0075] Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous velocity, acceleration and movement direction angle of the pedestrian movement trajectory to establish long-term slow-varying features.
[0076] In this step, the calculation methods of "instantaneous velocity, acceleration and movement direction angle" adopt the methods commonly used in the field, that is, the instantaneous velocity and acceleration are obtained by calculating the derivative after smoothing the trajectory through Kalman filter, and then taking the vehicle coordinate system as the reference coordinate system, and the movement direction angle is calculated through the geometric relationship of the trajectory. Briefly described as follows: Instantaneous velocity, after smoothing the trajectory through Kalman filter, calculate the displacement difference between adjacent frames divided by the time difference; Acceleration, by calculating the rate of change of velocity; Movement direction angle, obtained by calculating the direction vector of the pedestrian movement trajectory.
[0077] S106, Extract the short-term sudden change features of pedestrian movement
[0078] S106-1, Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous velocity, movement direction angle of the pedestrian movement trajectory, as well as the angle change and velocity change between adjacent trajectory points, and use the captured feature changes as short-term sudden change features for subsequent prediction of pedestrian crossing intention;
[0079] S106-2, Set an angle change threshold according to the normal walking angle of the pedestrian. If the angle change exceeds the set threshold, it is determined as a sudden turn;
[0080] S106-3, Set a speed change threshold according to the normal walking speed of the pedestrian. If the speed change exceeds the set threshold, it is determined as a sudden acceleration.
[0081] S2, perform max pooling on the pedestrian's pose features and average pooling on the scene semantic features to generate a multi-scale spatial feature map, and construct a spatial pyramid model of the pedestrian pose and the scene.
[0082] Among them:
[0083] ① The pedestrian's pose features include the coordinates of 17 pedestrian key points, limb angle vectors, and head orientation angle vectors.
[0084] The coordinates of 17 pedestrian key points refer to the head, shoulders, elbows, hips, knees, ankles, etc. of the pedestrian, and the head includes at least the left ear, right ear, and nose tip;
[0085] The limb angle vector is calculated based on the included angle of the elbow-shoulder-hip three points and is used to judge the movement trend of the pedestrian's upper limb;
[0086] The head orientation angle vector is calculated based on the geometric relationship of the left ear, right ear, and nose tip and is used to judge the pedestrian's line of sight direction.
[0087] ② The scene semantic features include but are not limited to zebra crossing signs, sidewalk signs, motor vehicle lane signs, and traffic light states.
[0088] S3, perform downsampling on the time context encoding vector in the time dimension, extract motion features of different time granularities, generate a multi-granularity time encoding vector, and construct a time pyramid model of the pedestrian's motion trajectory and speed.
[0089] In this step:
[0090] ① Preferably, perform adaptive cropping and normalization processing on the pedestrian area within the bounding box through the lightweight object detection model YOLOv5s-Tiny to generate a sequence of ROI images with a fixed size;
[0091] ② Preferably, use the lightweight MobileNetV3 Small network to perform semantic segmentation on the ROI image sequence to accurately extract the semantic features of the pedestrian;
[0092] ③ Preferably, use the Lite-HRNet network to extract the coordinates of the pedestrian's key points.
[0093] S4, construct a spatio-temporal context fusion model, including the following steps:
[0094] S401, perform weighting on the features of each layer of the spatial pyramid model and the time pyramid model respectively. For example, perform weighting on the pedestrian pose features and scene semantic features of different scales, and the weights are dynamically adjusted according to the fully connected layer and the convolutional layer for learning;
[0095] S402. Concatenate the weighted spatial features and temporal features to form a concatenated feature vector.
[0096] S403. Sum the concatenated feature vector hierarchically to obtain the final spatio-temporal context fusion feature.
[0097] Through the joint modeling of the spatial pyramid model and the temporal pyramid model, the problem of spatio-temporal feature fragmentation in the existing technology can be solved, the prediction accuracy of the model can be improved, the feature discrimination ability can be enhanced, and the generalization ability of the model can be promoted.
[0098] S5. To enhance the feature discrimination ability of autonomous driving in complex scenarios, this step describes the dynamic spatio-temporal attention mechanism. The spatial features and temporal features are processed by the dual-path attention module respectively to output the spatial attention weight and the temporal attention weight, which specifically includes the following steps:
[0099] S501. Establish a spatial attention model, input the multi-scale spatial feature map into the spatial attention model, and output the spatial attention weight. The spatial attention model is used to calculate the importance weight of each channel, and the formula is:
[0100]
[0101] In the formula,
[0102] is the spatial attention weight;
[0103] is the sigmoid function;
[0104] and are learning parameters;
[0105] is the global average pooling;
[0106] is the spatial feature.
[0107] S502. Establish a temporal attention model, input the multi-granularity temporal encoding vector into the temporal attention model, and output the temporal attention weight. The temporal attention model calculates the temporal weight through a gating mechanism (GRU unit), and the formula is:
[0108]
[0109] In the formula,
[0110] is the temporal attention weight;
[0111] is the sigmoid function;
[0112] is a learning parameter;
[0113] is a historical hidden state;
[0114] is the current frame feature.
[0115] S6. Add spatial weights to different levels of the spatial pyramid model and add temporal weights to different levels of the temporal pyramid model to enhance the attention of the spatio-temporal context fusion model, and then splice the spatial and temporal features to output an enhanced spatio-temporal feature vector.
[0116] Among them, splicing the spatial and temporal features includes the following steps:
[0117] S601. Spatial feature splicing
[0118] Splice the features of each level of the weighted spatial pyramid together along the channel dimension to form a fused spatial feature vector , and the expression is:
[0119]
[0120] In the formula,
[0121] is the fused spatial feature vector;
[0122] is the spatial attention weight of the nth level;
[0123] is the spatial feature of the nth level;
[0124] n is the number of levels of the spatial pyramid.
[0125] S602. Temporal feature splicing
[0126] Splice the features of each level of the weighted temporal pyramid together along the channel dimension to form a fused temporal feature vector , and the expression is:
[0127]
[0128] In the formula,
[0129] is the fused temporal feature vector;
[0130] is the time attention weight of the m-th layer;
[0131] is the time feature of the m-th layer;
[0132] m is the number of layers of the time pyramid.
[0133] S603, Spatial-Temporal Feature Concatenation
[0134] Concatenate the fused spatial feature vector and the time feature vector along the channel dimension to form the final fused and enhanced spatio-temporal feature vector , and the expression is:
[0135]
[0136] In the formula,
[0137] is the fused spatio-temporal feature vector;
[0138] is the fused spatial feature vector;
[0139] is the fused time feature vector.
[0140] S7, Build a lightweight classification network with depthwise separable convolutions, input the fused spatio-temporal feature vector , and output the pedestrian crossing intention probability value P ∈ [0, 1].
[0141] Among them, building a lightweight classification network with depthwise separable convolutions includes the following steps:
[0142] S701, Depthwise Convolution
[0143] Apply depthwise convolution operations independently to each channel of the input data. Each input channel is convolved using a separate 3×3 convolution kernel. The number of output channels generated by depthwise convolution is the same as the number of input channels, and each channel contains the convolution result of that channel.
[0144] S702, Pointwise Convolution
[0145] Apply pointwise convolution operations to the feature map generated by depthwise convolution. Pointwise convolution uses a 1×1 convolution kernel and performs convolution operations on all channels at each position. This operation performs a linear combination of each pixel point of the feature map and mixes the channel features generated by depthwise convolution. The number of output channels of pointwise convolution can be flexibly specified and is usually used to control the depth of the output feature map.
[0146] S703, Nonlinear Activation Function
[0147] After pointwise convolution, the ReLU function, a non-linear activation function, is applied.
[0148] S704, Pooling
[0149] Average pooling is performed to reduce the spatial dimension of the feature map while retaining important features.
[0150] S8, Perform a sliding window weighted average on the predicted probabilities of 5 consecutive frames to suppress transient noise, and calculate the confidence through the following formula:
[0151]
[0152] In the formula,
[0153] is the confidence;
[0154] is the sliding window size;
[0155] is the predicted probability value of the i-th frame within the sliding window;
[0156] is the average probability within the window.
[0157] The real-time prediction method for pedestrian crossing intention described in this embodiment can also improve the inference speed of the model by performing structured pruning, lightweight deployment, and priority scheduling on the model. For example:
[0158] 1. Remove redundant convolutional channels in MobileNetV3 (channel pruning rate 30%), and retain high-activation feature layers.
[0159] 2. Use TensorRT to convert the FP32 model to INT8 precision. After testing, the model size is compressed to 2.3MB, and the inference speed is increased by 3 times.
[0160] 3. Dynamically allocate computing resources according to the pedestrian risk level and distance:
[0161] High-risk pedestrians: Allocate 90% of the GPU computing power, prediction frequency 30Hz;
[0162] Medium- and low-risk pedestrians: Allocate 10% of the GPU computing power, prediction frequency 10Hz;
[0163] Adopt thread pool management to ensure that the single-frame processing delay ≤ 30ms.
[0164] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements fall within the scope of the present invention claimed.
Claims
1. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving, characterized in that, It includes the following steps: S1. Extract the time features of the pedestrian movement in the video captured by the in-vehicle camera, input them into a bidirectional LSTM network, and output a time context encoding vector; S2. Perform max pooling on the pedestrian pose features and average pooling on the scene semantic features to generate a multi-scale spatial feature map, and construct a spatial pyramid model of the pedestrian pose and the scene; S3. Downsample the time context encoding vector in the time dimension, extract the motion features of different time granularities, generate a multi-granularity time encoding vector, and construct a time pyramid model of the pedestrian motion trajectory and speed; S4. Weight each level feature of the spatial pyramid model and the time pyramid model respectively; splice the weighted spatial features and time features. The method is as follows: the features of each level of the weighted spatial pyramid and the time pyramid are spliced together according to the channel dimension to form a fused spatial feature vector and a time feature vector, and the fused spatial feature vector and the time feature vector are spliced together according to the channel dimension to form a spliced spatio-temporal feature vector; sum the spliced feature vectors by level and output the fused spatio-temporal feature vector; S5. Establish a classification network of depthwise separable convolution, input the fused spatio-temporal feature vector, and output the probability value of the pedestrian crossing intention.
2. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that, The time features include long-term slow-varying features and / or short-term sudden-change features. The long-term slow-varying features include the instantaneous speed, acceleration, and motion direction angle during the pedestrian movement. The short-term sudden-change features include sudden turning and sudden acceleration during the pedestrian movement.
3. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 2, characterized in that, The time features are extracted by the following method: S101. Set a sliding time window with a length of 5 frames; S102. Extract continuous frame data with a quantity of ≥ 20 frames from the video captured by the in-vehicle camera; S103. Process each frame image in the continuous frame data and real-time locate the pedestrian bounding box; S104. Perform Kalman filtering tracking on the center point of the pedestrian bounding box to generate a smooth pedestrian motion trajectory sequence; S105. Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, acceleration, and motion direction angle of the pedestrian motion trajectory, and establish long-term slow-varying features; S106. Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, motion direction angle of the pedestrian motion trajectory, as well as the angle change and speed change between adjacent trajectory points, and use the captured feature changes as short-term sudden-change features for subsequent pedestrian crossing intention prediction; Set an angle change threshold according to the normal walking angle of the pedestrian. If the change in the normal walking angle of the pedestrian exceeds the set threshold, it is determined that the pedestrian makes a sudden turn; Set a speed change threshold according to the normal walking speed of the pedestrian. If the change in the normal walking speed of the pedestrian exceeds the set threshold, it is determined that the pedestrian makes a sudden acceleration.
4. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 3, characterized in that, The pedestrian pose features include pedestrian key point coordinates, limb angle vectors, and head orientation angle vectors; the scene semantic features include at least one of zebra crossing signs, sidewalk signs, motor vehicle lane signs, and traffic light states.
5. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 4, characterized in that, The pedestrian area within the bounding box is adaptively cropped and normalized by the YOLOv5s-Tiny model to generate a sequence of ROI images with a fixed size; the ROI image sequence is semantically segmented by the MobileNetV3 Small network to extract the semantic features of the pedestrian; The key point coordinates of the pedestrian are extracted by the Lite-HRNet network.
6. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that, A spatial attention model and a temporal attention model are established to process spatial features and temporal features respectively, and the spatial attention weight and the temporal attention weight are output; Spatial weights are added to different levels of the spatial pyramid model, and temporal weights are added to different levels of the temporal pyramid model, and the spatial features and temporal features are concatenated to output a fused and enhanced spatio-temporal feature vector.
7. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 6, characterized in that, The input of the spatial attention model is a multi-scale spatial feature map, and the output is the spatial attention weight. The formula of the spatial attention model is: w s = σ(W2·ReLU(W1·GAP(F s ))) where w s is the spatial attention weight; σ is the sigmoid function; W1 and W2 are learning parameters; GAP is global average pooling; F s is the spatial feature; The input of the temporal attention model is a multi-granularity temporal encoding vector, and the output is the temporal attention weight. The formula of the temporal attention model is: w t = σ(W t · [h t-1 ; v t ) where, w t is the time attention weight; σ is the sigmoid function; W t is the learning parameter; h t-1 is the historical hidden state; v t is the current frame feature.
8. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that The concatenation of spatial and temporal features includes the following steps: The features of each level of the weighted spatial pyramid are concatenated together according to the channel dimension to form a fused spatial feature vector, and the expression is: S fused = Concatenate(w s1 ·S1, w s2 ·S2·...·w sn ·S n ) where S fused is the fused spatial feature vector; w sn is the spatial attention weight at the n-th level; S n is the spatial feature at the n-th level; n is the number of levels of the spatial pyramid; The features of each level of the weighted temporal pyramid are concatenated together according to the channel dimension to form a fused temporal feature vector, and the expression is: T fused = Concatenate(w t1 ·T1, w t2 ·T2·...·w tm ·t m ) where, T fused is the fused temporal feature vector; w tm is the temporal attention weight of the m-th level; T m is the temporal feature of the m-th level; m is the number of levels of the temporal pyramid; The fused spatial feature vector and the temporal feature vector are concatenated together according to the channel dimension to form a fused spatio-temporal feature vector, and the expression is: F fused = Concatenate(S fused , T fused ) where F fused is the fused spatio-temporal feature vector; S fused is the fused spatial feature vector; T fused is the fused temporal feature vector.
9. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that, The establishment of a depthwise separable convolution classification network includes the following steps: a. Depthwise convolution: Apply depthwise convolution operations independently to each channel of the input data. Each input channel is convolved using a separate 3×3 convolution kernel. The number of output channels generated by the depthwise convolution is the same as the number of input channels, and each channel contains the convolution result of that channel; b. Pointwise convolution: Apply pointwise convolution operations to the feature map generated by the depthwise convolution. The pointwise convolution uses a 1×1 convolution kernel and convolves all channels at each position. A linear combination is performed on each pixel point of the feature map to mix the channel features generated by the depthwise convolution; c. After the pointwise convolution, apply the nonlinear activation function ReLU function; d. Perform average pooling.
10. A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that, Set a sliding time window with a length of 5 frames, perform weighted average of the prediction probabilities of 5 consecutive frames by the sliding window, and calculate the confidence through the following formula: Where C is the confidence level; N is the sliding window size; P i is the predicted probability value of the i-th frame within the sliding window; is the mean probability within the window.
Citation Information
Patent Citations
Passenger taxi taking identification method and device, electronic equipment and storage medium
CN118230288A
Automatic trajectory prediction method based on graph spatial-temporal pyramid
WO2024193334A1