Visual space-time pedestrian passing intention prediction method for automatic driving

By constructing a spatial pyramid model and a time pyramid model, and combining a deep separable convolutional network to extract and fuse the time and spatial characteristics of pedestrian movement, the problem of pedestrian travel intention recognition in complex traffic environments is solved, and higher recognition accuracy and computing efficiency are achieved.

CN120047924AActive Publication Date: 2025-05-27NANCHANG INST OF TECH

Patent Information

Application Number
CN202510525077.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-05-27
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify pedestrians' intention to pass through in complex traffic environments, making it difficult to ensure the safety and stability of self-driving cars when interacting with pedestrians.

Method used

By extracting the temporal characteristics of pedestrian movement and inputting a bidirectional LSTM network, a time context encoding vector is generated; at the same time, a spatial pyramid model of pedestrian pose and scene and a time pyramid model of pedestrian movement trajectory and velocity are constructed, and spatial and temporal characteristics are integrated; finally, a classification network with deep separable convolution is used to predict pedestrian travel intentions.

Benefits of technology

It improves the accuracy of pedestrian intention recognition of self-driving cars in complex scenarios, enhances feature discrimination capabilities, and improves the generalization ability and computing efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047924A_ABST
    Figure CN120047924A_ABST
Patent Text Reader

Abstract

The invention records a visual space-time pedestrian passing intention prediction method for automatic driving, and the method comprises the following steps: S1, extracting pedestrian motion time features, inputting the features into a bidirectional LSTM network, and outputting a time context coding vector; s2, constructing a spatial pyramid model and a time pyramid model; s4, weighting each level feature of the space-time pyramid model, splicing the weighted space features and time features to form spliced space-time feature vectors, summing the spliced space-time feature vectors according to levels, and outputting fused space-time feature vectors; and S5, establishing a classification network of a depth separable convolution and outputting a pedestrian passing intention probability value. Through joint modeling of the space-time pyramid model, the problem of space-time feature splitting is solved. The spatial features and the time features are respectively processed through a double-path attention module, a spatial attention weight and a time attention weight are output, and the feature discrimination capability of automatic driving in a complex scene is improved by adopting space-time attention weight fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of autonomous driving, and particularly to a visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving. Background Art

[0002] Autonomous vehicles can eliminate or reduce factors such as fatigue, distraction, and drunk driving in human driving, thus significantly improving traffic safety. However, autonomous vehicles have difficulties in the safe and stable interaction with crossing pedestrians, which has become one of the important obstacles restricting the implementation of high-level autonomous vehicles on urban streets. Therefore, accurately and real-time predicting the future movement intention of pedestrians is the basic guarantee for solving the conflict between humans and vehicles and improving road driving safety.

[0003] Pedestrian crossing intention prediction is realized based on pedestrian detection and tracking. Generally, scholars researching pedestrian crossing intention prediction default that pedestrians have been detected and tracked. Currently, pedestrian crossing intention prediction is generally achieved through the following methods, for example:

[0004] Based on video images, relevant information about pedestrian crossing intention is obtained by extracting MCHOG features, and an SVM classifier is combined to determine whether the pedestrians on the street cross the road;

[0005] 18 human key point skeleton points are located through a pose estimation positioning algorithm, and a 396-dimensional feature vector is calculated by calculating the mathematical relationship between the above points. The 396-dimensional feature vector obtained from the image is input into the SVM, and the pedestrian movement state is obtained through the classifier;

[0006] The pedestrian frame data is cropped into two images to capture the head and leg information respectively, feature extraction is performed using AlexNet, and finally the classifier is used to predict the action intention of the pedestrian. This network proves that the combination accuracy of a CNN and a machine learning classifier far exceeds other traditional methods;

[0007] With the development of recurrent neural networks, some research scholars have proposed the ConvLSTM network. The ConvLSTM network takes images as input, preprocesses them with a pre-trained CNN, and inputs the extracted features into a convolutional LSTM. The last hidden state is input into a fully connected layer for prediction.

[0008] The above research methods are proposed based on spatio-temporal points, human geometric features, or motion information, with a single information source. Once the driving environment is complex or lacks information, the prediction effect is poor. Therefore, the current intention recognition algorithm is difficult to accurately recognize the pedestrian crossing intention by autonomous vehicles in a complex traffic environment. Summary of the Invention

[0009] To meet the need for accurate recognition of pedestrians' crossing intentions in autonomous vehicles, the present invention describes a visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving, which has the advantages of high computational efficiency and accurate intention recognition. The specific technical solutions are as follows:

[0010] A visual spatio-temporal pedestrian crossing intention prediction method for autonomous driving, comprising the following steps:

[0011] S1, Extract the temporal features of the pedestrian movement in the video captured by the on-vehicle camera, input them into a bidirectional LSTM network, and output a temporal context encoding vector;

[0012] S2, Perform max pooling on the pedestrian's pose features and average pooling on the scene semantic features to generate a multi-scale spatial feature map, and construct a spatial pyramid model of the pedestrian pose and the scene;

[0013] S3, Downsample the temporal context encoding vector in the temporal dimension, extract the motion features of different temporal granularities, generate a multi-granularity temporal encoding vector, and construct a temporal pyramid model of the pedestrian motion trajectory and speed;

[0014] S4, Weight each level feature of the spatial pyramid model and the temporal pyramid model respectively; splice the weighted spatial features and temporal features to form a spliced spatio-temporal feature vector; sum the spliced feature vectors by level and output a fused spatio-temporal feature vector;

[0015] S5, Establish a classification network with depthwise separable convolution, input the fused spatio-temporal feature vector, and output the pedestrian crossing intention probability value P∈[0,1].

[0016] Further, the temporal features include long-term slow-changing features and / or short-term sudden-changing features. The long-term slow-changing features include the instantaneous speed, acceleration, and motion direction angle of the pedestrian during movement, and the short-term sudden-changing features include sudden turning and sudden acceleration of the pedestrian during movement.

[0017] Further, the temporal features are extracted by the following method:

[0018] S101, Set a sliding time window with a length of 5 frames;

[0019] S102, Extract continuous frame data with a quantity ≧20 frames from the video captured by the on-vehicle camera;

[0020] S103, Process each frame image in the continuous frame data and real-time locate the pedestrian bounding box;

[0021] S104, Perform Kalman filter tracking on the center point of the pedestrian bounding box to generate a smooth pedestrian motion trajectory sequence;

[0022] S105. With the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, acceleration, and motion direction angle of the pedestrian's motion trajectory, and establish long-term slow-varying features.

[0023] S106. With the vehicle coordinate system as the reference coordinate system, calculate the instantaneous speed, motion direction angle of the pedestrian's motion trajectory, as well as the angle change and speed change between adjacent trajectory points, and use the captured feature changes as short-term mutation features for subsequent pedestrian crossing intention prediction; set an angle change threshold according to the normal walking angle of the pedestrian, and if the change in the normal walking angle of the pedestrian exceeds the set threshold, it is determined that the pedestrian has suddenly turned; set an angle change threshold according to the normal walking speed of the pedestrian, and if the change in the normal walking speed of the pedestrian exceeds the set threshold, it is determined that the pedestrian has suddenly accelerated.

[0024] Furthermore, the pose features of the pedestrian include the coordinates of the pedestrian's key points, the limb angle vector, and the head orientation angle vector; the scene semantic features include at least one of the zebra crossing sign, the sidewalk sign, the motor vehicle lane sign, and the traffic light state.

[0025] Furthermore, the YOLOv5s-Tiny model is used to perform adaptive cropping and normalization processing on the pedestrian area within the bounding box to generate a sequence of ROI images with a fixed size; the MobileNetV3 Small network is used to perform semantic segmentation on the ROI image sequence to extract the semantic features of the pedestrian; the Lite-HRNet network is used to extract the coordinates of the pedestrian's key points.

[0026] Furthermore, a spatial attention model and a temporal attention model are established to process spatial features and temporal features respectively, and output spatial attention weights and temporal attention weights; spatial weights are added to different levels of the spatial pyramid model, and temporal weights are added to different levels of the temporal pyramid model, and the spatial features and temporal features are concatenated to output a fused and enhanced spatio-temporal feature vector.

[0027] Furthermore, the input of the spatial attention model is a multi-scale spatial feature map, and the output is the spatial attention weight. The formula of the spatial attention model is:

[0028]

[0029] In the formula, is the spatial attention weight; is the sigmoid function; and are learning parameters; is the global average pooling; is the spatial feature;

[0030] The input of the temporal attention model is a multi-granularity temporal encoding vector, and the output is the temporal attention weight. The formula of the temporal attention model is:

[0031]

[0032] Wherein, is the time attention weight; is the sigmoid function; is the learning parameter; is the historical hidden state; is the current frame feature.

[0033] Furthermore, splicing the spatial and temporal features includes the following steps:

[0034] The weighted features of each level of the spatial pyramid are spliced together along the channel dimension to form a fused spatial feature vector, and the expression is:

[0035]

[0036] Wherein, is the fused spatial feature vector; is the spatial attention weight of the nth level; is the spatial feature of the nth level; n is the number of levels of the spatial pyramid;

[0037] The weighted features of each level of the temporal pyramid are spliced together along the channel dimension to form a fused temporal feature vector, and the expression is:

[0038]

[0039] Wherein, is the fused temporal feature vector; is the temporal attention weight of the mth level; is the temporal feature of the mth level; m is the number of levels of the temporal pyramid;

[0040] The fused spatial feature vector and the fused temporal feature vector are spliced together along the channel dimension to form a fused spatio-temporal feature vector, and the expression is:

[0041]

[0042] Wherein, is the fused spatio-temporal feature vector; is the fused spatial feature vector; is the fused temporal feature vector.

[0043] Furthermore, establishing a classification network of depthwise separable convolution includes the following steps:

[0044] a. Depth convolution: Apply depth convolution operations independently to each channel of the input data. For each input channel, a separate 3×3 convolutional kernel is used for convolution. The number of output channels generated by depth convolution is the same as the number of input channels, and each channel contains the convolution result of that channel;

[0045] b. Pointwise convolution: Apply pointwise convolution operations to the feature maps generated by depth convolution. Pointwise convolution uses a 1×1 convolutional kernel and performs convolution on all channels at each position, linearly combining each pixel point of the feature map to mix the feature of each channel generated by depth convolution;

[0046] c. After pointwise convolution, apply the ReLU function as the non - linear activation function;

[0047] d. Perform average pooling.

[0048] Furthermore, set a sliding time window with a length of 5 frames, perform weighted average of the prediction probabilities of 5 consecutive frames using the sliding window, and calculate the confidence through the following formula:

[0049]

[0050] In the formula, is the confidence; is the sliding window size; is the predicted probability value of the i - th frame within the sliding window; is the average probability within the window.

[0051] Based on the above technical solutions, the present invention has the following beneficial effects:

[0052] 1. By jointly modeling the spatial pyramid model of pedestrian pose and scene and the temporal pyramid model of pedestrian motion trajectory and speed, the problem of spatio - temporal feature fragmentation in the prior art is solved.

[0053] 2. Establish a dynamic spatio - temporal attention mechanism. The spatial and temporal features are processed separately through a dual - path attention module to output spatial attention weights and temporal attention weights, and a fusion strategy of spatial - temporal attention weights is adopted to improve the feature discrimination ability of autonomous driving in complex scenarios.

[0054] 3. Through the spatio - temporal context fusion model and the dynamic spatio - temporal attention mechanism, compared with the traditional LSTM baseline model, the prediction accuracy of the pedestrian crossing intention prediction method recorded in the present invention has been significantly improved. For example, the mean average precision (mAP) of the lightweight YOLOv5s - Tiny model in object detection has increased by 2.26%, which is comparable to the YOLOv8s model, but the number of parameters is only 53% of the YOLOv8s model, as shown in the following table:

[0055] Parameter comparison data table of the lightweight YOLOv5s-Tiny model and the YOLOv8s model

[0056] Model Mean Average Precision (mAP) Number of parameters (M) FLOPs (B) Yolov5s-Tiny Improved by 2.26% compared to the baseline model 4.2 8.7 Yolov8s 44.9 11.2 28.6 Brief Description of the Drawings

[0057] Figure 1 : Schematic diagram of the overall process of the method described in the present invention;

[0058] Figure 2 : Schematic diagram of the structure of the spatio-temporal feature pyramid and the attention mechanism. Detailed Description of the Invention

[0059] It should be noted that:

[0060] 1. Certain terms are used in the specification and claims to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. The specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction. Unless otherwise defined, the technical terms or scientific terms used in this disclosure should have the ordinary meaning understood by those of ordinary skill in the art within the field to which this disclosure belongs.

[0061] 2. In this embodiment, a monocular camera is used to capture videos and / or images. On-vehicle monocular cameras are currently widely used in the field of autonomous driving technology. Through a monocular camera, it is possible to capture images and also collect continuous video streams of the driving environment. Each static image of the captured images and / or videos is defined as an original image. This embodiment does not make improvements to the structure, principle, operation mode, etc. of the monocular camera, and will not be elaborated further.

[0062] As shown in Figure 1 and Figure 2 shown, this embodiment describes a method for predicting the intention of a pedestrian to cross in visual spatio-temporal for autonomous driving, including the following steps.‌

[0063] S1. Extract the temporal features of pedestrian movement from the video captured by the on-vehicle monocular camera, and input the temporal features into a bidirectional LSTM network to output a temporal context encoding vector.

[0064] The temporal features include long-term slow-changing features and / or short-term sudden-changing features. The long-term slow-changing features refer to the instantaneous speed, acceleration, and motion direction angle of the pedestrian during movement, and the short-term sudden-changing features refer to the sudden turning and sudden acceleration of the pedestrian during movement.

[0065] The extraction of the temporal features of pedestrian movement preferably adopts the following method or steps, specifically:

[0066] S101. Set a sliding time window

[0067] Set a sliding time window with a length of 5 frames to capture the short-term dynamic features of pedestrian movement.

[0068] S102, Extract consecutive frame data

[0069] Extract consecutive frame data from the video captured by the on-vehicle monocular camera, and the number of consecutive frame data ≥ 20 frames.

[0070] S103, Locate the pedestrian bounding box

[0071] Use the lightweight object detection model YOLOv5s-Tiny to process each frame in the consecutive frame data and locate the pedestrian bounding box in real time.

[0072] S104, Kalman filter tracking

[0073] Perform Kalman filter tracking on the center point of the pedestrian bounding box to generate a smooth sequence of pedestrian movement trajectories.

[0074] S105, Extract the long-term slow-varying features of pedestrian movement

[0075] Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous velocity, acceleration, and movement direction angle of the pedestrian movement trajectory to establish long-term slow-varying features.

[0076] In this step, the calculation method of "instantaneous velocity, acceleration, and movement direction angle" adopts the method commonly used in this field, that is, the instantaneous velocity and acceleration are obtained by calculating the derivative after smoothing the trajectory through Kalman filter, and then taking the vehicle coordinate system as the reference coordinate system, the movement direction angle is calculated through the geometric relationship of the trajectory. Briefly described as follows: Instantaneous velocity, after smoothing the trajectory through Kalman filter, calculate the displacement difference between adjacent frames divided by the time difference; Acceleration, by calculating the rate of change of velocity; Movement direction angle, obtained by calculating the direction vector of the pedestrian movement trajectory.

[0077] S106, Extract the short-term sudden change features of pedestrian movement

[0078] S106-1, Taking the vehicle coordinate system as the reference coordinate system, calculate the instantaneous velocity, movement direction angle, and the angle change and velocity change between adjacent trajectory points of the pedestrian movement trajectory, and use the captured feature changes as short-term sudden change features for subsequent pedestrian crossing intention prediction;

[0079] S106-2, Set an angle change threshold according to the normal walking angle of the pedestrian. If the angle change exceeds the set threshold, it is determined as a sudden turn;

[0080] S106-3, Set a speed change threshold according to the normal walking speed of the pedestrian. If the speed change exceeds the set threshold, it is determined as a sudden acceleration.

[0081] S2, perform max pooling on the pedestrian's pose features and average pooling on the scene semantic features to generate a multi-scale spatial feature map, and construct a spatial pyramid model of the pedestrian pose and the scene.

[0082] Among them:

[0083] ① The pedestrian's pose features include the coordinates of 17 pedestrian key points, limb angle vectors, and head orientation angle vectors.

[0084] The coordinates of 17 pedestrian key points refer to the head, shoulders, elbows, hips, knees, ankles, etc. of the pedestrian, and the head includes at least the left ear, right ear, and nose tip;

[0085] The limb angle vector is calculated based on the angle between the elbow-shoulder-hip three points and is used to judge the movement trend of the pedestrian's upper limb;

[0086] The head orientation angle vector is calculated based on the geometric relationship of the left ear, right ear, and nose tip and is used to judge the line of sight direction of the pedestrian.

[0087] ② The scene semantic features include but are not limited to zebra crossing signs, sidewalk signs, motor vehicle lane signs, and traffic light states.

[0088] S3, perform downsampling on the time context encoding vector in the time dimension, extract motion features of different time granularities, generate a multi-granularity time encoding vector, and construct a time pyramid model of the pedestrian motion trajectory and speed.

[0089] In this step:

[0090] ① Preferably, use the lightweight object detection model YOLOv5s-Tiny to perform adaptive cropping and normalization processing on the pedestrian area within the bounding box to generate a sequence of ROI images with a fixed size;

[0091] ② Preferably, use the lightweight MobileNetV3 Small network to perform semantic segmentation on the ROI image sequence to accurately extract the semantic features of the pedestrian;

[0092] ③ Preferably, use the Lite-HRNet network to extract the coordinates of the pedestrian key points.

[0093] S4, construct a spatio-temporal context fusion model, including the following steps:

[0094] S401, weight each level feature of the spatial pyramid model and the time pyramid model respectively. For example, weight the pedestrian pose features and scene semantic features of different scales, and the weights are dynamically adjusted according to learning in the fully connected layer and the convolutional layer;

[0095] S402. Concatenate the weighted spatial features and temporal features to form a concatenated feature vector;

[0096] S403. Sum the concatenated feature vector hierarchically to obtain the final spatio-temporal context fusion feature.

[0097] Through the joint modeling of the spatial pyramid model and the temporal pyramid model, the problem of spatio-temporal feature fragmentation in the existing technology can be solved, improving the model prediction accuracy, enhancing the feature discrimination ability, and promoting the model generalization ability.

[0098] S5. To enhance the feature discrimination ability of autonomous driving in complex scenarios, this step describes a dynamic spatio-temporal attention mechanism that processes spatial features and temporal features through a dual-path attention module respectively and outputs spatial attention weights and temporal attention weights. It specifically includes the following steps:

[0099] S501. Establish a spatial attention model, input the multi-scale spatial feature map into the spatial attention model, and output spatial attention weights. The spatial attention model is used to calculate the importance weights of each channel. The formula is:

[0100]

[0101] In the formula,

[0102] is the spatial attention weight;

[0103] is the sigmoid function;

[0104] and are learning parameters;

[0105] is the global average pooling;

[0106] is the spatial feature.

[0107] S502. Establish a temporal attention model, input the multi-granularity temporal encoding vector into the temporal attention model, and output temporal attention weights. The temporal attention model calculates the temporal weights through a gating mechanism (GRU unit). The formula is:

[0108]

[0109] In the formula,

[0110] is the temporal attention weight;

[0111] is the sigmoid function;

[0112] is a learning parameter;

[0113] is a historical hidden state;

[0114] is the current frame feature.

[0115] S6. Add spatial weights to different levels of the spatial pyramid model and add temporal weights to different levels of the temporal pyramid model to enhance the attention of the spatio-temporal context fusion model, and then splice the spatial and temporal features to output an enhanced spatio-temporal feature vector.

[0116] Among them, splicing the spatial and temporal features includes the following steps:

[0117] S601. Spatial feature splicing

[0118] Splice the features of each level of the weighted spatial pyramid together along the channel dimension to form a fused spatial feature vector , and the expression is:

[0119]

[0120] In the formula,

[0121] is the fused spatial feature vector;

[0122] is the spatial attention weight of the nth level;

[0123] is the spatial feature of the nth level;

[0124] n is the number of levels of the spatial pyramid.

[0125] S602. Temporal feature splicing

[0126] Splice the features of each level of the weighted temporal pyramid together along the channel dimension to form a fused temporal feature vector , and the expression is:

[0127]

[0128] In the formula,

[0129] is the fused temporal feature vector;

[0130] is the time attention weight of the m-th layer;

[0131] is the time feature of the m-th layer;

[0132] m is the number of layers of the time pyramid.

[0133] S603, Spatial-Temporal Feature Concatenation

[0134] Concatenate the fused spatial feature vector and the time feature vector along the channel dimension to form the final fused and enhanced spatio-temporal feature vector , and the expression is:

[0135]

[0136] In the formula,

[0137] is the fused spatio-temporal feature vector;

[0138] is the fused spatial feature vector;

[0139] is the fused time feature vector.

[0140] S7, Build a lightweight classification network with depthwise separable convolution, input the fused spatio-temporal feature vector , and output the pedestrian crossing intention probability value P ∈ [0, 1].

[0141] Among them, building a lightweight classification network with depthwise separable convolution includes the following steps:

[0142] S701, Depthwise Convolution

[0143] Apply depthwise convolution operation independently to each channel of the input data. Each input channel is convolved using a separate 3×3 convolution kernel. The number of output channels generated by depthwise convolution is the same as the number of input channels, and each channel contains the convolution result of that channel.

[0144] S702, Pointwise Convolution

[0145] Apply pointwise convolution operation to the feature map generated by depthwise convolution. Pointwise convolution uses a 1×1 convolution kernel and convolves all channels at each position. This operation performs a linear combination of each pixel point of the feature map and mixes the channel features generated by depthwise convolution. The number of output channels of pointwise convolution can be specified flexibly and is usually used to control the depth of the output feature map.

[0146] S703, Nonlinear Activation Function

[0147] After pointwise convolution, the ReLU function, a non-linear activation function, is applied.

[0148] S704, Pooling

[0149] Average pooling is performed to reduce the spatial dimension of the feature map while retaining important features.

[0150] S8, Perform a sliding window weighted average on the predicted probabilities of 5 consecutive frames to suppress transient noise and calculate the confidence through the following formula:

[0151]

[0152] In the formula,

[0153] is the confidence;

[0154] is the sliding window size;

[0155] is the predicted probability value of the i-th frame within the sliding window;

[0156] is the average probability within the window.

[0157] The real-time prediction method for pedestrian crossing intention described in this embodiment can also improve the inference speed of the model by performing structured pruning, lightweight deployment, and priority scheduling on the model. For example:

[0158] 1. Remove redundant convolutional channels in MobileNetV3 (channel pruning rate 30%) and retain high-activation feature layers.

[0159] 2. Use TensorRT to convert the FP32 model to INT8 precision. After testing, the model size is compressed to 2.3MB and the inference speed is increased by 3 times.

[0160] 3. Dynamically allocate computing resources according to the pedestrian risk level and distance:

[0161] High-risk pedestrians: Allocate 90% of the GPU computing power, prediction frequency 30Hz;

[0162] Medium- and low-risk pedestrians: Allocate 10% of the GPU computing power, prediction frequency 10Hz;

[0163] Adopt thread pool management to ensure that the single-frame processing delay ≤ 30ms.

[0164] The basic principles, main features and advantages of the present invention have been shown and described above. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and what is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving, characterized in that: The following steps are involved: S1, extracts the temporal features of pedestrian motion in the video captured by the vehicle camera, inputs them into the bidirectional LSTM network, and outputs the temporal context encoding vector; S2, performs maximum pooling on the pedestrian posture features and average pooling on the scene semantic features to generate a multi-scale spatial feature map and construct a spatial pyramid model of pedestrian posture and scene; S3, downsample the temporal context encoding vector in the time dimension, extract motion features of different time granularities, generate multi-granularity temporal encoding vectors, and construct a temporal pyramid model of pedestrian motion trajectory and speed; S4, weighting each level feature of the spatial pyramid model and the temporal pyramid model respectively; splicing the weighted spatial features and temporal features to form a spliced ​​spatiotemporal feature vector; summing the spliced ​​feature vectors by level, and outputting a fused spatiotemporal feature vector; S5, establish a deep separable convolutional classification network, input the fused spatiotemporal feature vector, and output the probability value of the pedestrian's intention to cross.

2. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that: The temporal features include long-term slowly changing features and / or short-term suddenly changing features. The long-term slowly changing features include the instantaneous speed, acceleration and moving direction angle of the pedestrian when moving. The short-term suddenly changing features include sudden turns and sudden accelerations of the pedestrian when moving.

3. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 2, characterized in that: The temporal features are extracted by the following method: S101, setting a sliding time window with a length of 5 frames; S102, extracting continuous frame data of 20 frames or more from the video captured by the vehicle-mounted camera; S103, processing each frame image in the continuous frame data, and locating the pedestrian boundary box in real time; S104, performing Kalman filter tracking on the center point of the pedestrian boundary box to generate a smooth pedestrian motion trajectory sequence; S105, using the vehicle coordinate system as a reference coordinate system, calculating the instantaneous speed, acceleration and motion direction angle of the pedestrian's motion trajectory, and establishing a long-term slow-changing feature; S106, using the vehicle coordinate system as a reference coordinate system, calculating the instantaneous speed, the moving direction angle, and the angle change and speed change between adjacent trajectory points of the pedestrian's moving trajectory, and using the captured feature changes as short-term mutation features for subsequent pedestrian crossing intention prediction; An angle change threshold is set according to the normal walking angle of the pedestrian. If the normal walking angle of the pedestrian changes beyond the set threshold, it is judged as a sudden turn of the pedestrian. The angle change threshold is set according to the normal walking speed of the pedestrian. If the normal walking speed of the pedestrian changes beyond the set threshold, it is judged as a sudden acceleration of the pedestrian.

4. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 3, characterized in that: The posture features of pedestrians include the coordinates of pedestrian key points, body angle vectors, and head orientation angle vectors; the scene semantic features include at least one of zebra crossing signs, sidewalk signs, motor vehicle lane signs, and traffic light status.

5. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 4, characterized in that: The pedestrian area within the bounding box is adaptively cropped and normalized using the YOLOv5s-Tiny model to generate a ROI image sequence with a fixed size. The ROI image sequence is semantically segmented using the MobileNetV3 Small network to extract the semantic features of pedestrians. The key point coordinates of pedestrians are extracted through the Lite-HRNet network.

6. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that: Establish a spatial attention model and a temporal attention model to process spatial features and temporal features respectively, and output spatial attention weights and temporal attention weights; Spatial weights are added to different levels of the spatial pyramid model, and temporal weights are added to different levels of the temporal pyramid model. The spatial features and temporal features are concatenated to output the fused and enhanced spatiotemporal feature vector.

7. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 6, characterized in that: The input of the spatial attention model is a multi-scale spatial feature map, and the output is the spatial attention weight. The formula of the spatial attention model is: In the formula, is the spatial attention weight; is the sigmoid function; and is the learning parameter; is global average pooling; For spatial characteristics; The input of the temporal attention model is a multi-granularity temporal encoding vector, and the output is the temporal attention weight. The temporal attention model formula is: In the formula, is the temporal attention weight; is the sigmoid function; is the learning parameter; Hidden state for history; is the current frame feature.

8. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that: The stitching of spatial and temporal features consists of the following steps: The weighted features of each level of the spatial pyramid are spliced ​​together according to the channel dimension to form a fused spatial feature vector, which is expressed as: In the formula, is the fused spatial feature vector; is the spatial attention weight of the nth level; is the spatial feature of the nth level; n is the number of levels of the spatial pyramid; The weighted time pyramid features at each level are spliced ​​together according to the channel dimension to form a fused time feature vector, which is expressed as: In the formula, is the fused time feature vector; is the temporal attention weight of the mth level; is the time feature of the mth level; m is the number of levels in the time pyramid; The fused spatial feature vector and temporal feature vector are concatenated together according to the channel dimension to form a fused spatiotemporal feature vector, which is expressed as: In the formula, is the fused spatiotemporal feature vector; is the fused spatial feature vector; is the fused time feature vector.

9. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that: Building a depthwise separable convolutional classification network consists of the following steps: a. Depth convolution: Depth convolution operation is applied independently to each channel of the input data. Each input channel is convolved using a separate 3×3 convolution kernel. The number of output channels generated by depth convolution is the same as the number of input channels, and each channel contains the convolution result of that channel. b. Point-by-point convolution: Apply point-by-point convolution to the feature map generated by deep convolution. Point-by-point convolution uses a 1×1 convolution kernel, performs convolution operations on all channels at each position, linearly combines each pixel of the feature map, and mixes the channel features generated by deep convolution. c. After point-by-point convolution, apply the nonlinear activation function ReLU function; d. Perform average pooling.

10. The visual spatiotemporal pedestrian crossing intention prediction method for autonomous driving according to claim 1, characterized in that: Set a sliding time window with a length of 5 frames, perform sliding window weighted average on the prediction probability of 5 consecutive frames, and calculate the confidence using the following formula: In the formula, is the confidence level; is the sliding window size; is the predicted probability value of the i-th frame in the sliding window; is the mean probability within the window.

Citation Information

Patent Citations

  • Multi-dimensional convolutional neural network learner modeling method for multi-source heterogeneous data

    CN112529054A

  • End-to-end automatic driving behavior decision-making method and system and terminal equipment

    CN113139446A

  • Construction method and application of 3D human body posture estimation model

    CN113205595A

  • Automatic driving track prediction method based on space-time pyramid

    CN115049130A

  • Pedestrian crossing intention recognition method based on multi-source information fusion

    CN117173663A

Cited By

  • Automatic driving forward collision early warning method based on capsule network fusion perception

    CN120913177A