Method, device and storage medium for predicting pedestrian crossing intention

By constructing a pedestrian interaction graph model, combining visual, semantic and space-time dynamic features, the graph convolution network is used to predict pedestrian crossing intentions, which solves the problem of insufficient prediction accuracy in complex traffic scenarios and achieves higher prediction accuracy.

CN117152793BActive Publication Date: 2025-08-15CHONGQING CHANGAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311090808.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-28
Publication Date
2025-08-15
Estimated Expiration
2043-08-28

AI Technical Summary

Technical Problem

The prior art fails to fully consider the influence of multiple interactive objects when predicting pedestrian crossing intentions in complex traffic scenarios, resulting in low accuracy of prediction results.

Method used

By obtaining video data of traffic scenes, identifying the target object and constructing a pedestrian interaction graph model, combining the graph convolution network, considering the visual characteristics, semantic characteristics and space-time dynamic characteristics of the target pedestrian and its multiple interaction objects, a pedestrian interaction graph model is constructed to predict the intention to cross the street.

Benefits of technology

The accuracy of pedestrian crossing intention prediction in complex traffic scenarios is improved, and the influence of multiple interactive objects is taken into account, which enhances the accuracy of prediction results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117152793B_ABST
    Figure CN117152793B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and storage medium for predicting the intention of pedestrians to cross the street. The method includes: obtaining video data of the current traffic scene and identifying multiple target objects; determining the target pedestrian among the multiple target objects and multiple interactive objects corresponding to the target pedestrian, respectively determining the feature sets of the target pedestrian and the multiple interactive objects, and constructing a pedestrian interaction graph model based on the feature sets of the target pedestrian and the multiple interactive objects; inputting the pedestrian interaction graph model into a graph convolution-based network to obtain a prediction result of the target pedestrian's intention to cross the street. Among them, the multiple interactive objects include mobile interactive objects and non-mobile interactive objects, the feature sets of the mobile interactive objects and the target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of the non-mobile interactive objects include semantic features and visual features. The present application can predict the intention of pedestrians to cross the street in complex traffic scenes, and the accuracy of the prediction results is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of pedestrian crossing intention prediction in autonomous driving, and specifically to a method, device and storage medium for predicting pedestrian crossing intention. Background Art

[0002] With the rapid development of autonomous driving technology in recent years, research related to autonomous driving has become a hot topic. Accurately predicting pedestrians' intention to cross the street based on video data captured by onboard video equipment is one of the current challenges.

[0003] To address this, existing technologies first model pedestrian interactions from a single perspective, focusing on visual and dynamic features, and extracting interaction features. They then predict pedestrians' intention to cross the street based on these interaction features. For example, a method for predicting pedestrians' crossing intentions using an autonomous vehicle and a method for predicting pedestrian trajectories based on autonomous aerial photography respectively model pedestrian interactions based on the visual features of key areas such as the pedestrian's head and torso, and the pedestrian's dynamic features. However, in complex traffic scenarios, methods that rely solely on visual or dynamic information to model interactions and obtain interaction features are often insufficient. The main reasons for this are: First, pedestrian behavior and intentions in traffic scenarios are influenced by multiple interacting objects, and modeling interactions solely from the pedestrian perspective can seriously affect the accuracy of prediction results. Second, in complex traffic scenarios, pedestrian interaction features are complex and changeable, and modeling based solely on a single feature can also affect the accuracy of prediction results.

[0004] Therefore, the method used in the existing technology to predict pedestrian crossing intention has the problems of not considering complex traffic scenarios, having a small adaptability range and low accuracy of prediction results. Summary of the Invention

[0005] The purpose of the embodiments of the present application is to provide a method, device and storage medium for predicting pedestrian crossing intentions, so as to solve the problems that the methods used in the prior art for predicting pedestrian crossing intentions do not take complex traffic scenarios into consideration, have a small scope of adaptability and low accuracy of prediction results.

[0006] To achieve the above objectives, the present application provides, in a first aspect, a method for predicting pedestrian crossing intention, the method comprising:

[0007] Obtain video data of the current traffic scene;

[0008] Identify multiple target objects in video data;

[0009] determining a target pedestrian among multiple target objects and multiple interaction objects corresponding to the target pedestrian;

[0010] Determine the feature sets of the target pedestrian and multiple interaction objects respectively;

[0011] Construct a pedestrian interaction graph model based on the feature sets of the target pedestrian and multiple interaction objects;

[0012] The pedestrian interaction graph model is input into a graph convolution-based network to obtain the prediction result of the target pedestrian's crossing intention;

[0013] Among them, multiple interactive objects include mobile interactive objects and non-mobile interactive objects. The feature sets of mobile interactive objects and target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of non-mobile interactive objects include semantic features and visual features.

[0014] In this embodiment of the present application, the feature set for determining the target pedestrian and multiple interaction objects includes:

[0015] The video image regions corresponding to the target pedestrian and multiple interactive objects are input into the visual feature extraction module to obtain high-level invisible features corresponding to the target object and multiple interactive objects respectively;

[0016] The high-dimensional latent features of the target pedestrian and multiple interacting objects are converted into the same size through the spatial pyramid model to obtain the visual features of the target pedestrian and multiple interacting objects respectively.

[0017] In the embodiment of the present application, the feature sets for respectively determining the target pedestrian and the multiple interaction objects include:

[0018] Filtering out mobile interactive objects from multiple interactive objects;

[0019] Obtain historical trajectory data of mobile interaction objects and target pedestrians respectively;

[0020] The historical trajectory data of the preset length of the mobile interactive object and the target pedestrian are respectively input into the gated recurrent unit to obtain the spatiotemporal dynamic characteristics of the target pedestrian and the mobile interactive object.

[0021] In an embodiment of the present application, constructing a pedestrian interaction graph model based on a feature set of a target pedestrian and multiple interaction objects includes:

[0022] Based on the feature set, the interaction features of the target pedestrian and each interaction object are extracted respectively to obtain multiple interaction features;

[0023] Construct the target adjacency matrix of the pedestrian interaction graph model;

[0024] Multiple interaction features are used as node feature vectors of different target nodes, and a pedestrian interaction graph model is constructed based on graph theory and the target adjacency matrix.

[0025] In an embodiment of the present application, when the interactive object is a mobile interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes:

[0026] Splicing the visual features and semantic features of the target pedestrian and the mobile interactive object to obtain the spliced features;

[0027] The concatenated features are fused with the spatiotemporal dynamic features through the self-attention mechanism to obtain interactive features.

[0028] In an embodiment of the present application, when the interactive object is a non-moving interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes:

[0029] The visual features and semantic features of target pedestrians and non-moving objects are fused through the self-attention mechanism to obtain interactive features.

[0030] In the embodiment of the present application, the target adjacency matrix for constructing the pedestrian interaction graph model includes:

[0031] constructing a first adjacency matrix according to the spatial distribution between the plurality of interacting objects;

[0032] constructing a second adjacency matrix according to the relative heading angles between the multiple interacting objects;

[0033] The first adjacency matrix and the second adjacency matrix are fused to obtain a target adjacency matrix.

[0034] In the embodiment of the present application, determining a target pedestrian among multiple target objects includes:

[0035] Filter out pedestrians from the target objects;

[0036] Determine the detection box area of the pedestrian;

[0037] When the detection box area is greater than or equal to the set area threshold, the pedestrian is determined as a target pedestrian.

[0038] In an embodiment of the present application, the interactive object corresponding to the target pedestrian is a target object among multiple target objects whose Euclidean distance to the target pedestrian is less than or equal to a set distance threshold.

[0039] A second aspect of the present application provides a device for predicting pedestrian crossing intention, comprising:

[0040] a memory configured to store instructions; and

[0041] The processor is configured to call instructions from the memory and implement the above-mentioned method for predicting the intention of pedestrians crossing the street when executing the instructions.

[0042] A third aspect of the present application provides a machine-readable storage medium having stored thereon instructions for enabling a machine to execute the above-mentioned method for predicting pedestrian crossing intentions.

[0043] Beneficial effects of this application:

[0044] (1) This application models the target pedestrian and its corresponding interactive target, taking into account that the target pedestrian's intention to cross the street in the traffic scene will be affected by multiple interactive objects. The modeling of the target pedestrian and its multiple interactive objects is conducive to improving the accuracy of the prediction results;

[0045] (2) This application models the model based on multiple dimensions including visual features, semantic features, and spatiotemporal dynamic features, taking into account the complex and changeable characteristics of pedestrian interaction in complex traffic scenarios, which is conducive to improving the accuracy of prediction results.

[0046] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present application but do not constitute a limitation on the embodiments of the present application. In the accompanying drawings:

[0048] Figure 1 A flowchart of a method for predicting pedestrian crossing intention provided in an embodiment of the present application;

[0049] Figure 2 A flowchart of a method for extracting interactive features provided in a specific embodiment of the present application;

[0050] Figure 3 A flowchart of a method for fusing multi-source features of a mobile interactive object provided in a specific embodiment of the present application;

[0051] Figure 4 A flowchart of a method for multi-source feature fusion of non-moving interactive objects provided in a specific embodiment of the present application;

[0052] Figure 5 A schematic diagram showing that the relative orientation angles between different interactive objects provided in a specific embodiment of the present application are in the first quadrant;

[0053] Figure 6 A schematic diagram showing that the relative orientation angles between different interactive objects provided in a specific embodiment of the present application are in the second quadrant;

[0054] Figure 7A schematic diagram showing that the relative orientation angles between different interactive objects provided in a specific embodiment of the present application are in the third quadrant;

[0055] Figure 8 A schematic diagram showing that the relative orientation angles between different interactive objects provided in a specific embodiment of the present application are in the fourth quadrant;

[0056] Figure 9 A schematic diagram of a graph convolution-based encoding and decoding network provided in a specific embodiment of the present application;

[0057] Figure 10 A flowchart of a method for predicting pedestrian crossing intentions provided in a specific embodiment of the present application;

[0058] Figure 11 A structural block diagram of a device for predicting pedestrian crossing intention provided in an embodiment of the present application.

[0059] Among them, 110 is a memory; 120 is a processor. DETAILED DESCRIPTION

[0060] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the specific implementation methods described herein are only used to illustrate and explain the embodiments of the present application and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0061] It should be noted that if the embodiments of the present application involve directional indications (such as up, down, left, right, front, back, etc.), the directional indications are only used to explain the relative position relationship, movement status, etc. between the various components under a certain specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0062] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present application, the descriptions of "first", "second", etc. are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the fact that they can be implemented by ordinary technicians in this field. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0063] Figure 1 This is a flow chart of a method for predicting pedestrian crossing intentions provided in an embodiment of the present application. Figure 1 As shown, an embodiment of the present application provides a method for predicting pedestrian crossing intention, which may include the following steps:

[0064] Step 101: Obtain video data of the current traffic scene;

[0065] Step 102: identifying multiple target objects in the video data;

[0066] Step 103: Determine a target pedestrian among the multiple target objects and multiple interaction objects corresponding to the target pedestrian;

[0067] Step 104: Determine feature sets of the target pedestrian and multiple interaction objects respectively;

[0068] Step 105: construct a pedestrian interaction graph model based on the feature sets of the target pedestrian and multiple interaction objects;

[0069] Step 106: Input the pedestrian interaction graph model into a graph convolution-based network to obtain a prediction result of the target pedestrian's intention to cross the street;

[0070] Among them, multiple interactive objects include mobile interactive objects and non-mobile interactive objects. The feature sets of mobile interactive objects and target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of non-mobile interactive objects include semantic features and visual features.

[0071] In an embodiment of the present application, in order to predict the intention of pedestrians to cross the street in the current traffic scene, video data of the current real traffic scene can be obtained first. The video data can be obtained by image acquisition methods such as vehicle-mounted cameras, and then the target objects in the video data are identified to obtain multiple target objects in the video. Among them, the target object refers to a common object in a traffic scene selected according to needs, including pedestrians, vehicles, traffic signs, and traffic lights. In one example, the video data can be input into a detection and tracking module based on a deep learning model to identify the target object of interest. In one example, each target object corresponds to multiple pixels in each frame of the video data, and the coordinates of the upper left and lower right pixel points of each target object represent the detection box of the target object, such as: Where t is the t-th frame image, i is the i-th object identified, is the coordinate of the upper left corner of the detection box of the i-th target object identified in the t-th frame, The coordinates of the upper right corner of the detection box of the i-th target object identified in the t-th frame. This facilitates subsequent calculations.

[0072] In an embodiment of the present application, a target pedestrian among multiple target objects and multiple interactive objects corresponding to the target pedestrian can be determined based on the detection frame of the target object. Among them, the target pedestrian refers to a pedestrian who may have the intention to cross the street in the current traffic scene, and the interactive object corresponding to the target pedestrian refers to an object that can affect the target pedestrian's intention to cross the street. One target pedestrian can correspond to multiple interactive objects. The multiple interactive objects can include mobile interactive objects and non-mobile interactive objects. Mobile interactive objects can include pedestrians, cars, cyclists, etc. Non-mobile interactive objects can include zebra crossings, traffic lights, traffic signs, pedestrian signals, etc.

[0073] Furthermore, the feature sets of the target pedestrian and each of its corresponding interactive objects are determined separately. In the embodiment of the present application, the feature sets of the target pedestrian and all its corresponding interactive objects include semantic features and visual features. The semantic features are feature vectors that characterize different types of target objects, and the visual features are latent features that contain visual information. In addition, since the position of the non-mobile interactive object in the video data does not change, that is, the corresponding detection frame information does not change, while the position of the mobile interactive object is variable, the target pedestrian is also a type of mobile interactive object. Therefore, the feature set of the target pedestrian and its corresponding mobile interactive object can also include spatiotemporal dynamic features. The spatiotemporal dynamic features refer to the temporal dependency between the historical trajectory data of different frames of the target pedestrian or mobile interactive object in the video data.

[0074] In an embodiment of the present application, after extracting feature sets of a target pedestrian and its multiple interaction objects, a pedestrian interaction graph model can be constructed based on the feature sets of the target pedestrian and its interaction objects, thereby obtaining a prediction result of the target pedestrian's crossing intention. In one example, an adjacency matrix representing the spatial relationship between the target pedestrian and its interaction objects can be obtained based on their spatial distribution and relative orientation. Then, based on the obtained interaction features and the adjacency matrix, a pedestrian interaction graph containing rich interaction features is constructed. Finally, the pedestrian interaction graph is input into a graph convolution-based network to obtain a pedestrian crossing intention prediction result.

[0075] Through the above technical solution, the video data of the current traffic scene is first obtained and multiple target objects in the video data are identified; then the target pedestrian and multiple interactive objects corresponding to the target pedestrian are determined among the multiple target objects, and the feature sets of the target pedestrian and the multiple interactive objects are determined respectively, and then a pedestrian interaction graph model is constructed based on the feature sets of the target pedestrian and the multiple interactive objects; finally, the pedestrian interaction graph model is input into a network based on graph convolution to obtain a prediction result of the target pedestrian's intention to cross the street. Among them, the multiple interactive objects include mobile interactive objects and non-mobile interactive objects, and the feature sets of mobile interactive objects and target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of non-mobile interactive objects include semantic features and visual features. This application models the target pedestrian and its corresponding interactive target, taking into account that the target pedestrian's intention to cross the street in the traffic scene will be affected by multiple interactive objects. Modeling is performed based on multiple dimensions of visual features, semantic features and spatiotemporal dynamic features, which can predict the pedestrian's intention to cross the street in complex traffic scenes, and the accuracy of the prediction results is high.

[0076] In the embodiment of the present application, step 103 of determining a target pedestrian among multiple target objects may include:

[0077] Filter out pedestrians from the target objects;

[0078] Determine the detection box area of the pedestrian;

[0079] When the detection box area is greater than or equal to the set area threshold, the pedestrian is determined as a target pedestrian.

[0080] Specifically, the set area threshold refers to the minimum value of the detection frame area of the target pedestrian. The controller can screen the target pedestrians according to the distance from the vehicle. Since the pixel area occupied by the pedestrian in the image is inversely proportional to the distance between the camera and the pedestrian, that is, the closer the distance between the pedestrian and the camera, the larger the area occupied by the pedestrian in the image, and vice versa. Therefore, the target pedestrian can be determined based on the detection frame area. After the detection and tracking module identifies all the target objects in the current traffic scene, the pedestrians among the target objects are screened out, and the detection frame area of each pedestrian is calculated separately. When the detection frame area of the pedestrian is greater than or equal to the set area threshold, the pedestrian is screened as the target pedestrian, otherwise it is ignored. In this way, the predicted target can be quickly determined.

[0081] In an embodiment of the present application, the interactive object corresponding to the target pedestrian is a target object among multiple target objects whose Euclidean distance to the target pedestrian is less than or equal to a set distance threshold.

[0082] Specifically, setting the distance threshold refers to the minimum value of the Euclidean distance between the interactive object and the target pedestrian. Since the target pedestrian's intention to cross the street is often affected by other objects in the same traffic scene, the corresponding interactive object can be determined based on the target pedestrian, and one target pedestrian often corresponds to multiple interactive objects. After determining the target pedestrian, the controller can respectively calculate the Euclidean distance between the pixel coordinates of the center point of the target pedestrian detection frame and the pixel coordinates of the center point of the detection frame of other target objects in the same scene. When the Euclidean distance is less than or equal to the set distance threshold, the target object is filtered as the interactive object of the target pedestrian, otherwise it is ignored. Similarly, the interactive targets of different target pedestrians in the scene are determined according to the above method. It can be understood that when there are multiple target pedestrians in a scene, the target pedestrians may also be each other's interactive objects. In this way, the interactive objects corresponding to the target pedestrian can be quickly screened out.

[0083] In an embodiment of the present application, in order to determine the semantic features of the target pedestrian and its interactive object, the common types of pedestrian interactive objects in the traffic scene can be numbered first. For example, the numbers of each object can be as follows: pedestrians are numbered 1, cars are numbered 2, cyclists are numbered 3, zebra crossings are numbered 4, traffic lights are numbered 5, pedestrian courtesy signs are numbered 7, and pedestrian signal lights are numbered 8. The semantic feature vectors representing different types of interactive targets are then initialized to 8-bit all-0 vectors, and the elements at the corresponding positions of the semantic feature vectors are set to 1 according to the type numbers of different targets to obtain the corresponding semantic feature vector sets. For example, the semantic feature vector of pedestrians is 10000000, and the semantic feature vector of cars is 01000000. In this way, the type of target object can be quickly determined through semantic features.

[0084] In this embodiment of the present application, step 104, determining the feature sets of the target pedestrian and the multiple interactive objects respectively, may include:

[0085] The video image regions corresponding to the target pedestrian and multiple interactive objects are input into the visual feature extraction module to obtain high-level invisible features corresponding to the target object and multiple interactive objects respectively;

[0086] The high-dimensional latent features of the target pedestrian and multiple interacting objects are converted into the same size through the spatial pyramid model to obtain the visual features of the target pedestrian and multiple interacting objects respectively.

[0087] Specifically, to determine the visual features of the target pedestrian and each of their corresponding interactive objects, visual features can be extracted from the video image region corresponding to the target pedestrian or their interactive objects. The controller can input the video image region into the visual feature extraction module to obtain high-dimensional latent features containing visual information. This visual information includes information such as the pedestrian's facial expression, traffic light color, and pedestrian gestures. The high-dimensional latent features are the feature matrices output by the corresponding visual feature extraction module.

[0088] In an example, the visual feature extraction module can be composed of two layers of two-dimensional convolutional neural networks and two layers of average pooling layers, wherein the convolution kernel size of the two layers of convolutional neural networks is 5×5 and the step size is 1; the pooling window size of the two layers of average pooling layers is 2×2 and the pooling step size is 1.

[0089] Furthermore, because the video image regions of the target pedestrian and the corresponding interactive objects have different sizes, the output sizes of the visual feature module will be inconsistent. To facilitate subsequent feature fusion, a spatial pyramid model can be used to convert the invisible features of the target pedestrian and its multiple interactive objects into the same size. This will yield the visual features of the target pedestrian and its multiple interactive objects.

[0090] In an example, the spatial pyramid model can be composed of three layers of spatial pooling layers. The first pooling layer adjusts the input latent features to a size of 5×5×dim_h, where dim_h is the number of spatial channels of the latent features. The second pooling layer adjusts the input latent features to a size of 3×3×dim_h, where dim_h is defined as above. The third pooling layer adjusts the input latent features to a size of 1×1×dim_h, where dim_h is defined as above. Finally, the outputs of the three pooling layers are concatenated into latent features of a size of 35×dim_h.

[0091] In this embodiment of the present application, the feature set for determining the target pedestrian and multiple interaction objects includes:

[0092] Filtering out mobile interactive objects from multiple interactive objects;

[0093] Obtain historical trajectory data of mobile interaction objects and target pedestrians respectively;

[0094] The historical trajectory data of the preset length of the mobile interactive object and the target pedestrian are respectively input into the gated recurrent unit to obtain the spatiotemporal dynamic characteristics of the target pedestrian and the mobile interactive object.

[0095] In the embodiment of the present application, for a mobile interactive object, its corresponding spatiotemporal dynamic features can be extracted. The target pedestrian can also be considered as a mobile interactive object. Furthermore, for the target pedestrian and its corresponding mobile interactive object, its center position information can be calculated based on its detection frame coordinates, and the center position information can be used as its trajectory information. Among them, the detection frame coordinates of the target pedestrian or mobile interactive object can be expressed as: Then its corresponding historical trajectory can satisfy formula (1):

[0096]

[0097] in, is the center coordinate of the i-th target in the t-th frame image.

[0098] Furthermore, the preset length historical trajectory data of the target pedestrian and each mobile interactive object are respectively input into their respective gated recurrent units (GRUs). The temporal dependency between the historical trajectory data of the target pedestrian and the interactive object in different frames is obtained through the gating mechanism in the gated recurrent unit, and the extracted temporal dependency is used as the spatiotemporal dynamic features of the target pedestrian and its interactive object. Among them, the preset length refers to the number of frames of the video data, and the preset length can be set according to actual needs. For example, the preset length can be 8 frames. In one example, the input dimension of each gated recurrent unit can be 2, and the hidden layer dimension can be 16.

[0099] In this way, by extracting the spatiotemporal dynamic features of the target pedestrian and its mobile interaction objects, it can provide a basis for subsequent modeling based on multi-source feature fusion interaction and improve the effectiveness of modeling.

[0100] In an embodiment of the present application, constructing a pedestrian interaction graph model based on a feature set of a target pedestrian and multiple interaction objects includes:

[0101] Based on the feature set, the interaction features of the target pedestrian and each interaction object are extracted respectively to obtain multiple interaction features;

[0102] Construct the target adjacency matrix of the pedestrian interaction graph model;

[0103] Multiple interaction features are used as node feature vectors of different target nodes, and a pedestrian interaction graph model is constructed based on graph theory and the target adjacency matrix.

[0104] In an embodiment of the present application, to construct a pedestrian interaction graph model, node features of the pedestrian interaction graph model can be first extracted, and a target adjacency matrix for the pedestrian interaction graph model can be constructed. The pedestrian interaction graph model can then be constructed based on the two. Specifically, multiple interaction features can be first obtained based on the feature sets of the target pedestrian and each interacting object. In one example, for both the target pedestrian and the mobile interacting object, since their feature sets include visual features, semantic features, and spatiotemporal dynamic features, the visual features and semantic features can be concatenated to obtain concatenated features. The concatenated features can then be fused with the spatiotemporal dynamic features using a self-attention mechanism to obtain interaction features. In another example, for non-mobile interacting objects, whose feature sets include only visual features and semantic features, the self-attention mechanism can be used to directly fuse their visual features with the semantic features to obtain interaction features. Furthermore, a target adjacency matrix for the pedestrian interaction model graph can be constructed. In one example, a target adjacency matrix can be constructed based on the spatial distribution and relative heading angles between different interacting objects. Finally, the interaction features are used as node feature vectors for different target nodes, and the pedestrian interaction graph model is constructed based on graph theory and the target adjacency matrix.

[0105] In an embodiment of the present application, when the interactive object is a mobile interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes:

[0106] Splicing the visual features and semantic features of the target pedestrian and the mobile interactive object to obtain the spliced features;

[0107] The concatenated features are fused with the spatiotemporal dynamic features through the self-attention mechanism to obtain interactive features.

[0108] In the embodiment of the present application, the feature sets of target pedestrians and mobile interactive objects each include visual features, semantic features, and spatiotemporal dynamic features. For target pedestrians and mobile interactive objects, the visual features and semantic features of each target pedestrian and interactive object can be first concatenated. That is, the visual features and semantic features of the target pedestrian are concatenated, and the visual features and semantic features of different mobile interactive objects are concatenated to obtain a concatenated feature. Then, a self-attention mechanism is used to update the concatenated features using the concatenated features and spatiotemporal dynamic features as input to obtain an interaction feature.

[0109] In an embodiment of the present application, when the interactive object is a non-moving interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes:

[0110] The visual features and semantic features of target pedestrians and non-moving objects are fused through the self-attention mechanism to obtain interactive features.

[0111] In an embodiment of the present application, the feature set of a non-moving interactive object includes visual features and semantic features. For a non-moving interactive object, the self-attention mechanism can be directly used to fuse its visual features with its semantic features to obtain interactive features.

[0112] Figure 2 This is a flow chart of a method for extracting interactive features provided in a specific embodiment of the present application. Figure 2 As shown, the interactive object is first determined based on the input video image detection and tracking results, and whether the interactive object is a mobile interactive object is judged. If the interactive object is non-mobile, the visual features of the non-mobile interactive object are obtained through a two-layer convolutional neural network with a convolution kernel size of 5×5, a two-layer average pooling layer with a convolution kernel size of 2×2, and a spatial pyramid pooling network (SPP-Net). The semantic features of the non-mobile interactive object are obtained by using the semantic feature vector library of the interactive object type name. The interactive features are then fused using the self-processing mechanism. If the interactive object is mobile, the same method is used to determine the visual and semantic features of the mobile interactive object. The spatiotemporal dynamic features of the mobile interactive object are obtained using the gated recurrent unit (GRU). The semantic features and visual features are then concatenated to obtain the concatenated features. The concatenated features are then fused with the spatiotemporal dynamic features using the self-processing mechanism to obtain the interactive features.

[0113] Figure 3 This is a flow chart of a method for fusion of multi-source features of mobile interactive objects provided in a specific embodiment of the present application. Figure 3 As shown in the figure, for the target pedestrian and mobile interaction object, the concatenated features are first mapped into vectors Q and K respectively through two layers of fully connected networks. Then, the spatiotemporal dynamic features are mapped into vector V through a single layer of fully connected networks. The elements in Q and K are multiplied one by one, and the attention value α is calculated using the Softmax function. Finally, the dot product operation of the attention value α and the vector V is performed to obtain the final interaction feature, which is used as the graph node feature in the pedestrian interaction graph model. Figure 4 This is a flow chart of a method for multi-source feature fusion of non-moving interactive objects provided in a specific embodiment of the present application. Figure 4 As shown in FIG, the fusion scheme of visual features and semantic features for non-moving interactive objects is the same as above.

[0114] In the embodiment of the present application, the target adjacency matrix for constructing the pedestrian interaction graph model includes:

[0115] constructing a first adjacency matrix according to the spatial distribution between the plurality of interacting objects;

[0116] constructing a second adjacency matrix according to the relative heading angles between the multiple interacting objects;

[0117] The first adjacency matrix and the second adjacency matrix are fused to obtain a target adjacency matrix.

[0118] In an embodiment of the present application, in order to effectively model the dynamic interaction behaviors between different interactive objects including the target pedestrian, the embodiment of the present application constructs the target adjacency matrix based on the spatial distribution and relative orientation angles between the interactive objects.

[0119] In an embodiment of the present application, a method for constructing a first adjacency matrix based on the spatial distribution between different interactive objects is as follows:

[0120] First, calculate the Euclidean distance between the center points of different interactive objects according to formula (2). For example, calculate the Euclidean distance d between the center points of interactive object i and interactive object j. ij :

[0121]

[0122] in, are the center coordinates of interaction object i and interaction object j at time t respectively.

[0123] Then, the first adjacency matrix A[i,j] is constructed according to formula (3) based on the Euclidean distance between the center points of different interaction objects. d :

[0124]

[0125] Where ψ1 is a constant value of 0.0001 to avoid i,j A[i,j] appears when it is 0 d Infinite value occurs.

[0126] In an embodiment of the present application, a method for constructing a second adjacency matrix based on the relative heading angles between different interactive objects is as follows:

[0127] Figure 5 A schematic diagram of a specific embodiment of the present application showing that the relative orientation angles between different interactive objects are in the first quadrant. Figure 6 A schematic diagram showing that the relative orientation angles between different interactive objects provided in a specific embodiment of the present application are in the second quadrant. Figure 7 A schematic diagram of a specific embodiment of the present application showing that the relative orientation angles between different interactive objects are in the third quadrant. Figure 8 A schematic diagram of the relative orientation angles between different interactive objects provided in a specific embodiment of the present application is in the fourth quadrant. Taking any two interactive objects i and j as an example, first connect the position coordinate points of interactive objects i and j at time t Get the straight line as the horizontal baseline and follow the Figure 5、 Figure 6 、 Figure 7 and Figure 8 The X-axis and Y-axis and the corresponding four quadrants are defined in the manner shown, where the four quadrants are marked as I, II, III and IV. Then, the position coordinates of the interactive objects i and j at time t+1 are used. Get the angles relative to the X axis, such as Figure 5 、 Figure 6 、 Figure 7 and Figure 8 As shown, the angle is the heading angle of the interactive object, recorded as

[0128] After obtaining the heading angles of any two interactive objects i and j, the heading angle difference between the two agents is calculated. Here, the heading angles of the interactive objects i and j are both in the first quadrant for specific explanation: Figure 5 As shown, When the heading angles of agents i and j are in other quadrants, the difference in heading angles satisfies formula (4):

[0129]

[0130] Calculate the difference in heading angles between the two interacting targets Then, the second adjacency matrix based on the heading angles of different interactive objects is constructed by formula (5):

[0131]

[0132] Where ψ2 is a constant value of 0.0001 to avoid Appears when 0 Infinite value occurs.

[0133] After obtaining the first adjacency matrix and the second adjacency matrix based on the Euclidean distance and relative heading angle between pedestrians and different interacting objects, the target adjacency matrix A[i,j] is obtained by fusion using formula (6).

[0134]

[0135] Among them, ω d , are two hyperparameters, and ω d =0.5,

[0136] Figure 9 A schematic diagram of a graph convolution-based encoding and decoding network provided in a specific embodiment of the present application. In a specific embodiment of the present application, after obtaining a pedestrian interaction graph model containing interaction information of a target pedestrian and his interaction objects, as shown in FIG. Figure 9 As shown, this is input into an encoder-decoder network consisting of three layers of graph convolution, a multilayer perceptron (MLP), and softmax to obtain a prediction result on whether the pedestrian has the intention to cross the street. The prediction result includes two outputs: the pedestrian has the intention to cross the street and the pedestrian does not have the intention to cross the street. In one example, the data dimension of the first layer of the three-layer graph convolution output is 512, the output dimension of the second layer is 256, and the output dimension of the third layer is 64. The multilayer perceptron consists of four fully connected layers, and the dimensions of the output features of each fully connected layer are 32, 16, 8, and 2 respectively. Finally, the output of the multilayer perceptron module is input into the softmax to obtain the final prediction result, where the softmax output data dimension is 2, respectively representing whether the pedestrian has the intention to cross the street.

[0137] Figure 10 This is a flow chart of a method for predicting pedestrians’ intention to cross the street, provided in a specific embodiment of the present application. Figure 10 As shown, a specific embodiment of the present application provides a method for predicting pedestrian crossing intention. The method may include: inputting a video image into a detection and tracking module, then processing the output data of the detection and tracking module through a visual feature extraction module, a spatiotemporal dynamic feature extraction module, and a semantic feature extraction module to obtain interaction features, thereby obtaining a node feature vector. Furthermore, based on the spatial distribution and relative orientation of pedestrians and their interaction objects, the Euclidean distance and relative orientation angle are calculated to obtain an adjacency matrix. A pedestrian interaction graph model is then constructed based on the node feature vectors and the adjacency matrix, and the pedestrian interaction graph model is input into a graph convolution model to ultimately obtain a pedestrian crossing intention prediction result.

[0138] In this way, compared with the modeling method that does not consider the differences in the categories of interacting objects, the above technical solution can consider the different impacts of inter-class differences of different interacting objects on pedestrian behavior intentions, thereby more reasonably modeling pedestrian interactions and providing a guarantee for improving the accuracy of pedestrian trajectory or intention prediction.

[0139] Figure 11 This is a structural block diagram of a device for predicting pedestrian crossing intentions provided in an embodiment of the present application. Figure 11 As shown, an embodiment of the present application provides a device for predicting pedestrian crossing intention, which may include:

[0140] Memory 110 configured to store instructions; and

[0141] The processor 120 is configured to call instructions from the memory 110 and implement the above-mentioned method for predicting pedestrian crossing intention when executing the instructions.

[0142] Specifically, in the embodiment of the present application, the processor 120 may be configured to:

[0143] Obtain video data of the current traffic scene;

[0144] Identify multiple target objects in video data;

[0145] determining a target pedestrian among the multiple target objects and multiple interaction objects corresponding to the target pedestrian;

[0146] Determine the feature sets of the target pedestrian and multiple interaction objects respectively;

[0147] Construct a pedestrian interaction graph model based on the feature sets of the target pedestrian and multiple interaction objects;

[0148] The pedestrian interaction graph model is input into a graph convolution-based network to obtain the prediction result of the target pedestrian's crossing intention;

[0149] Among them, multiple interactive objects include mobile interactive objects and non-mobile interactive objects. The feature sets of mobile interactive objects and target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of non-mobile interactive objects include semantic features and visual features.

[0150] Furthermore, the processor 120 may be configured to:

[0151] The video image regions corresponding to the target pedestrian and multiple interactive objects are input into the visual feature extraction module to obtain high-level invisible features corresponding to the target object and multiple interactive objects respectively;

[0152] The high-dimensional latent features of the target pedestrian and multiple interacting objects are converted into the same size through the spatial pyramid model to obtain the visual features of the target pedestrian and multiple interacting objects respectively.

[0153] Furthermore, the processor 120 may be configured to:

[0154] Filtering out mobile interactive objects from multiple interactive objects;

[0155] Obtain historical trajectory data of mobile interaction objects and target pedestrians respectively;

[0156] The historical trajectory data of the preset length of the mobile interactive object and the target pedestrian are respectively input into the gated recurrent unit to obtain the spatiotemporal dynamic characteristics of the target pedestrian and the mobile interactive object.

[0157] Furthermore, the processor 120 may be configured to:

[0158] Based on the feature set, the interaction features of the target pedestrian and each interaction object are extracted respectively to obtain multiple interaction features;

[0159] Construct the target adjacency matrix of the pedestrian interaction graph model;

[0160] Multiple interaction features are used as node feature vectors of different target nodes, and a pedestrian interaction graph model is constructed based on graph theory and the target adjacency matrix.

[0161] Furthermore, the processor 120 may be configured to:

[0162] When the interactive object is a mobile interactive object, based on the feature set, the interaction features of the target pedestrian and each interactive object are extracted separately, including:

[0163] Splicing the visual features and semantic features of the target pedestrian and the mobile interactive object to obtain the spliced features;

[0164] The concatenated features are fused with the spatiotemporal dynamic features through the self-attention mechanism to obtain interactive features.

[0165] Furthermore, the processor 120 may be configured to:

[0166] In the case where the interactive object is a non-moving interactive object, based on the feature set, the interaction features of the target pedestrian and each interactive object are extracted separately, including:

[0167] The visual features and semantic features of target pedestrians and non-moving objects are fused through the self-attention mechanism to obtain interactive features.

[0168] Furthermore, the processor 120 may be configured to:

[0169] constructing a first adjacency matrix according to the spatial distribution between the plurality of interacting objects;

[0170] constructing a second adjacency matrix according to the relative heading angles between the multiple interacting objects;

[0171] The first adjacency matrix and the second adjacency matrix are fused to obtain a target adjacency matrix.

[0172] Furthermore, the processor 120 may be configured to:

[0173] Filter out pedestrians from the target objects;

[0174] Determine the detection box area of the pedestrian;

[0175] When the detection box area is greater than or equal to the set area threshold, the pedestrian is determined as a target pedestrian.

[0176] In an embodiment of the present application, the interactive object corresponding to the target pedestrian is a target object among multiple target objects whose Euclidean distance to the target pedestrian is less than or equal to a set distance threshold.

[0177] Through the above technical solution, the video data of the current traffic scene is first obtained and multiple target objects in the video data are identified; then the target pedestrian and multiple interactive objects corresponding to the target pedestrian are determined among the multiple target objects, and the feature sets of the target pedestrian and the multiple interactive objects are determined respectively, and then a pedestrian interaction graph model is constructed based on the feature sets of the target pedestrian and the multiple interactive objects; finally, the pedestrian interaction graph model is input into a network based on graph convolution to obtain a prediction result of the target pedestrian's intention to cross the street. Among them, the multiple interactive objects include mobile interactive objects and non-mobile interactive objects, and the feature sets of mobile interactive objects and target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature sets of non-mobile interactive objects include semantic features and visual features. This application models the target pedestrian and its corresponding interactive target, taking into account that the target pedestrian's intention to cross the street in the traffic scene will be affected by multiple interactive objects. Modeling is performed based on multiple dimensions of visual features, semantic features and spatiotemporal dynamic features, which can predict the pedestrian's intention to cross the street in complex traffic scenes, and the accuracy of the prediction results is high.

[0178] An embodiment of the present application also provides a machine-readable storage medium having stored thereon instructions for enabling a machine to execute the above-mentioned method for predicting pedestrian crossing intentions.

[0179] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0180] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0181] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0182] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0183] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0184] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0185] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0186] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0187] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement or improvement made within the spirit and principles of the present application should be included within the scope of the claims of the present application.

Claims

1. A method for predicting pedestrian crossing intention, characterized in that: include: Obtain video data of the current traffic scene; identifying a plurality of target objects in the video data; determining a target pedestrian among the multiple target objects and multiple interactive objects corresponding to the target pedestrian; Determining feature sets of the target pedestrian and the multiple interactive objects respectively; Constructing a pedestrian interaction graph model based on the feature sets of the target pedestrian and the multiple interaction objects; Inputting the pedestrian interaction graph model into a graph convolution-based network to obtain a prediction result of the target pedestrian's intention to cross the street; Among them, the multiple interactive objects include mobile interactive objects and non-mobile interactive objects, the feature sets of the mobile interactive objects and the target pedestrians both include semantic features, spatiotemporal dynamic features and visual features, and the feature set of the non-mobile interactive objects includes the semantic features and the visual features.

2. The method for predicting pedestrian crossing intention according to claim 1, characterized in that: The feature set for determining the target pedestrian and the plurality of interactive objects includes: Inputting the video image regions corresponding to the target pedestrian and the multiple interactive objects into a visual feature extraction module to obtain high-level invisible features corresponding to the target object and the multiple interactive objects respectively; The high-dimensional latent features of the target pedestrian and the multiple interactive objects are converted into the same size through a spatial pyramid model to obtain the visual features of the target pedestrian and the multiple interactive objects respectively.

3. The method for predicting pedestrian crossing intention according to claim 1, characterized in that: The feature sets for respectively determining the target pedestrian and the plurality of interactive objects include: Filtering out a mobile interactive object from the plurality of interactive objects; respectively obtaining historical trajectory data of the mobile interactive object and the target pedestrian; The historical trajectory data of the preset length of the mobile interactive object and the target pedestrian are respectively input into the gated recurrent unit to obtain the spatiotemporal dynamic characteristics of the target pedestrian and the mobile interactive object.

4. The method for predicting pedestrian crossing intention according to claim 1, characterized in that: The constructing of a pedestrian interaction graph model according to the feature sets of the target pedestrian and the plurality of interaction objects comprises: Based on the feature set, respectively extracting interaction features of the target pedestrian and each interaction object to obtain a plurality of interaction features; Constructing a target adjacency matrix of the pedestrian interaction graph model; The multiple interaction features are respectively used as node feature vectors of different target nodes, and the pedestrian interaction graph model is constructed based on graph theory and the target adjacency matrix.

5. The method for predicting pedestrian crossing intention according to claim 4, characterized in that: In a case where the interactive object is the mobile interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes: splicing the visual features and semantic features of the target pedestrian and the mobile interactive object to obtain spliced features; The spliced features are fused with the spatiotemporal dynamic features through a self-attention mechanism to obtain the interactive features.

6. The method for predicting pedestrian crossing intention according to claim 4, characterized in that: In a case where the interactive object is a non-moving interactive object, extracting the interaction features of the target pedestrian and each interactive object based on the feature set includes: The visual features and semantic features of the target pedestrian and the non-moving object are fused through a self-attention mechanism to obtain the interaction feature.

7. The method for predicting pedestrian crossing intention according to claim 4, characterized in that: The target adjacency matrix for constructing the pedestrian interaction graph model includes: constructing a first adjacency matrix according to the spatial distribution between the plurality of interacting objects; constructing a second adjacency matrix according to the relative heading angles between the plurality of interactive objects; The first adjacency matrix and the second adjacency matrix are fused to obtain the target adjacency matrix.

8. The method for predicting pedestrian crossing intention according to claim 1, characterized in that: Determining a target pedestrian among the multiple target objects includes: Filtering out pedestrians from the target objects; Determining a detection frame area of the pedestrian; When the area of the detection frame is greater than or equal to a set area threshold, the pedestrian is determined as the target pedestrian.

9. The method for predicting pedestrian crossing intention according to claim 1, characterized in that: The interactive object corresponding to the target pedestrian is a target object among the multiple target objects whose Euclidean distance to the target pedestrian is less than or equal to a set distance threshold.

10. A device for predicting pedestrians' intention to cross the street, characterized in that: include: a memory configured to store instructions; as well as A processor is configured to call the instructions from the memory and implement the method for predicting a pedestrian's intention to cross the street according to any one of claims 1 to 9 when executing the instructions.

11. A machine-readable storage medium, characterized in that The machine-readable storage medium stores instructions for causing a machine to execute the method for predicting a pedestrian's intention to cross the street according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Street-crossing pedestrian group multi-modal trajectory prediction method for autonomous vehicle

    CN114898293A

  • Pedestrian crossing intention recognition method based on double-flow adaptive graph convolutional neural network

    CN116630873A