Riding track and intention prediction method and system, storage medium and equipment
Through the Transformer model combining vehicle position and visual sensing information, multi-task learning of cyclist trajectory and intention is achieved, solving the problem of neglecting interaction between bicycles and VRUs in existing VRU behavior predictions, and improving the accuracy of cyclist prediction and the safety of autonomous driving.
Patent Information
- Application Number
- CN202510305246.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-11
AI Technical Summary
The existing VRU behavior prediction methods are difficult to meet the cyclist's trajectory and intention prediction needs in autonomous driving scenarios, mainly due to the lack of sufficient input characteristics and the neglect of the interaction between the bicycle and the VRU, and trajectory prediction and intention prediction are regarded as independent problems.
Using the Transformer model, through self-attention and cross-attention mechanisms, combining vehicle position information and on-board visual sensing information, the interaction between the cyclist and bicycle is modeled, multi-task learning of cyclist trajectory and intention is achieved, and the future trajectory and intention of the cyclist are output.
It improves the accuracy of cyclist trajectory and intention prediction, adapts to different groups of cyclists, and enhances the collision risk warning capabilities of autonomous driving vehicles.
Smart Images

Figure CN120296339A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to a method, a system, a storage medium and a device for predicting a riding trajectory and an intention. Background Art
[0002] The statements in this part only provide background technical information related to the present invention, and do not necessarily constitute prior art.
[0003] During the driving of an autonomous vehicle, there are cyclists and pedestrians around. To ensure safety, generally, VRU (Vulnerable Road User) behavior prediction technology is used to predict the trajectories and intentions of pedestrians to help the autonomous vehicle avoid pedestrians, while less attention is paid to cyclists. And the current technology has the following problems:
[0004] 1. In the realistic traffic mixed scenario, the autonomous vehicle and its surrounding relevant traffic participants (including cyclists and pedestrians) form an interdependent whole, and their behaviors affect each other's decisions. In the standard VRU behavior prediction problem, the observer's perspective is generally a fixed perspective, only considering the interaction between VRU groups, without considering the interaction between the ego vehicle and VRUs.
[0005] 2. In terms of input information, predicting the motion trajectory of VRUs in the traffic environment requires using multiple information such as their positions, postures, motions and environmental factors. Most of the existing behavior prediction models only use the historical trajectory information of VRUs, resulting in insufficient input features of the model.
[0006] 3. There is an inherent connection between the trajectory and intention of a cyclist. Intention generates action, and the trajectory of the action reflects the intention. However, in the existing research on VRU behavior prediction, the trajectory prediction task and intention prediction are mostly separated and regarded as two independent problems to a large extent. Few methods consider the inherent connection between the intention of VRUs and their motion trajectories.
[0007] In summary, the above problems make the existing VRU behavior prediction methods difficult to meet the requirements of predicting the trajectories and intentions of cyclists in the autonomous driving scenario. Summary of the Invention
[0008] To solve the technical problems existing in the above background art, the present invention provides a method, system, storage medium and device for predicting cycling trajectories and intentions. The problem of predicting cyclists' trajectories and intentions is regarded as a multi-task learning problem, and Transformer is selected as the basis to construct a prediction model. Using vehicle pose information and on-vehicle visual sensing information as inputs, feature information that can fully reflect cyclists' intentions and behaviors is extracted from them. Then, in the encoding stage processed by Transformer, two self-attention modules are used to model the interactions between cyclists and the interactions between the observer (ego vehicle) and cyclists respectively; in the decoding stage processed by Transformer, the predicted trajectories and intentions of cyclists are output simultaneously.
[0009] To achieve the above object, the present invention adopts the following technical solutions:
[0010] The first aspect of the present invention provides a method for predicting cycling trajectories and intentions, including the following steps:
[0011] Obtain the position and heading of the observer at the current moment to form the pose at the current moment, and use the difference in pose between adjacent moments as the action a of the observer t ;
[0012] Obtain the image information of the cyclist, and after preprocessing, obtain the position, speed, orientation information of the cyclist and the information of the road where the cyclist is located at the same moment to form a state vector X t ;
[0013] Use a cyclist state encoder to model the interactions between cyclists in the current frame, input the state vector of cyclists at time t, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between cyclists x ;
[0014] Use a past trajectory encoder to model the interactions between the observer and the cyclist, input the action a of the current observer at time t t and the trajectory prediction value of the cyclist, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between the two y ;
[0015] G x and G y Use a cyclist trajectory / intention decoder to obtain the future trajectory Y t of the cyclist and the future intention H t+1 under the condition of the observer's action a t+1 .
[0016] As a further implementation, the action a of the observer t is the difference in pose between adjacent moments, that is, a t = p t – pt-1 = [ΔC x , ΔC y , Δθ] T , where the pose of the observer at time t is p t = [C x , C y , θ] T , (C x , C y ) is the center coordinate of the observer, and θ is the heading angle.
[0017] As a further implementation, the trajectory of the cyclist is , where u and v are the position coordinates of the cyclist in the world coordinate system, is its speed; the intention of the cyclist is defined as {turn left, go forward, turn right}
[0018] As a further implementation, in the cyclist state vector X t , each cyclist state is as follows: where u i , v i are the position coordinates of cyclist i in the world coordinate system, is the speed, that is, the coordinate difference between two consecutive frames, ori i is the orientation of cyclist i, and road i is the attribute of the road where cyclist i is located.
[0019] As a further implementation, the cyclist state encoder and self-attention mechanism in the Transformer model are used to model the interaction between cyclists in the current frame. The cyclist state vector at time t is input, and a d-dimensional intermediate embedding vector G x is output; as shown in the following formula: where represents the model of the encoder, are the learnable model parameters, and X t is the cyclist state vector.
[0020] As a further implementation, the past trajectory encoder and self-attention mechanism in the Transformer model are used to model the interaction between the observer and the cyclist. The action a t of the current observer at time t and the predicted value of the cyclist's trajectory are input, and a d-dimensional intermediate embedding vector G y is output; as shown in the following formula: G y = E φ (a t , Yt ; φ); where, E φ (·) represents an encoder model with learnable parameter φ, and Y t is the cyclist's trajectory.
[0021] As a further implementation, using the cyclist trajectory / intention decoder and cross-attention mechanism in the Transformer model, it is realized that under the condition of the observer's action a t , the state of the cyclist changes from Y t to Y t+1 , H t+1 The prediction is specifically:
[0022]
[0023] where, represents a decoder model with learnable parameter ψ, Y t+1 is the future trajectory, H t+1 is the future intention, G x is the intermediate embedding vector between cyclists, and G y is the intermediate embedding vector between the observer and the cyclist.
[0024] As a further implementation, by performing sine-cosine position encoding on the embedding vector R x to add the time position information in the sequential input.
[0025] The second aspect of the present invention provides a system required to implement the above method, including:
[0026] An observer background module, configured to: obtain the position and heading of the observer at the current moment, form the pose at the current moment, and use the difference between the poses at adjacent moments as the observer's action a t ;
[0027] A cyclist state module, configured to: obtain the image information of the cyclist, and after preprocessing, obtain the position, speed, orientation information of the cyclist and the information of the road where the cyclist is located at the same moment, and form a state vector X t ;
[0028] A cyclist state encoder, configured to: use the cyclist state encoder to model the interaction between cyclists in the current frame, input the cyclist state vector at time t, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between cyclists x ;
[0029] A past trajectory encoder, configured to: use the past trajectory encoder to model the interaction between the observer and the cyclist, input the action a t of the current observer at time t and the trajectory prediction value Y of the cyclistt Output a d-dimensional intermediate embedding vector G that reflects the interaction relationship between the two y ;
[0030] A trajectory / intention decoder, configured to: G x and G y Using the rider trajectory / intention decoder, obtain the future trajectory Y of the rider t under the condition of the observer's action a t+1 and the future intention H t+1 .
[0031] The third aspect of the present invention provides a computer-readable storage medium.
[0032] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the prediction method of the riding trajectory and intention as described above.
[0033] The fourth aspect of the present invention provides a computer device.
[0034] A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the steps in the prediction method of the riding trajectory and intention as described above.
[0035] Compared with the prior art, the above one or more technical solutions have the following beneficial effects:
[0036] 1. In an actual traffic scenario, traffic participants consist of an observer (an autonomous vehicle with environmental perception function) and an observed object (a group of riders). They drive towards their respective goals and interact with each other. During the process of the observer obtaining the rider's state, the observer's own actions are also changing. However, most traditional behavior prediction methods do not consider the observer's actions and simply consider the behavior prediction of the observed object. Therefore, the observer is considered to be in a fixed state, resulting in it being unsuitable for predicting riders in the context of autonomous driving. This solution takes the observer's action a t as a condition to predict the future trajectory Y t+1 and intention H t+1 of the group of riders, considers the correlation relationship between the rider's state X t and the observer's action a t , and both the observer and the observed object are in a moving state. Using pose information and visual information as inputs, multi-dimensional features that can fully reflect the intentions and behaviors of the riders are extracted, and through encoding and decoding processes, the predicted trajectories and intentions of the group of riders are obtained, which can be extended to other VRU behavior prediction problems to meet the trajectory and intention prediction requirements of VRUs.
[0037] 2. A method for predicting the behavior of a group of cyclists considering the actions of observers based on an attention mechanism, which is applicable to the collision risk warning of autonomous vehicles.
[0038] 3. It also has inherent adaptability to different numbers of inputs (i.e., the number of observed cyclists N), so it can predict the trajectories and intentions of groups of cyclists with varying numbers, improving its practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0040] Figure 1 It is a schematic diagram of the prediction model architecture for the cycling trajectories and intentions provided by one or more embodiments of the present invention;
[0041] Figure 2 It is a schematic diagram of the extraction of the orientation characteristics of cyclists provided by one or more embodiments of the present invention;
[0042] Figure 3 It is a schematic diagram of the process of extracting the road characteristics of cyclists provided by one or more embodiments of the present invention;
[0043] FIG. 4(a) is a schematic diagram of the architecture of the state encoder self-attention layer provided by one or more embodiments of the present invention;
[0044] FIG. 4(b) is a schematic diagram of the architecture of the past trajectory encoder self-attention layer provided by one or more embodiments of the present invention;
[0045] FIG. 4(c) is a schematic diagram of the architecture of the cyclist trajectory / intention decoder cross-attention layer provided by one or more embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0047] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0048] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0049] Vulnerable Road User (VRU) behavior prediction refers to predicting the behavior of vulnerable road users (such as pedestrians, cyclists, etc.) through technical means to improve traffic safety.
[0050] In a realistic traffic mixed scenario, an autonomous vehicle and its surrounding relevant traffic participants form an interdependent whole, and their respective behaviors affect each other's decisions. In the standard VRU behavior prediction problem, the observer's perspective is generally a fixed perspective, only considering the interaction between VRU groups and not considering the interaction between the ego vehicle and VRUs.
[0051] In terms of input information, predicting the motion trajectory of VRUs in a traffic environment requires using multiple information such as their position, attitude, motion, and environmental factors. Most existing trajectory prediction models only use the historical trajectory information of VRUs, resulting in insufficient input features for the model.
[0052] There is an inherent connection between the trajectory and intention of a cyclist. Intention generates action, and the trajectory of the action reflects the intention. However, in the existing research on VRU behavior prediction, the trajectory prediction task and intention prediction are mostly separated and are regarded as two independent problems to a large extent. Few methods consider the inherent connection between the intention of VRUs and their motion trajectories.
[0053] Therefore, the following embodiments provide a prediction method, system, storage medium, and device for cycling trajectories and intentions, regarding the problem of cyclist trajectory and intention prediction as a multi-task learning problem, and selecting Transformer as the basis to construct a prediction model. Using vehicle pose information and on-vehicle visual sensing information as input, extracting feature information that can fully reflect the intention and behavior of cyclists. Then, in the encoding stage processed by Transformer, two self-attention modules are used to model the interaction between cyclists and the interaction between the observer (ego vehicle) and cyclists respectively; in the decoding stage processed by Transformer, the predicted trajectory and intention of cyclists are output simultaneously.
[0054] In the following embodiments, "trajectory" specifically refers to "predicted trajectory" and is represented by Y; as for the "historical trajectory" in the general sense before time t, since it is included in the sequence of state X, it is collectively referred to as "state".
[0055] Example 1:
[0056] Prediction problem definition. On a road with mixed traffic, traffic participants consist of an observer (an autonomous vehicle with environmental perception function) and observed objects (visible groups of cyclists), which are moving towards their respective goals and interacting with each other. At time t, the observer observes the states of N cyclists through sensors such as on-vehicle cameras Track these N cyclists within a time window centered on t.
[0057] Different from other trajectory prediction algorithms centered on observed objects, in this problem, the observer is also an actor and interacts with other traffic participants (cyclists). Therefore, the goal of the method is to predict the future trajectories t and intentions of cyclists given a set of cyclist states X t and the action a of the observer and intentions . That is, conditional on the observer's action a t , predict the future trajectories Y t+1 and intentions H t+1 of the group of cyclists.
[0058] In this problem, it is assumed that the road surface is flat and the camera direction of the observer is parallel to it. The pose of the observer at time t is given by p t = [C x , C y , θ] T , where (C x , C y ) is the center coordinate of the observer and θ is the heading angle. The pose of the observer can be obtained by GPS+IMU or other sensors, such as vision-based methods (e.g., SLAM), laser point cloud-based methods (e.g., 3D-NDT), etc. The action of the observer is defined as the difference in poses at adjacent times: a t = p t – p t-1 = [ΔC x , ΔC y , Δθ] T .
[0059] The trajectory of a cyclist is defined as , where u, v are the position coordinates of the cyclist in the world coordinate system, is its speed; the intention of the cyclist is defined as {turn left, go straight, turn right}.
[0060] To solve the above problems, in this embodiment, Transformer is selected as the basis to build a prediction model because its self-attention mechanism can perform fast feature extraction on input information in parallel and is good at capturing the implicit interaction relationships in sequence data. In the standard attention mechanism, the input vector is first mapped into a query matrix Q, a key matrix K, and a value matrix V, and the attention is calculated as follows:
[0061]
[0062] where A is the attention matrix, and d k represents the dimension of the Q matrix. In the self-attention mechanism, Q and K, V are mapped from the same vector; in the cross-attention mechanism, and K, V are mapped from different vectors.
[0063] As Figure 1 shown, the prediction model includes an input module and a Transformer processing module. The input module extracts the features of the rider and constructs the rider state vector X t , which is used as the input of the Transformer Encoder; the Transformer processing module maps the implicit interaction relationships in the input data based on the attention mechanism to achieve the trajectory prediction and intention prediction of the rider group.
[0064] In the input module, data preprocessing is performed to obtain the state of each rider to represent their current movement and the background they are in. To fully reflect the intention and behavior characteristics of the rider, in addition to the position and speed information, this embodiment also uses the orientation information of the rider and the road information where they are located. Therefore, for each rider state x t in X i t is defined as:
[0065]
[0066] where u i , v i are the position coordinates of rider i in the world coordinate system, is its speed, defined as the coordinate difference between two consecutive frames, ori i is the orientation of rider i (Orientation), and the orientation sequence information of the rider can well reflect their potential movement intention. road i is the road attribute of rider i.
[0067] The state parameters can be obtained in the following way: The Bounding box of each rider is obtained through the existing object detection algorithm, and the center coordinates of the Bounding box are the pixel coordinates of the rider in the image coordinate system; Through the joint calibration of the lidar and the camera, the coordinate correspondence between the world coordinate system and the image coordinate system can be obtained, and then the pixel coordinates can be converted into world coordinates.
[0068] As Figure 2 shown, there are 8 types of rider orientations: {E, NE, N, NW, W, SW, S, SE}, where N represents the positive direction of the road. The orientation of the rider can be obtained by training an orientation classifier based on a CNN network model, such as the Faster R-CNN model combining ResNet50 and FPN.
[0069] road i is the road attribute of rider i. The road attribute has an important impact on the behavior of the rider. To obtain the road attribute of the location where the rider is, as Figure 3 shown, first perform semantic segmentation processing on the input image, assign semantic labels to each pixel, and obtain a pixel-level road traffic element segmentation (Road Segmentation) map. Then, according to the position of the center of the bottom of the Bounding box, query the semantic segmentation map to obtain the road attribute of the location where the rider is. Road semantic segmentation can use the MScaleOCR network model, which is trained on the KITTI-STEP dataset. During training, the road attribute is simply divided into 2 categories: intersection and road section. The output of the network is a probability value.
[0070] As Figure 1 shown, the Transformer processing module consists of a rider state encoder, a past trajectory encoder, and a rider trajectory / intention decoder, and their attention layer structure is shown in Figure 4.
[0071] (1) The rider state encoder uses the self-attention mechanism to model the interaction between riders in the current frame. The input is the rider state X t at time t, and the output is a d-dimensional intermediate embedding vector G x that reflects the interaction relationship between riders.
[0072]
[0073] Among them represents the model of the encoder, are learnable model parameters.
[0074] As Figure 1As shown in the figure, the rider status encoder consists of a multi-layer perceptron (MLP) and a multi-head self-attention layer. First, the input rider status vector is embedded into a d-dimensional feature vector by the MLP layer and used as the input to the self-attention layer. Then, based on the self-attention mechanism shown in Equation 1, the interactions between riders are encoded into an intermediate embedding vector
[0075] To address the problem of changes in the number of riders in the observer's field of view (such as riders joining or leaving the group or being blocked), an attention mask M is set for the self-attention matrix A of the feature vector R x as follows:
[0076] A′ = A eM (4)
[0077] where e represents the Hadamard product. The elements m in M ij are set as follows:
[0078]
[0079] Before inputting into the self-attention layer, the embedding vector R x is subjected to sinusoidal positional encoding to add temporal position information in the sequential input.
[0080] (2) Past trajectory encoder. In the standard trajectory prediction problem, the observer's perspective is generally fixed, while in this problem, both the observer and the observed are in motion, and the rider behavior to be predicted is closely intertwined with the motion (i.e., actions) of the observer itself. Therefore, the past trajectory encoder models the interaction between the observer and the rider based on the self-attention mechanism, taking the current action a of the observer at time t t and the predicted value Y of the rider's trajectory at the current moment t as inputs, and outputs a d-dimensional intermediate embedding vector G y reflecting the interaction relationship between the two.
[0081] G y = E φ (a t , Y t ; φ) (6)
[0082] where E φ (·) represents an encoder model with learnable parameters φ.
[0083] Similarly to the rider state encoder, this module first embeds the observer action and the predicted rider trajectory into d-dimensional vectors , which are used as the input to the self-attention layer. Then, based on the self-attention mechanism shown in Equation (1), the interaction between the observer and the rider trajectory is encoded into an intermediate vector
[0084] (3) Rider Trajectory / Intention Decoder. The rider trajectory / intention decoder is based on the cross-attention mechanism and realizes the prediction of the rider state from Y t to Y t to Y t+1 , H t+1 under the condition of the observer action a
[0085]
[0086] where represents the decoder model with learnable parameters ψ.
[0087] The rider trajectory / intention decoder consists of a cross-attention layer (Multi-Head Cross-Attention) and an FFN. The cross-attention layer integrates the rider group interaction embedding feature G x output by the state encoder and the observer-rider interaction embedding feature G y output by the past trajectory encoder, and outputs a vector containing the future trajectory embedding feature and the intention embedding feature In the attention calculation based on Equation (1), the key matrix K and the value matrix V are mapped from G x , and the query matrix Q is mapped from G y . The output R o is decoded by the FFN as the intention H t+1 of N riders and the trajectory Y t+1 .
[0088] Loss Function. In the training phase, the rider trajectory and intention prediction problems can be reduced to a multi-task learning problem. Corresponding loss functions are designed for the trajectory prediction task and the intention prediction task respectively, and then weighted summation is performed to combine the loss functions of each task into a total loss function.
[0089] For a group of N riders with a prediction duration of T, the loss function of the trajectory prediction task is defined as:
[0090]
[0091] where y is the predicted position of the cyclist, is the actual position of the cyclist.
[0092] The loss function for the intention prediction task selects the cross-entropy function, which is defined as:
[0093]
[0094] where H is the predicted intention distribution, is the true intention distribution, and K is the number of intention categories. Here, K = 3.
[0095] The total loss function is:
[0096]
[0097] where ω is the weight of each task.
[0098] Optionally, the prediction model parameters are set as follows:
[0099] Both the self-attention layer and the cross-attention layer use the multi-head mechanism, with the number of heads set to 8; the number of layers of the encoder and decoder is set to 3; the dimension d of the embedding vector is set to 32; the number of hidden neurons in the MLP is set to 16, and the Relu activation function and BatchNormalization are adopted.
[0100] For the sake of convenience of expression, only the single-step prediction from time t to time t+1 is described above; using the autoregressive mechanism, it is not difficult to generalize to the case where the historical trajectory duration M>1 and the prediction duration T>1:
[0101] ① At time t, for the input Y on the encoder side of the past trajectory t , let each The prediction model uses the historical state information {X t-M+1 , X t-M , …, X t} of the cyclist in the past M moments and the observer action information a t to predict the predicted trajectory Y t+1 at time t+1, and the predicted intention H t+1 ;
[0102] ② Output the predicted trajectory Y t+1 to the past trajectory encoder, let Y t = Y t+1 , and then combine R x and a t to predict the predicted trajectory Y t+2 at time t+2, and the predicted intention H t+2 ;
[0103] ③Iteratively perform such operations until the predicted trajectory Y of the rider at time t+T is obtained t+T and the predicted intention H t+T , and finally obtain the trajectory and intention of the rider over the future time length T.
[0104] In addition, the historical trajectory duration M and the prediction duration T need to be set by oneself.
[0105] The above solution only uses vehicle pose information and on-vehicle vision sensing information as inputs, extracts multi-dimensional features that can fully reflect the intentions and behaviors of riders, and obtains the predicted trajectories and intentions of the rider group through Transformer processing, which is of great significance for the collision risk warning of autonomous vehicles.
[0106] Most current VRU behavior prediction methods do not consider the actions of the observer and simply consider the behavior prediction of the observed object. The present invention provides a method for predicting the behavior of a rider group based on the attention mechanism while considering the actions of the observer, and takes into account the internal connection between the trajectory and the intention. This method can be extended to general VRU behavior prediction.
[0107] This method also has inherent adaptability to different numbers of inputs (i.e., the number of observed riders N), so it can predict the trajectories and intentions of rider groups with varying numbers, improving its practicality.
[0108] Embodiment 2:
[0109] A system for implementing the above method includes:
[0110] An observer background module, configured to: obtain the position and heading of the observer at the current moment, form the pose at the current moment, and use the difference in poses between adjacent moments as the action a of the observer t ;
[0111] A rider state module, configured to: obtain the image information of the rider, and after preprocessing, obtain the position, speed, orientation information of the rider and the information of the road where the rider is located at the same moment, and form a state vector X t ;
[0112] A rider state encoder, configured to: use the rider state encoder to model the interaction between riders in the current frame, input the rider state vector at time t, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between riders x ;
[0113] A past trajectory encoder, configured to: use the past trajectory encoder to model the interaction between the observer and the rider, input the action a of the current observer at time t tand the predicted values of the rider's trajectory, output a d-dimensional intermediate embedding vector G that reflects the interaction relationship between the two y ;
[0114] A trajectory / intention decoder, configured as: G x and G y Using the rider trajectory / intention decoder, obtain the future trajectory Y t of the rider and the future intention H t+1 under the condition of the observer's action a t+1 .
[0115] In an actual traffic scenario, traffic participants consist of an observer (an autonomous vehicle with environmental perception function) and the observed objects (a group of riders). They drive towards their respective goals and interact with each other. While the observer obtains the state of the rider, its own actions are still changing. Most traditional behavior prediction methods do not consider the observer's actions and simply consider the behavior prediction of the observed objects. Therefore, it is considered that the observer is in a fixed state, resulting in it being unsuitable for predicting the behavior of riders in the context of autonomous driving. In this embodiment, with the observer's action a t as the condition, predict the future trajectory Y t+1 and intention H t+1 of the group of riders, consider the correlation between the rider state X t and the observer's action a t , and both the observer and the observed objects are in a moving state. Using pose information and visual information as inputs, extract multivariate features that can fully reflect the intentions and behaviors of the riders. Through encoding and decoding processes, obtain the predicted trajectory and intention of the group of riders, which can be extended to other VRU behavior prediction problems to meet the trajectory and intention prediction requirements of VRUs.
[0116] Embodiment Three:
[0117] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in the method for predicting the riding trajectory and intention as described in Embodiment One above.
[0118] Embodiment Four:
[0119] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in the method for predicting the riding trajectory and intention as described in Embodiment One above.
[0120] The steps or modules involved in the second to fourth embodiments above correspond to those in the first embodiment. For specific implementation manners, reference may be made to the relevant description part of the first embodiment. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and cause the processor to execute any method in the present invention.
[0121] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for predicting cycling trajectories and intentions, characterized in that: including the following steps: Obtain the position and heading of the observer at the current moment to form the pose at the current moment, and use the difference in poses between adjacent moments as the action a of the observer t ; Obtain the image information of the cyclist. After preprocessing, obtain the position, speed, orientation information of the cyclist and the information of the road where the cyclist is located at the same moment, and form a state vector X t ; Using a rider status encoder, the rider status vector at time t is input, and a d-dimensional intermediate embedding vector G that reflects the interaction relationship between riders is output x ; Using the past trajectory encoder, input the action a of the current observer at time t t and the predicted trajectory value Y of the cyclist t , and output a d-dimensional intermediate embedding vector G that reflects the interaction relationship between the two y ; The obtained intermediate embedding vector G x and G y Using the rider trajectory / intention decoder, obtain the future trajectory Y t of the rider and the future intention H t+1 under the condition of the observer's action a t+1 .
2. The prediction method of riding trajectory and intention according to claim 1, characterized in that, The action a of the observer t is the pose difference between adjacent moments, i.e., a t = p t – p t-1 = [ΔC x , ΔC y , Δθ] T , where the pose of the observer at time t is p t = [C x , C y , θ] T , (C x , C y ) is the central coordinate of the observer, and θ is the heading angle.
3. The prediction method of riding trajectory and intention according to claim 1, characterized in that The trajectory of the cyclist is where u and v are the position coordinates of the cyclist in the world coordinate system, and is its speed; the intention of the cyclist is {turn left, go straight, turn right}.
4. The prediction method of riding trajectory and intention according to claim 1, characterized in that Rider state vector X t In which, each rider state is as follows: Among them, u i , v i are the position coordinates of rider i in the world coordinate system, is the speed, that is, the coordinate difference between two consecutive frames, ori i is the orientation of rider i, and road i is the attribute of the road where rider i is located.
5. The prediction method of riding trajectory and intention according to claim 1, characterized in that Using the rider state encoder and self-attention mechanism in the Transformer model, model the interaction between riders in the current frame. Input the rider state vector at time t and output a d-dimensional intermediate embedding vector G that reflects the interaction relationship between riders x ; as shown in the following formula: Among them, represents the model of the encoder, is the learnable model parameter, and X t is the rider state vector.
6. The prediction method of riding trajectory and intention according to claim 1, characterized in that Model the interaction between the observer and the cyclist using the past trajectory encoder and self-attention mechanism in the Transformer model, and input the action a of the current observer at time t t and the predicted trajectory value of the cyclist, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between the two y ; as shown in the following formula: G y = E φ (a t , Y t ; φ); where E φ (·) represents an encoder model with learnable parameters φ, and Y t is the cyclist's trajectory.
7. The method for predicting riding trajectory and intention according to claim 1, characterized in that When the number of cyclists in the observer's field of view changes, an attention mask M is set for the self-attention matrix A of the feature vector R x .
8. A prediction system for cycling trajectories and intentions, characterized in that, including: An observer background module is configured to: obtain the position and heading of the observer at the current moment, form the pose at the current moment, and use the difference between the poses at adjacent moments as the action a of the observer t ; The rider status module is configured to: obtain the image information of the rider, and after preprocessing, obtain the position, speed, orientation information of the rider and the information of the road where the rider is located at the same moment, and form a state vector X t ; A rider status encoder, configured to: utilize the rider status encoder to input a rider status vector at time t and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between riders x ; A past trajectory encoder, configured to: utilize the past trajectory encoder to input the action a of the current observer at time t t and the predicted value of the rider's trajectory, and output a d-dimensional intermediate embedding vector G reflecting the interaction relationship between the two y ; A trajectory / intention decoder, configured to: obtain an intermediate embedding vector G x and G y Using the rider trajectory / intention decoder, obtain the future trajectory Y t of the rider and the future intention H t+1 under the condition of the observer's action a t+1 .
9. A computer-readable storage medium having a computer program stored thereon, and when the program is executed by a processor, the steps in the method for predicting a riding trajectory and intention according to any one of claims 1-7 above are implemented.
10. A computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the method for predicting a riding trajectory and intention according to any one of claims 1-7 are implemented.