An intelligent driving planning method for pedestrian interaction scenarios
Through neural network model and interaction enhancement technology, the personified planning trajectory is generated imitated by human expert data, and the existing intelligent driving planning methods are solved, and the trajectory quality and low interactivity of existing intelligent driving planning methods in complex pedestrian interaction scenarios are achieved, and efficient and anthropomorphic intelligent driving path planning is achieved.
Patent Information
- Application Number
- CN202411431989.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-10-14
AI Technical Summary
The existing intelligent driving planning methods have poor trajectory quality, low interactivity and insufficient anthropomorphism in complex pedestrian interaction scenarios.
The neural network model is used to combine experience replay pooling and interaction enhancement technology to generate anthropomorphic planning trajectories by imitating human expert driving data, and to realize dynamic path planning of autonomous driving trolleys in pedestrian interaction scenarios through the autoregressive path planning model.
It improves the planning efficiency and trajectory quality in complex pedestrian interaction scenarios, can respond to pedestrian behavior changes in real time, and optimizes intelligent driving efficiency and comfort.
Smart Images

Figure CN119296077B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving, and particularly to an intelligent driving planning method for pedestrian interaction scenarios. Background Art
[0002] As an important part of future transportation, the development of autonomous vehicles not only indicates a significant improvement in traffic safety and traffic efficiency but also represents the rise of sustainable development and new business models. In this process, intelligent driving planning technology plays a crucial role, especially in complex high-interaction driving scenarios, and a typical case is the pedestrian interaction scenario in a park. Traditional rule-based methods, such as state machines and decision trees, although they can provide effective solutions in specific situations, have limited capabilities in dealing with large-scale state spaces and dynamically changing traffic environments.
[0003] Model Predictive Control (MPC), as an advanced control strategy, is increasingly applied to the intelligent driving planning of autonomous vehicles. MPC optimizes the driving trajectory in real time by establishing a dynamic model of the vehicle, taking into account the vehicle's dynamic constraints and real-time perceived environmental information, thereby achieving precise control of the vehicle's movement. However, traditional MPC has many defects in the path planning task of intelligent driving. For example, the computational complexity of MPC is high, resulting in poor real-time performance, especially in high-precision control problems; secondly, MPC requires an accurate system model to be designed for prediction and control, which means a high system modeling ability is needed. However, considering the uncertainty of pedestrian movement and the lack of high-precision maps in the real environment, it is almost impossible to accurately design the system model of the driving environment where the vehicle is located, ultimately resulting in low decision-making performance. Summary of the Invention
[0004] The present invention provides an intelligent driving planning method for pedestrian interaction scenarios to solve the technical problems of poor trajectory quality, low interactivity, and lack of anthropomorphism in existing intelligent driving planning methods in complex interaction scenarios.
[0005] To solve the above technical problems, the present invention provides the following technical solutions:
[0006] On the one hand, the present invention provides an intelligent driving planning method for pedestrian interaction scenarios, including:
[0007] S1, build a pedestrian interaction driving scenario in the park environment, and define vehicle sensor observations, vehicle perception states, objective functions, and vehicle underlying longitudinal and lateral controllers;
[0008] S2, Initialize the parameters of the neural network model and initialize the experience replay pool; among them, the neural network model is used to generate candidate planning trajectories; the experience replay pool is used to store the driving data sequences of human experts, and each driving data sequence consists of vehicle sensor observations, vehicle perception states, and objective function values at multiple moments;
[0009] S3, The human expert drives the vehicle in the pedestrian interaction driving scenario in the park environment, completes the vehicle navigation task under various pedestrian densities and various road conditions, generates the corresponding driving data sequences, and stores them in the experience replay pool;
[0010] S4, Use the driving data sequences in the experience replay pool to train the neural network model until the preset training termination condition is reached;
[0011] S5, Deploy the trained neural network model in the pedestrian interaction driving scenario in the park environment, plan the vehicle running trajectory, and use the vehicle's underlying longitudinal and lateral controllers to control the vehicle to run according to the planned trajectory;
[0012] S6, Determine whether the performance of the neural network model meets the standard. If there is a situation where the neural network model does not meet the standard in a certain scenario, collect more driving data sequences in the non-compliant scenarios through human experts, store them in the experience replay pool, and repeat S4 to S5 until the neural network model meets the standard in all scenarios;
[0013] S7, The training ends, and the last updated neural network model is obtained as the result of the training.
[0014] Furthermore, the construction of the pedestrian interaction driving scenario in the park environment includes:
[0015] Build a pedestrian interaction driving scenario in the park environment based on the Carla simulator. According to the actual park scenario, construct a corresponding OpenDRIVE format map, randomly generate multiple pedestrians with different shapes and behaviors, and randomly assign starting positions and target positions to these pedestrians, so that they walk along possible paths. When the pedestrians reach the predetermined target positions, assign new target positions to them, so that the generated pedestrians move in the park.
[0016] Furthermore, the vehicle sensor observations include the image information collected by the vehicle's surround cameras and the point cloud information collected by the vehicle's lidar; among them, the point cloud information collected by the vehicle's lidar is converted into a bird's-eye view with a fixed resolution, forming a square grid centered on the vehicle with a preset side length. The grid is divided into multiple grid blocks of the same size, and the height information of each grid block is quantified into a preset number of levels;
[0017] The vehicle perception state includes the vehicle motion state, the pedestrian motion state, and the lane line position state; among them, the vehicle motion state includes the global coordinates, speed, acceleration, orientation, width, and length of the vehicle; the pedestrian motion state includes the global coordinates, speed, acceleration, and orientation of the pedestrian; the lane line position state includes multiple two-dimensional coordinate points of each lane line.
[0018] The objective function is expressed as:
[0019] c = α1c safe + α2c speed + α3c lane + α4c comfort
[0020] Among them, c represents the objective function; α1, α2, α3, and α4 are preset weighting coefficients. Among them, N obs is the number of pedestrians, d safe is the preset safety distance, d i is the distance between the i-th pedestrian and the vehicle; c speed = β1v car + β2a car where, v car is the speed of the vehicle, a car is the acceleration of the vehicle, β1 and β2 are preset weight coefficients; c lane = -d center 2 where, d center is the distance between the vehicle and the center line of the road; c comfort = -(β3Δa car 2 + β4Δδ car 2 ) where, Δa car and Δδ car are the acceleration and the steering angle change rate of the vehicle respectively, and β3 and β4 are preset weight coefficients.
[0021] Furthermore, only the data within the preset range around the vehicle are retained for the vehicle motion state, the pedestrian motion state, and the lane line position state; and for the lane line position state within the preset range around the vehicle, all are retained. For the pedestrian motion state within the preset range around the vehicle, first sort the pedestrians according to the distance from the pedestrian to the vehicle, and then retain the motion states of the preset number of pedestrians closest to the vehicle. When the number is insufficient, it is filled with 0.
[0022] Furthermore, the vehicle bottom layer longitudinal and lateral controller is a PID controller.
[0023] Further, the neural network model parameters include a scene perception model, an interaction-enhanced environment model, and a path planning model; among them, the scene perception model is used to extract the vehicle perception state from vehicle sensor observations; the interaction-enhanced environment model is used to model the interaction relationship between the vehicle, pedestrians, and the ground Figure 3 and predict the future vehicle perception state and the objective function value; the path planning model is used to output the candidate planned trajectory of the vehicle; during training, first use the driving data sequence in the experience replay pool to train the scene perception model and the interaction-enhanced environment model in sequence, and then use the driving data sequence in the experience replay pool to train the path planning model based on the behavior cloning method, so that the trained path planning model can generate the vehicle planned trajectory for a future period of time in an autoregressive manner; among them, during the training process, the gradient descent method is used to update the parameters of the scene perception model, the interaction-enhanced environment model, and the path planning model.
[0024] Further, the scene perception model includes a camera branch network and a radar branch network; among them, the camera branch network is composed of a 4-layer convolutional neural network and a 1-layer fully connected neural network, and is used to extract feature information from the images collected by the vehicle's surround cameras to obtain camera features The radar branch network is composed of a 4-layer convolutional neural network and a 1-layer fully connected neural network, and is used to extract feature information from the point cloud data collected by the vehicle's lidar to obtain radar features For the obtained camera features and radar features Use a Transformer encoder and a 1-layer fully connected neural network to extract multi-modal fusion features Among them, the Transformer encoder is composed of multiple self-attention mechanisms, and its initial input is and The concatenation of, and the final output is the multi-modal fusion feature Multi-modal fusion feature Pass through multiple decoder heads to output the predicted values of the vehicle motion state The predicted values of the pedestrian motion state And the predicted values of the lane line position state Among them, each decoder head is composed of a 2-layer fully connected neural network.
[0025] Further, the input of the interaction-enhanced environment model is the vehicle motion state, the pedestrian motion state, and the lane line position state; for the vehicle motion state, first use a 2-layer fully connected neural network for encoding, and then perform self-attention mechanism fusion of temporal features at the temporal level to obtain the vehicle motion state temporal feature f car; For the pedestrian motion state, first encode it using a two-layer fully connected neural network, and then perform self-attention mechanism fusion on the temporal dimension to obtain the temporal feature of the pedestrian motion state f obs ; For the lane line position state, encode it using a two-layer fully connected neural network to obtain the encoded information f of the lane line position state line ;
[0026] After obtaining f car 、f obs and f line , the interaction-enhanced environment model first uses the self-attention mechanism to model the interaction between the vehicle and the pedestrian based on f car and f obs , and then uses the cross-attention mechanism to model the interaction between the vehicle and the map, and between the pedestrian and the map:
[0027]
[0028] where, I inter represents the interaction result between the vehicle and the pedestrian; I map represents the interaction result between the vehicle and the map, and between the pedestrian and the map; SA inter represents the self-attention mechanism at the level of traffic participants, that is, vehicles and pedestrians, and is used to model the interaction relationship between vehicles and pedestrians; Q, K, V represent the query, key value, and value vector in the attention mechanism; Concat(f car , f obs ) represents the concatenation result of f car and f obs ; CA line represents the cross-attention mechanism at the level of traffic participants and the road, and is used to model the attention of traffic participants to lane line elements;
[0029] After obtaining the interaction results between the vehicle and the pedestrian, the vehicle and the map, and the pedestrian and the map, the interaction-enhanced environment model decodes these interaction results to generate three outputs. The first output is the prediction of the vehicle motion state and the pedestrian motion state at the next moment. By predicting the vehicle motion state and the pedestrian motion state at the next moment and respectively supplementing the prediction results of the vehicle motion state and the pedestrian motion state into the existing vehicle operation state sequence and pedestrian operation state sequence, the interaction-enhanced environment model can autoregressively generate the future vehicle motion state sequence and pedestrian motion state sequence: Among them, the decoder for outputting the vehicle motion state prediction result and the decoder for outputting the pedestrian motion state prediction result are both composed of a two-layer fully connected neural network; the second output is the prediction of the future trajectory of the pedestrian; among them, the decoder for outputting the future trajectory prediction result of the pedestrian is composed of a two-layer fully connected neural network; the third output is the prediction of the objective function value; among them, the decoder for outputting the objective function value prediction result is composed of a two-layer fully connected neural network.
[0030] Further, the input of the path planning model is f car 、f obs and f line ;
[0031] The path planning model first obtains a query vector through f car , uses f obs as the key vector and value vector, and uses the cross-attention mechanism to obtain the interaction features between the vehicle and the pedestrian; then uses the interaction features between the vehicle and the pedestrian as the query vector, uses f line as the key vector and value vector, and uses the cross-attention mechanism to obtain the interaction features between the vehicle and the lane boundary; fuses the interaction features between the vehicle and the pedestrian and the interaction features between the vehicle and the lane boundary, and gives them to the planning decoder to output the final motion planning trajectory; among them, the planning decoder uses a gated recurrent unit to autoregressively generate the future planned trajectory points.
[0032] Further, the path planning model first generates multiple planning trajectories based on the multiple shooting method, then uses the interaction-enhanced environment model to simulate the multiple planning trajectories generated by the path planning model, and estimates the sum of the objective function values of each time step of each trajectory, and selects the trajectory with the largest sum of the objective function values of each time step; uses the selected trajectory as the actual running trajectory, and uses the vehicle bottom longitudinal and lateral controllers in the Carla simulator to control the vehicle to complete the trajectory.
[0033] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0034] On the other hand, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by a processor to implement the above method.
[0035] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0036] First of all, the environment model obtained by learning in the present invention can adapt to system changes and non-linearity, and the interaction enhancement technology further establishes the interaction relationship between the vehicle, pedestrians and the map, and models the driving scenario more accurately; secondly, by imitating and learning the trajectories of human experts, the path planning model in the present invention can provide a high-quality anthropomorphic initial solution for the trajectory, which helps to improve the efficiency and final quality of the planning, especially in complex driving environments such as pedestrian interaction. Finally, the present invention can realize the dynamic path planning of the autonomous vehicle in the pedestrian interaction scenario. This method can not only respond to the changes in pedestrian behavior in real time, but also optimize the intelligent driving efficiency and comfort in the pedestrian interaction scenario on the premise of ensuring safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 is a flowchart of an intelligent driving planning method for a pedestrian interaction scenario provided by an embodiment of the present invention;
[0039] Figure 2 is a schematic diagram of a training framework of an intelligent driving planning method for a pedestrian interaction scenario provided by an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of a pedestrian interaction driving scenario simulation in a park environment provided by an embodiment of the present invention;
[0041] Figure 4 is a schematic diagram of a neural network of an interaction-enhanced environment model provided by an embodiment of the present invention;
[0042] Figure 5 is a schematic diagram of a neural network of a trajectory planning model provided by an embodiment of the present invention;
[0043] Figure 6 is a system block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0045] First of all, it should be noted that in the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "exemplarily" is intended to present the concept in a concrete way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.
[0046] First embodiment
[0047] In view of the problems of poor trajectory quality, low interactivity and lack of anthropomorphism in existing intelligent driving planning methods in complex interactive scenarios, this embodiment follows the idea of MPC, introduces the learned environment model and path planning model, and implements an intelligent driving planning method for pedestrian interaction scenarios. Its system architecture includes: a campus pedestrian interaction scenario built based on the Carla simulator, which is used to support the training and verification of the algorithm, define the input and objective function required for the algorithm to run, and the underlying controller of the vehicle; an experience replay pool, which is used to store the experience collected by experts in the simulation scenario as training data for the algorithm; a parameter initializer, which is used to initialize the neural network model parameters of the algorithm; a scene perception model, which is used to process the sensor observations of the car and extract the features of the car itself and the surrounding environment (including road boundaries and pedestrians); an interactively enhanced environment model, which is used to model the interactive relationships between pedestrians, between pedestrians and cars, and between traffic participants including cars and pedestrians and road boundaries, and predict changes in the environment around the car and target values based on these interactive relationships; a path planning model, which is used to output candidate planning paths in the next few seconds based on the features obtained from the upstream model. During the operation of the algorithm, the parameter initializer is first used to initialize the network parameters of the scene perception model, the interactively enhanced environmental model, and the path planning model, and the experience replay pool is initialized. Then the following steps are repeated until the algorithm performance reaches the expected level: human experts drive the car in the park simulation scene, complete the car's navigation tasks under various pedestrian densities and road conditions, collect relevant driving data, and store it in the experience replay pool; use the data in the experience pool to perform multiple rounds of updates on the scene perception model, the interactively enhanced environmental model, and the path planning model; use the above three models in the simulation environment to complete the intelligent driving planning task in the pedestrian interaction scene, check the algorithm performance, and human experts collect data again for driving scenarios that do not meet the performance standards, replenish the experience replay pool, and conduct the next training until the performance in all scenarios meets the standards.
[0048] Specifically, the execution process of this method is as followsFigure 1 As shown in the figure, it includes the following steps:
[0049] S1. Build a pedestrian interaction driving scenario in the park environment, and define vehicle sensor observations, vehicle perception states, objective functions, and vehicle bottom-layer longitudinal and lateral controllers;
[0050] Among them, in this embodiment, a pedestrian interaction driving scenario in the park environment is built based on the Carla simulator, and a set of pedestrian generation and interaction mechanisms are created, effectively improving the authenticity of the simulation environment. Specifically, in this embodiment, according to the real road structure of a certain park, a map in OpenDRIVE format is made using Roadrunner. To simulate a real pedestrian interaction scenario, 100 pedestrians with different shapes and behaviors are generated in the map, and starting positions and target positions are randomly assigned to these pedestrians, making them walk along possible paths. When pedestrians reach the predetermined target positions, new target positions are assigned to them, so that the generated pedestrians move within the park. By introducing dynamic traffic flow and uncertain pedestrian behavior patterns, this simulation scenario can simulate different interaction situations, enhancing the authenticity and complexity of the scenario. Its visualization interface is as shown in Figure 3 the figure.
[0051] In addition, it should be noted that the position characteristics of the road boundary lines around the car and the movement characteristics of the car and pedestrians need to be further processed before being input to the downstream environment model and path planning model. Among them, for all features, first only the features within 40m around the ego vehicle are retained; for the road position features, all are retained. For the movement state features of pedestrians, they are first sorted according to the distance to the car, and then the 20 nearest pedestrian movement features are retained. If there are not enough, they are filled with 0.
[0052] Furthermore, the information included in the vehicle sensor observation o is: the vehicle surround camera o cam and the vehicle lidar o lidar ; among them, the input dimension of the surround camera is Among them, the number of surround cameras is 6, the number of image channels (RGB) of each camera is 3, and the resolution of a single camera image is 1280x720. The high-resolution input can provide rich visual information, which helps to perform accurate environmental perception and object recognition; the lidar point cloud is converted into a bird's-eye view with a fixed resolution. Considering the points 20 meters in front and behind and 20 meters on the left and right of the ego vehicle, a square grid with a side length of 40 meters is formed. The grid is divided into blocks of 0.125 meters by 0.125 meters, and the resolution of the bird's-eye view is 320x320. Further considering quantifying the height information of each grid block into 4 possible levels, the final input dimension of the lidar is Such processing enables the overhead view of the lidar to capture the relative height changes of the ground and different obstacles, enhancing the understanding of the three-dimensional structure of the surrounding environment and providing more accurate spatial information for the vehicle.
[0053] Further, the information included in the vehicle perception state s is: the car motion state s car , the pedestrian motion state s obs , and the lane line position state s line . The dimension of the car motion state is including the global coordinates of the car (x car , y car ), speed v car , acceleration a car , orientation θ car , and the width and length of the vehicle body (width, length); the dimension of the pedestrian motion state is where N obs is the number of pedestrians, usually set to 20. The motion state of a single pedestrian includes the global coordinates of the pedestrian (x obs , y obs ), speed v obs , acceleration a obs , and orientation θ obs ; the dimension of the lane line position state is where, N line is the number of lane lines, usually set to 200. A single lane line contains 20 two-dimensional coordinate points. For the pedestrian state, if there are more than 20 within the perception range, the 20 closest ones are selected according to the distance from the car. If the number is insufficient, it is filled with 0; for the lane line state, if the number is insufficient, it is also filled with 0. The above state information can directly obtain the true value through the simulator when collecting human expert data, but when testing the trained model, it is prohibited to use the simulator to obtain the true value, and only the predicted values of this information can be obtained through the environment perception model.
[0054] Further, the vehicle objective function c is calculated based on the vehicle perception state and control actions, mainly including the safety objective c safe , the efficiency objective c speed , the lane keeping objective c lane and the comfort objective c comfort and other 4 parts. Specifically, the safety objective c safe aims to prevent the car from colliding with pedestrians and keep away from pedestrians as much as possible, defined as where, N obs =20 is the number of pedestrians, d safe =1m is the preset safety distance, d i is the distance between the i-th pedestrian and the car; the efficiency objective c speedAim to optimize the driving path and speed of the vehicle to reduce the driving time, defined as c speed = β1v car + β2a car , where v car is the speed of the vehicle, a car is the acceleration of the vehicle, and β1 and β2 are weight coefficients, usually set to 1 and 0.2 respectively; the lane-keeping target c lane aims to ensure that the vehicle drives stably within the lane and avoid deviating from the center of the lane, which is achieved by minimizing the deviation between the vehicle and the center of the lane, defined as c lane = -d center 2 , where d center is the distance between the vehicle and the center line of the road; the comfort target c comfort aims to reduce the discomfort of the vehicle passengers, which is achieved by smoothing the changes in the acceleration and steering angle of the vehicle, defined as c comfort = -(β3Δa car 2 + β4Δδ car 2 ), where Δa car and Δδ car are the change rates of the acceleration and steering angle of the vehicle respectively, and β3 and β4 are weight coefficients, usually set to 0.1 and 0.1 respectively. Generally speaking, the objective function c is the weighted sum of 4 sub-objective functions: c = α1c safe + α2c speed + α3c lane + α4c comfort , where α1 to α4 are the weighted coefficients of each item, usually set to 1, 0.1, 0.5, and 0.1.
[0055] Furthermore, the vehicle's underlying lateral and longitudinal controller adopts a Proportional-Integral-Derivative (PID) controller. The vehicle's underlying lateral and longitudinal controller can generate two control actions: acceleration a and steering angle δ, which are used to track the future trajectory of the vehicle output by the path planning model.
[0056] S2, Initialize the neural network model parameters and initialize the experience replay pool;
[0057] Among them, the neural network model is used to implement tasks such as mapping and pedestrian detection, predict the state transition of the environment, and generate candidate planning trajectories; the experience replay pool is used to store the driving data sequences of human experts. Among them, the capacity of the experience replay pool is limited, and the storage rule is the first-in-first-out rule; a single expert data point at time t can be expressed as (o t , s t , c t), each data sequence can be represented by , where H represents the length of the data sequence, and o t is the sensor observation of the vehicle at time t, including the surround-view camera and lidar; s t is the state of the vehicle at time t, including the motion state characteristics such as the positions, orientations, and speeds of the vehicle and pedestrians, as well as the lane line position characteristics; c t is the objective function of the vehicle at time t, indicating the benefit of the vehicle's current driving.
[0058] Specifically, in this embodiment, the neural network model parameters include three parts: a scene perception model, an interaction-enhanced environment model, and a path planning model. Among them, the scene perception model is used to extract the vehicle perception state from the vehicle sensor observations; the interaction-enhanced environment model is used to model the interaction relationships among the vehicle, pedestrians, and the ground Figure 3 and predict the future vehicle perception state and objective function value; the path planning model is used to output the candidate planning trajectories of the vehicle; the interaction-enhanced environment model and the path planning model are composed of fully connected neural networks. Next, in combination with Figure 2 、 Figure 4 and Figure 5 each model will be introduced in detail:
[0059] The scene perception model is used to obtain the driving scene state around the vehicle from the surround-view camera and on-vehicle lidar. As Figure 2 shown, it uses two branches to extract features from the camera input and radar input respectively:
[0060]
[0061] Among them, Extracter cam and Extracter lidar represent the camera branch network and the radar branch network respectively. Extracter cam consists of 4 convolutional neural networks and 1 fully connected neural network, and extracts the visual features of 6 surround-view cameras respectively. Finally Similarly, Extracter lidar also consists of 4 convolutional neural networks and 1 fully connected neural network, and finally
[0062] After obtaining and a multi-modal fusion feature is extracted using a Transformer encoder and 1 fully connected neural network (MLP). and Concatenation, and finally output the multi-modal fusion features
[0063]
[0064] Multi-modal fusion features Through multiple decoder heads, output the predicted values of the vehicle motion state, pedestrian motion state, and lane line position state
[0065]
[0066] Among them, Decoder car , Decoder obs , Decoder line are the decoders for the vehicle, pedestrian, and lane line states respectively, and are all composed of 2-layer fully connected neural networks in this embodiment.
[0067] Furthermore, the interaction-enhanced environment model first further encodes the lane line position feature sequence and the motion feature sequences of the vehicle and the pedestrian extracted by the perception module respectively using an encoder; secondly, for the extracted features, multiple self-attention mechanisms and cross-attention mechanisms are introduced to model the temporal features and spatial interaction features; after fusing the above features, perform prediction tasks on all features at the next moment, perform trajectory prediction tasks on pedestrian-related features, and perform regression prediction tasks on each sub-objective in the objective function to realize the training of the model. Specifically, the structure of the interaction-enhanced environment model is as Figure 4 shown, and its input is the motion state sequences of the vehicle and the pedestrian in the expert data: where H is the sequence length, and the lane line position state For the state sequences of the vehicle, pedestrian and the state of the lane line, 2-layer fully connected neural networks (Multilayer Perceptron, MLP) are respectively used for encoding. For the encoded results of the vehicle and the pedestrian, additional self-attention mechanisms are used at the temporal level to fuse the temporal features:
[0068]
[0069] In the above formula, f car , f car , f car represent the encoded results of the vehicle, pedestrian, and lane line respectively, SA time represents the self-attention mechanism at the temporal level, and Q, K, V represent the query, key, and value vectors in the attention mechanism.
[0070] After obtaining the encoding result, the environment model first uses the self-attention mechanism to model the interaction between the vehicle and the pedestrian, and then uses the cross-attention mechanism to model the interaction between the vehicle and the map, as well as between the pedestrian and the map:
[0071]
[0072] In the above formula, I inter represents the interaction result between the vehicle and the pedestrian with each other, and I map represents the interaction result between the vehicle and the map, as well as between the pedestrian and the map. SA inter represents the traffic participants, that is, the self-attention mechanism at the level of the car and the pedestrian, and is used to model the interaction relationship between the car and the pedestrian with each other; CA line represents the cross-attention mechanism between the traffic participants and the road level, and is used to model the attention of the traffic participants to the lane line elements.
[0073] After obtaining the interaction result, the environment model decodes and generates three outputs according to these interaction results. The first output is the prediction of the motion feature at the next moment. By predicting the feature at the next moment and supplementing it to the existing state sequence, the environment model can autoregressively generate the future state sequence:
[0074]
[0075] In the above formula, Decoder pred_car and Decoder pred_obs are predictors of the car and pedestrian features respectively, and are each composed of a 2-layer fully connected neural network in this embodiment.
[0076] The second output is the prediction of the future trajectory of the pedestrian. By introducing the future trajectory prediction task of the pedestrian, the network supervision signal is enhanced, and the model has a certain post hoc interpretability:
[0077]
[0078] In the above formula represents the prediction of the future trajectory of the pedestrian, and its ground truth is extracted from the dataset; Decoder pred_traj is the trajectory predictor, and is composed of a 2-layer fully connected neural network in this embodiment.
[0079] The third output is the prediction of the objective function value. Considering that the objective function contains 4 sub-objectives, the model also predicts the values of the four sub-objective functions respectively:
[0080]
[0081] In the above formula Predictions for safety, driving efficiency, lane keeping, and comfort goals respectively; Decoder cost_safe , Decoder cost_speed , Decoder cost_lane , Decoder cost_comfort They are predictors for each sub-goal, and in this embodiment, they are all composed of a two-layer fully connected neural network.
[0082] The interaction-enhanced environment model can take into account the interactions between various elements in the driving scenario, better establish the dynamic relationship between the vehicle, pedestrians, and the road environment, enabling the autonomous driving system to more accurately predict and respond to changes in the surrounding environment, and thus simulate and evaluate complex traffic scenarios.
[0083] Furthermore, the input of the path planning model is the sequential encoding f of the car car , the sequential encoding f of the pedestrian obs , and the position encoding f of the road line . The path planning model first obtains a query vector from the current features of the car, uses the pedestrian features as the key and value vectors, and uses the cross-attention mechanism to obtain the interaction features between the car and the pedestrian; then uses the interaction features obtained in the previous step as the query vector, uses the lane line features as the key and value vectors, and uses the cross-attention mechanism to obtain the interaction features between the car and the lane boundary; fuses the above two interaction features and gives them to the planning decoder to output the final motion planning trajectory. The planning decoder uses a gated recurrent unit to generate future planned trajectory points in an autoregressive manner. Specifically, the structure of the path planning model is as shown in Figure 5 .
[0084] Among them, it should be noted that the planning of the car also needs to consider the interaction with pedestrians and map information, which are both implemented using the cross-attention mechanism in this embodiment:
[0085]
[0086] In the above formula, I car-obs is the interaction relationship between the car and the pedestrian, representing the car's attention to the pedestrian; I car-line is the interaction relationship between the car and the road, representing the car's attention to the lane line. CA car-obs and CA car-line are the cross-attention mechanisms between the car and the pedestrian and between the car and the road respectively.
[0087] After obtaining the above two interaction results, they are input into the planning decoder. In this example, the planning decoder is composed of a Gated Recurrent Unit (GRU). The working mode of the GRU at time t can be expressed as:
[0088]
[0089] In the above formula, h t represents the latent state at time t. At the initial time, that is, at time 0, the latent state is formed by concatenating I car-obs and I car-line . The latent state at subsequent times is generated by the GRU; f t+1 is the planning intermediate feature at time t + 1. Through this feature, a Gaussian distribution can be constructed and sampled to obtain the coordinate value offset Δw t+1 of the planning trajectory point at time t + 1. The final predicted planning trajectory can be expressed as:
[0090]
[0091] In the above formula, T is the step length of the vehicle planning trajectory, and the true value p car of the planning trajectory is extracted from the dataset.
[0092] S3. The human expert drives the vehicle in the pedestrian interaction driving scenario in the park environment, completes the vehicle navigation task under various pedestrian densities and various road conditions, generates the corresponding driving data sequence, and stores it in the experience replay pool;
[0093] S4. Use the driving data sequence in the experience replay pool to train the neural network model until the preset training termination condition is reached;
[0094] Specifically, in this embodiment, the training process of the model is as follows: Sample the expert data sequence samples from the experience replay pool, and train the scene perception model and the interaction-enhanced environment model, so that the scene perception model can obtain the state features of the vehicle, pedestrians, and lane lines from the sensor observations, and the environment model can model the vehicle-pedestrian interaction environment, predict the state changes, and estimate the objective function value. After the scene perception model is trained and can accurately obtain the state information from the sensor input, start training the interaction-enhanced environment model. Specifically: Use the expert dataset to train the path planning model based on the behavior cloning method, so that the path planning model can generate the vehicle planning trajectory for a future period of time in an autoregressive manner.
[0095] During the training process, the gradient descent method is used to update the parameters of the scene perception model, the interaction-enhanced environment model, and the path planning model, and it is judged in real time whether the preset training termination condition is reached, that is, whether the number of epochs of training iteration reaches the upper limit. If not, the training continues until the preset training termination condition is reached. Among them, the number of epochs of training iteration, usually also known as epoch, is set to 60 in this embodiment.
[0096] S5. Deploy the trained neural network model in the pedestrian interaction driving scenario in the park environment, plan the vehicle operation trajectory, and use the vehicle's underlying longitudinal and lateral controllers to control the vehicle to run according to the planned trajectory.
[0097] Further, the path planning model first generates multiple planned trajectories based on the multiple shooting method, then uses the interaction-enhanced environment model to simulate the multiple trajectories, and estimates the sum of the objective function values of each time step of each trajectory, and selects the trajectory with the largest sum of the objective function values of each time step; the selected trajectory is used as the actual running trajectory, and the underlying longitudinal and lateral PID controllers are used in the Carla simulator to control the vehicle to complete the trajectory.
[0098] Specifically, since the offset of the planned trajectory point coordinate values is sampled from a distribution, according to the multiple shooting method, multiple samplings of the planning model can obtain multiple candidate planned trajectories. In this embodiment, the number of sampled candidate planned trajectories is set to 5. When the future trajectory of the vehicle is determined, the future motion state characteristics of the vehicle are also determined accordingly. Based on the fixed historical motion characteristics of the vehicle and pedestrians, and the future motion characteristics of the vehicle, the environment model can predict the future motion characteristics of pedestrians in an autoregressive form and estimate the future objective function values. By calculating the objective function values of the 5 candidate trajectories, the trajectory with the largest sum is selected as the actual running trajectory.
[0099] S6. Judge whether the performance of the neural network model meets the standard. If there is a situation where the neural network model does not meet the standard in a certain scenario, more driving data sequences in the non-compliant scenario are collected by human experts, stored in the experience replay pool, and S4 to S5 are repeated until the neural network model meets the standard in all scenarios.
[0100] S7. The training ends, and the last updated neural network model is obtained as the result of the training.
[0101] The intelligent driving planning for the pedestrian interaction scenario can be realized by using the finally trained model.
[0102] In summary, this embodiment provides an intelligent driving planning method for pedestrian interaction scenarios. First, the learned environment model can adapt to system changes and non-linearity, and the interaction enhancement technology further establishes the interaction relationship among the vehicle, pedestrians, and the map, enabling more accurate modeling of driving scenarios. Second, by imitating and learning the trajectories of human experts, the path planning model can provide anthropomorphic high-quality initial trajectory solutions, which helps improve the efficiency and final quality of planning, especially in complex driving environments such as pedestrian interaction. Finally, the intelligent driving planning method of this embodiment can achieve dynamic path planning for autonomous vehicles in pedestrian interaction scenarios. This method can not only respond in real time to changes in pedestrian behavior but also optimize the intelligent driving efficiency and comfort in pedestrian interaction scenarios while ensuring safety.
[0103] Second Embodiment
[0104] This embodiment provides an electronic device, as Figure 6 shown. The electronic device includes: a processor and a memory; wherein, the processor and the memory can be connected through a communication bus; at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the above first embodiment. In addition, the electronic device may further include a transceiver, and the processor and the transceiver can be connected through a communication bus, and the transceiver is used for communicating with other devices.
[0105] Next, in combination with Figure 6 the following, a specific introduction to each component of the electronic device will be given:
[0106] Among them, the processor is the control center of the electronic device. The electronic device may include multiple processors, and each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here can be a single processor or a collective term for multiple processing elements. For example, the processor can be one or more central processing units (CPUs), or other general-purpose processors, application specific integrated circuits (ASICs), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0107] In a specific implementation, as an embodiment, the processor may include one or more CPUs. For example Figure 6 CPU0 and CPU1 shown in [figure reference], of course, this is only an exemplary illustration.
[0108] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.
[0109] Optionally, the memory may be a read-only memory (ROM) or other type of static storage device that can store static information and instructions, a random access memory (RAM) or other type of dynamic storage device that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory may be integrated with the processor or may exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 6 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.
[0110] The transceiver may include a receiver and a transmitter ( Figure 6 not shown separately in the figure). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. The transceiver may be integrated with the processor or may exist independently and be coupled to the processor through the interface circuit of the electronic device ( Figure 6 not shown in the figure), and the embodiments of the present invention do not make specific limitations in this regard.
[0111] In addition, it should be noted that Figure 6 the structure of the electronic device shown in the figure does not constitute a limitation on the device. The actual device may include more or fewer components than shown in the figure, or combine some components, or have a different component layout. In addition, the technical effects achieved by the electronic device when executing the method of the first embodiment above may refer to the technical effects described in the first embodiment above, so they will not be elaborated here.
[0112] Third Embodiment
[0113] This embodiment provides a computer-readable storage medium in which at least one instruction is stored. The instruction is loaded and executed by a processor to implement the method of the first embodiment above. Among them, the computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.
[0114] In addition, it should be noted that the present invention can be provided as a method, apparatus, or computer program product. Therefore, the embodiments of the present invention can take the form of all or part of a hardware embodiment, all or part of a software embodiment, or an embodiment combining software and hardware aspects. Moreover, when implemented using software, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center containing one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.
[0115] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified functions in Figure 1 one process or multiple processes and / or blocks Figure 1the functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes and / or boxes Figure 1 one process or more processes and / or boxes Figure 1 the steps of the functions specified in one box or more boxes.
[0117] It should also be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or terminal device comprising the said element. In addition, the term "and / or" is only a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. Among them, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context. "At least one" means one or more, and "a plurality" means two or more. "At least one of the following (items)" or similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b or c can mean: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0118] In addition, it can be understood that in various embodiments of the present invention, the magnitudes of the serial numbers of the above processes do not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0119] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or in combination with computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0120] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of functional modules / units is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. One can select some or all of the units according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in each embodiment of the present invention, the functional units can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0121] If the method is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0122] Finally, it should be noted that the above description is only a preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those of ordinary skill in the art, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle described in the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. An intelligent driving planning method for pedestrian interaction scenarios, characterized in that: include: S1, build a pedestrian interactive driving scenario in a campus environment, and define vehicle sensor observation, vehicle perception status, objective function, and vehicle underlying lateral and longitudinal controllers; S2, initialize the neural network model parameters and the experience replay pool; the neural network model is used to generate candidate planning trajectories; the experience replay pool is used to store driving data sequences of human experts, each of which consists of vehicle sensor observations, vehicle perception states, and objective function values at multiple moments; S3, human experts drive vehicles in the pedestrian interaction driving scenario in the campus environment, complete vehicle navigation tasks under various pedestrian densities and various road conditions, generate corresponding driving data sequences, and store them in the experience replay pool; S4, training the neural network model using the driving data sequence in the experience replay pool until a preset training termination condition is reached; S5, deploys the trained neural network model in the pedestrian interactive driving scenario in the campus environment, plans the vehicle's running trajectory, and uses the vehicle's underlying lateral and longitudinal controllers to control the vehicle to run according to the planned trajectory; S6, judging whether the performance of the neural network model meets the standard. If the performance of the neural network model does not meet the standard in a certain scenario, more driving data sequences in the scenario that does not meet the standard are collected by human experts, stored in the experience replay pool, and S4 to S5 are repeated until the performance of the neural network model meets the standard in all scenarios; S7, the training is finished, and the last updated neural network model is obtained as the training result; The vehicle sensor observation includes image information collected by the vehicle surround view camera and point cloud information collected by the vehicle laser radar; wherein the point cloud information collected by the vehicle laser radar is converted into a bird's-eye view with a fixed resolution, forming a square grid centered on the vehicle with a preset length as the side length, dividing the grid into a plurality of grid blocks of the same size, and quantizing the height information of each grid block into a preset number of levels; The vehicle perception state includes the vehicle motion state, the pedestrian motion state and the lane line position state; wherein the vehicle motion state includes the vehicle's global coordinates, speed, acceleration, direction and the width and length of the vehicle body; the pedestrian motion state includes the pedestrian's global coordinates, speed, acceleration and direction; the lane line position state includes multiple two-dimensional coordinate points of each lane line; The objective function is expressed as: c=α1c safe +α2c speed +α3c lane +α4c comfort Wherein, c represents the objective function; α1, α2, α3, α4 are preset weighting coefficients; Among them, N obs is the number of pedestrians, d safe is the preset safety distance, d i is the distance between the ith pedestrian and the vehicle; c speed =β1v car +β2a car , where v car is the speed of the vehicle, a car is the acceleration of the vehicle, β1 and β2 are the preset weight coefficients; c lane =-d center 2 , where d center is the distance between the vehicle and the center line of the road; c comfort =-(β3Δa car 2 +β4Δδ car 2 ), where Δa car and Δδ car are the acceleration and steering angle change rate of the vehicle, β3 and β4 are the preset weight coefficients; The neural network model parameters include a scene perception model, an interactively enhanced environment model, and a path planning model; wherein the scene perception model is used to extract the vehicle perception state from the vehicle sensor observation; the interactively enhanced environment model is used to model the interaction relationship between the vehicle, pedestrians, and the map, and predict the future vehicle perception state and objective function value; the path planning model is used to output the planning trajectory of the vehicle candidate; During training, the scene perception model and the interactively enhanced environment model are first trained in sequence using the driving data sequence in the experience replay pool, and then the path planning model is trained based on the behavior cloning method using the driving data sequence in the experience replay pool, so that the trained path planning model can generate the vehicle planning trajectory for a period of time in the future in an autoregressive manner; wherein, during the training process, the parameters of the scene perception model, the interactively enhanced environment model and the path planning model are updated using the gradient descent method.
2. The intelligent driving planning method for pedestrian interaction scenarios according to claim 1, characterized in that: The construction of a pedestrian interactive driving scene in a park environment includes: Based on the Carla simulator, a pedestrian interactive driving scenario in a campus environment is built. According to the actual campus scenario, the corresponding OpenDRIVE format map is constructed, and multiple pedestrians with different shapes and behaviors are randomly generated. The starting position and target position are randomly assigned to these pedestrians so that they walk along possible paths. When the pedestrians reach the predetermined target position, a new target position is assigned to them, so that the generated pedestrians move in the campus.
3. The intelligent driving planning method for pedestrian interaction scenarios according to claim 1, characterized in that: The vehicle motion state, pedestrian motion state and lane line position state only retain the data within the preset range around the vehicle; and all lane line position states within the preset range around the vehicle are retained. For the pedestrian motion state within the preset range around the vehicle, the pedestrians are first sorted according to the distance from the pedestrian to the vehicle, and then the motion states of a preset number of pedestrians closest to the vehicle are retained, and if the number is insufficient, it is padded with 0.
4. The intelligent driving planning method for pedestrian interaction scenarios according to claim 1, characterized in that: The vehicle bottom layer transverse and longitudinal controller is a PID controller.
5. The intelligent driving planning method for pedestrian interaction scenarios according to claim 1, characterized in that: The scene perception model includes a camera branch network and a radar branch network; wherein the camera branch network is composed of a 4-layer convolutional neural network and a 1-layer fully connected neural network, which is used to extract features from the image information collected by the vehicle surround view camera to obtain camera features. The radar branch network consists of a 4-layer convolutional neural network and a 1-layer fully connected neural network, which is used to extract features from the point cloud information collected by the vehicle laser radar to obtain radar features. For the obtained camera features and radar signature Use Transformer encoder and 1-layer fully connected neural network to extract multimodal fusion features The Transformer encoder is composed of a multi-layer self-attention mechanism, and its initial input is and The final output is the multimodal fusion feature Multimodal fusion features After passing through multiple decoder heads, the predicted value of the vehicle's motion state is output Prediction of pedestrian motion status And the predicted value of the lane line position state Among them, each decoder head consists of a 2-layer fully connected neural network.
6. The intelligent driving planning method for pedestrian interaction scenarios according to claim 5, characterized in that: The input of the interactive enhanced environment model is the vehicle motion state, pedestrian motion state and lane line position state. For the vehicle motion state, a two-layer fully connected neural network is first used for encoding, and then a self-attention mechanism is used to fuse the temporal features at the temporal level to obtain the vehicle motion state temporal feature f car ; For the pedestrian motion state, a two-layer fully connected neural network is first used for encoding, and then a self-attention mechanism is used to fuse the temporal features at the temporal level to obtain the pedestrian motion state temporal features f obs ; For the lane line position state, a 2-layer fully connected neural network is used for encoding to obtain the lane line position state encoding information fl line ; In getting f car 、f obs and f line After that, the interactively enhanced environment model first uses the self-attention mechanism based on f car and f obs , modeling the interaction between cars and people, and then using the cross-attention mechanism to model the interaction between cars and maps, and between pedestrians and maps: Among them, I inter Represents the interaction result between the car and the person; I map Represents the interaction between the car and the map, and between pedestrians and the map; SA inter represents traffic participants, that is, the self-attention mechanism at the vehicle and pedestrian level, which is used to model the interaction between vehicles and pedestrians; Q, K, V represent the query, key value and value vector in the attention mechanism; Concat(f car ,f obs ) indicates f car and f obs The splicing result of CA line A cross-attention mechanism representing the traffic participants and the road level is used to model the attention of traffic participants to lane line elements; After obtaining the interaction results between the car and the person, the car and the map, and the pedestrian and the map, the interactively enhanced environment model generates three outputs based on the decoding of these interaction results. The first output is the prediction of the vehicle motion state and the pedestrian motion state at the next moment. By predicting the vehicle motion state and the pedestrian motion state at the next moment and adding the predicted result of the vehicle motion state to the existing vehicle running state sequence, and adding the predicted result of the pedestrian motion state to the existing pedestrian running state sequence, the interactively enhanced environment model can autoregressively generate future vehicle motion state sequences and pedestrian motion state sequences: wherein the decoder for outputting the vehicle motion state prediction result and the decoder for outputting the pedestrian motion state prediction result are respectively composed of 2 layers of fully connected neural networks; the second output is the future trajectory prediction of pedestrians; wherein the decoder for outputting the future trajectory prediction result of pedestrians is composed of 2 layers of fully connected neural networks; the third output is the objective function value prediction; wherein the decoder for outputting the objective function value prediction result is composed of 2 layers of fully connected neural networks.
7. The intelligent driving planning method for pedestrian interaction scenarios according to claim 6, characterized in that: The input of the path planning model is f car 、f obs and f line ; The path planning model first passes through f car Get a query vector and convert f obs As the key vector and value vector, the cross attention mechanism is used to obtain the interaction features between vehicles and pedestrians; the interaction features between vehicles and pedestrians are then used as the query vector, and f line As key-value vectors and value vectors, the cross-attention mechanism is used to obtain the interaction features between the vehicle and the lane boundary; the interaction features between the vehicle and the pedestrian and the interaction features between the vehicle and the lane boundary are fused and handed over to the planning decoder to output the final motion planning trajectory; wherein the planning decoder uses a gated recurrent unit to generate future planned trajectory points in an autoregressive manner.
8. The intelligent driving planning method for pedestrian interaction scenarios according to claim 7, characterized in that: The path planning model first generates multiple planning trajectories based on the multiple targeting method, and then uses the interactive enhanced environment model to simulate the multiple planning trajectories generated by the path planning model, and estimates the sum of the objective function values of each trajectory at each time step, and selects the trajectory with the largest sum of the objective function values at each time step; the selected trajectory is used as the actual running trajectory, and the vehicle's underlying lateral and longitudinal controller is used in the Carla simulator to control the vehicle to complete the trajectory.
Citation Information
Patent Citations
Automatic driving vehicle and track planning and model training method, device and equipment
CN117746360A
Intelligent driving test method based on multi-background traffic participant interaction
CN117906973A