Bev feature space-based end-to-end automatic driving method and system
By adopting an end-to-end method based on BEV feature space in the autonomous driving system and combining with the deep Q network algorithm, the problem of poor performance of the existing technology in complex urban environments is solved, and autonomous driving with stronger perception and generalization capabilities is achieved.
Patent Information
- Application Number
- CN202510196413.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
Existing autonomous driving technologies do not perform well in complex urban environments, especially in dense traffic scenarios, and it is difficult to evaluate the movement of an agent through simple sparse reward signals, resulting in limited generalization capabilities.
The end-to-end autonomous driving method based on the BEV feature space is adopted, and the observation information of the driving route is collected in real time, and the surrounding viewing image is input to the pre-trained BEV visual perception module, a feature map is generated from the BEV perspective, and the vehicle status information is spliced and input into the depth Q network algorithm to make autonomous driving decisions.
Implement model-free vision-based autonomous driving in complex urban environments, with stronger perception and generalization capabilities, and can handle complex driving scenarios more effectively.
Smart Images

Figure CN120057036A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of autonomous driving, and relates to an end-to-end autonomous driving method and system based on a BEV feature space. Background Art
[0002] In the past few decades, autonomous driving technology has attracted increasing attention due to its great potential to change people's travel modes. However, this technology still faces many technical obstacles that must be overcome. Urban driving is one of the most challenging problems, mainly due to the complex urban environment and dynamic driving behaviors, such as surrounding vehicles and moving pedestrians, etc. Given a starting point and a target location, the goal of autonomous urban driving is to successfully complete the route within a limited time and meet the requirements of predefined conditions, such as no collisions, etc.
[0003] Traditional rule-based control methods are difficult to handle all edge cases and avoid violations during driving. In recent years, imitation learning has been widely applied to autonomous driving, but this method has problems of data bias and distribution shift, resulting in limited generalization ability of the agent. Deep reinforcement learning is another promising model-free control technology and has achieved remarkable success in various complex tasks, such as video games and robot control. However, existing autonomous driving frameworks based on deep reinforcement learning perform poorly in dense traffic scenarios because of the complexity of the urban environment and the difficulty for deep reinforcement learning agents to evaluate their actions generated in long-term and complex behaviors only through simple sparse reward signals. Summary of the Invention
[0004] The purpose of the present invention is to provide an end-to-end autonomous driving method and system based on a BEV feature space, which can achieve model-free vision-based autonomous driving in a complex urban environment and have stronger perception and generalization abilities.
[0005] To solve the above technical problems, the present invention is implemented by adopting the following technical solutions.
[0006] In a first aspect, the present invention proposes an end-to-end autonomous driving method based on a BEV feature space, including:
[0007] Real-time collecting observation information of the driving route, where the observation information includes panoramic images, vehicle state information, and environmental information;
[0008] Inputting the panoramic images into a pre-trained BEV visual perception module to obtain a feature map in the BEV view;
[0009] Concatenating the vehicle state information and the obtained feature map in the BEV view to obtain a feature z for autonomous driving decision-making;
[0010] Input the feature z into a pre-trained deep Q-network algorithm to obtain the decision of continuous actions of the autonomous vehicle;
[0011] The autonomous vehicle executes corresponding actions according to the obtained decision, thereby realizing autonomous driving.
[0012] Combined with the first aspect, further, the vehicle state information includes the position, direction, speed, and acceleration of the autonomous vehicle; the environmental information includes waypoint information.
[0013] Combined with the first aspect, further, the inputting the surround-view image into a pre-trained BEV visual perception module to obtain a feature map from the BEV perspective includes:
[0014] The BEV visual perception module includes an image encoding network and a BEV perception feature space;
[0015] Input the surround-view image into the image encoding network for extracting image features to obtain an image feature map;
[0016] Use the information of the image feature map, combined with convolution, sampling, and pooling operations, to construct a BEV perception feature space, and further obtain a feature map from the BEV perspective.
[0017] Combined with the first aspect, further, the inputting the surround-view image into the image encoding network for extracting image features to obtain image feature information includes:
[0018] The image encoding network includes an FPN and a ResNet as the backbone network; input the surround-view image into the ResNet to obtain feature maps of different scales, and input the obtained feature maps of different scales into the FPN for fusion to obtain a fused feature map. Specifically: perform upsampling and downsampling operations on the feature maps of different scales through the FPN so that the feature maps of different scales have the same spatial size, and then perform horizontal connection and fusion operations to obtain a fused image feature map.
[0019] Combined with the first aspect, further, the training method of the BEV visual perception module includes:
[0020] Collect the observation information of the driving route for training;
[0021] Input the surround-view image for training into the image encoding network to obtain an image feature map for training;
[0022] Use the information of the image feature map for training, combined with convolution, sampling, and pooling operations, to construct a BEV perception feature space for training, and further obtain a feature map from the BEV perspective for training;
[0023] Concatenate the vehicle state information for training and the feature map in the BEV view for training, further process the features through a fully connected layer, and obtain the feature z for autonomous driving decision-making in training;
[0024] Construct a perception task branch and a behavior cloning task branch, and obtain the total loss to train the BEV visual perception model.
[0025] Combined with the first aspect, further, the perception task branch includes an object detection task and a semantic segmentation task;
[0026] Input the feature map in the BEV view for training into the SSD detection head to obtain the result of object detection; input the feature map in the BEV view for training into the ASPP dilated convolution module to obtain the result of semantic segmentation.
[0027] Combined with the first aspect, further, use a multi-layer perceptron to complete the regression value prediction tasks for steering, throttle, and brake;
[0028] Calculate the losses of each branch; use 、 to represent the object detection task and the semantic segmentation task of the feature map in the BEV view respectively, and construct a cross-entropy loss to calculate the losses of the object detection task and the semantic segmentation task:
[0029] ;
[0030] ;
[0031] Among them, is the loss of the object detection task, is the loss of the semantic segmentation task of the feature map in the BEV view; 、 represent the ground truth values of object detection and semantic segmentation of the feature map in the BEV view respectively; 、 represent the number of objects in the object detection task and the semantic segmentation task of the image respectively has no practical meaning, representing the sum starting from the first item, and the total number of items is 、 ; represents the input spatial features in the BEV view;
[0032] Use 、 、 to represent three regression tasks, namely the steering regression task, the throttle regression task, and the brake regression task, and construct a mean squared error loss MSE to calculate the losses of each regression task:
[0033]
[0034]
[0035]
[0036] Among them, 、 、 represent the losses of the steering regression task, throttle regression task, and brake regression task respectively, 、 、 represent the true values of steering, throttle, and brake respectively, and z represents the input decision high-dimensional feature;
[0037] The total loss combines the losses of the above five tasks, and trains the BEV visual perception model by performing multi-task learning and minimizing the total loss :
[0038] ;
[0039] Among them, 、 、 、 are the weights of the object detection task loss, semantic segmentation task loss, steering regression task loss, throttle regression task loss, and brake regression task loss respectively.
[0040] Combined with the first aspect, further, the observation information for training also includes the true values of the object detection task, semantic segmentation task, brake, throttle, and steering, and these data can be directly obtained through the autonomous driving simulator.
[0041] Combined with the first aspect, further, the method for obtaining the observation information of the driving route for training is to collect the observation information on the driving route from the autonomous driving simulator.
[0042] Combined with the first aspect, further, using the information of the image features, combined with convolution, sampling, and pooling operations, a perceptual BEV feature space is constructed, and then a feature map in the BEV view is obtained, including:
[0043] Based on the obtained image feature map, using the depth estimation algorithm of LSS, the obtained image feature map is dimensionally increased to construct a frustum, and the image depth is predicted;
[0044] Generate a three-dimensional point cloud using the image feature map, depth information map, and camera parameter information, and the implementation formula is as follows:
[0045] ;
[0046] Among them, Represent image coordinates; Represent point cloud coordinates; Both are matrices used to transform the reference frame of point cloud coordinates, Used to transform the reference frame of the point cloud to the image, Used to correct the transformed reference frame; Is the camera internal parameter, through Realize the transformation from point cloud coordinates to image coordinates;
[0047] Pool the 3D point cloud along the vertical direction to obtain the preliminary BEV space features;
[0048] Based on the obtained preliminary BEV space features, use the ResNet network and FPN to fuse features of different scales, and finally obtain the optimized BEV space features, that is, obtain the feature map from the BEV perspective.
[0049] Combined with the first aspect, further, the spliced vehicle state information and the feature map from the BEV perspective are further processed through a fully connected layer to obtain the feature z for training autonomous driving decisions, including:
[0050] Flatten the optimized BEV space features, splice them with the vehicle state information to form a new decision high-dimensional feature;
[0051] Input the decision high-dimensional feature into the fully connected layer to obtain the high-dimensional feature z of autonomous driving decisions.
[0052] Combined with the first aspect, further, input the feature z into the pre-trained deep Q-network algorithm to obtain the decisions of continuous actions of the autonomous driving vehicle, including:
[0053] Discretize the continuous action space , that is , , Respectively represent steering, throttle, and brake;
[0054] For the steering wheel space; discretize it uniformly from -1 to 1 into 33 intervals, -1 corresponds to turning completely to the left, and 1 corresponds to turning completely to the right; for the throttle and brake spaces, combine them and discretize them into 3 different actions: accelerating (throttle is 0.6, brake is 0), moving forward (throttle is 0, brake is 0), and decelerating (throttle is 0, brake is 1);
[0055] Shape the reward;
[0056] Reward formulation consists of two aspects: 1) When a predefined event is triggered, the agent receives a sparse reward; 2) At each timestamp, the agent receives a dense reward. For the sparse reward, three adverse events that incur penalty rewards are defined: 1) Collision with a static object; 2) Collision with a vehicle or pedestrian; 3) Route deviation. In addition, a positive event that gives a reward is defined: Successful completion of the route.
[0057] Among them, the dense reward includes a deviation degree reward , a deviation distance reward , and a speed reward ; The calculation method of the deviation degree reward is as follows:
[0058] ;
[0059] Among them, represents the route deviation degree, is the maximum threshold set to 90°;
[0060] The calculation method of the deviation distance reward is as follows:
[0061] ;
[0062] Among them, represents the route deviation distance, represents the maximum threshold of the route deviation distance;
[0063] The calculation method of the speed reward is as follows:
[0064] ;
[0065] Among them, when no dynamic objects (i.e., vehicles and pedestrians) are detected, according to traffic rules, and are set as the minimum and maximum recommended speeds, is defined as the target speed, and its calculation formula is ;
[0066] Initialize the online Q-network and the target Q-network, create an experience replay buffer, and given the decision-making high-dimensional features from the BEV perception module at the moment, use them as the initial state of the DQN agent at this moment and store them in the experience replay buffer;
[0067] According to the policy, select an action, randomly select an action for exploration with a probability, and select the action with the maximum Q value with a probability. The formula is as follows:
[0068] 。
[0069] Combined with the first aspect, further, the training method of the deep Q-network algorithm is as follows: According to the feature z provided by the BEV perception module, the deep Q-network outputs steering, throttle, and brake values, and the autonomous driving vehicle performs corresponding actions based on this output, interacts with the environment of the autonomous driving simulator (i.e., interacts with the environmental information in the observation information for training). Subsequently, the environmental state changes, and the state information at this moment is stored in the experience replay buffer. Then, a new round of training is carried out until the optimal deep Q-network algorithm is obtained.
[0070] It should be noted that: in the actual use process (i.e., after the BEV perception module and the deep Q-network algorithm are trained), after the feature z is input into the trained deep Q-network algorithm, the deep Q-network outputs action decisions of steering, throttle, and brake values. The deep Q-network outputs steering, throttle, and brake values, and the autonomous driving vehicle performs corresponding actions based on this output and interacts with the environmental information in the real-time collected observation information.
[0071] Combined with the first aspect, furthermore, the training method of the deep Q-network algorithm specifically includes:
[0072] Based on the action instruction, the agent executes and obtains a new state and reward, and at the same time stores the transition tuple (state, action, reward, new state) into the experience replay buffer;
[0073] Randomly sample a batch of samples from the experience buffer to train the online Q-network and the target Q-network, and at the same time calculate the difference between the target Q value and the predicted Q value, and update the network parameters. The Q value calculation formula is:
[0074]
[0075] where, represents the Q value of the action under the state , represents the reward, represents the discount factor, represents the next state, represents the next action;
[0076] After a certain number of steps, copy the parameters of the online Q-network to the target Q-network to keep the target network stable;
[0077] Enter the action selection decision again and start iterating until the established goal is reached, that is, until the vehicle reaches the set end point.
[0078] Second aspect, the present invention proposes an end-to-end autonomous driving system based on a BEV feature space, including:
[0079] An observation information acquisition module, configured to collect observation information of the driving route in real time, where the observation information includes panoramic images, vehicle state information, and environmental information;
[0080] A BEV visual perception module, configured to input the panoramic image into a pre-trained BEV visual perception module to obtain a feature map from the BEV perspective; splice the vehicle state information and the obtained feature map from the BEV perspective to obtain a feature z for autonomous driving decision-making;
[0081] A reinforcement learning module, configured to input the feature z into a pre-trained deep Q-network algorithm to obtain a decision on the continuous actions of the autonomous driving vehicle; the autonomous driving vehicle executes corresponding actions according to the obtained decision, thereby realizing autonomous driving.
[0082] In combination with the second aspect, further, the BEV visual perception module includes an image encoding network module and a BEV perception feature space module; the image encoding network module is configured to input the panoramic image into an image encoding network for extracting image features to obtain an image feature map; the BEV perception feature space module is configured to use the information of the image feature map, combine convolutional, sampling, and pooling operations to construct a BEV perception feature space, and further obtain a feature map from the BEV perspective.
[0083] Third aspect, the present invention proposes a computer-readable storage medium, on which a computer program is stored, characterized in that: when the computer program is executed by a processor, the steps of the above-mentioned end-to-end autonomous driving method based on a BEV feature space are realized.
[0084] Fourth aspect, the present invention proposes a computer device, characterized in that it includes:
[0085] A memory for storing a computer program;
[0086] A processor for executing the computer program to realize the steps of the above-mentioned end-to-end autonomous driving method based on a BEV feature space.
[0087] Fifth aspect, the present invention proposes a computer program product, including a computer program, characterized in that: when the computer program is executed by a processor, the steps of the above-mentioned end-to-end autonomous driving method based on a BEV feature space are realized.
[0088] Compared with the prior art, the beneficial effects achieved by the present invention:
[0089] (1)The method of the present invention is based on the Bird's Eye View (BEV) feature space and is used to achieve vision-based autonomous urban driving. It can realize model-free vision-based autonomous driving in complex urban environments and has stronger perception and generalization capabilities.
[0090] (2)The system of the present invention includes an image encoding network module, a BEV perception feature space module, and a reinforcement learning module. The image encoding network module is responsible for extracting key features of the surround-view images; the BEV perception feature space module uses these features to transform the driving scene into a feature map from the BEV perspective; in order to comprehensively construct the driving scene information, we introduce BEV multi-task branches (i.e., the perception task branch and the behavior cloning task branch) to complete various vision tasks; based on the BEV features, combined with the current scene and vehicle state information, a high-dimensional autonomous driving decision representation is further formed; in order to ensure that the system of the present invention can make reasonable decisions in a dynamic traffic environment, we use the Deep Q-Network (DQN) algorithm to realize the combination of environmental perception and decision control, and construct an end-to-end autonomous driving system, providing a new solution for urban driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0091] Figure 1 It is a framework flowchart of the end-to-end autonomous driving method in Embodiment 1;
[0092] Figure 2 It is a structural diagram of the BEV vision perception module in the end-to-end autonomous driving method in Embodiment 1;
[0093] Figure 3 It is a structural diagram of the image encoding module and the BEV feature space construction module in the BEV vision perception module in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0094] The technical solutions of the present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0095] The term "and / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " generally represents an "or" relationship between the associated objects before and after.
[0096] Embodiment 1
[0097] As Figure 1As shown in the figure, an end-to-end autonomous driving method based on the BEV feature space in this embodiment includes the following steps:
[0098] Step 1: Collect diverse observation information of the training route from the autonomous driving simulator, including panoramic images, vehicle state information, and environmental information;
[0099] Step 2: Input the panoramic image into the network composed of ResNet and FPN to extract image features;
[0100] Step 3: Use the image feature information, combined with convolution, sampling, and pooling operations, to construct a BEV feature space;
[0101] Step 4: Concatenate the vehicle state information and the feature map in the BEV view, and further process the features through a fully connected layer to form a high-dimensional representation for autonomous driving decision-making;
[0102] Step 5: Based on Steps 2 to 3, construct a perception task branch and a behavior cloning branch. The perception task branch includes object detection and semantic segmentation; the behavior cloning branch includes the prediction of action steering values, throttle values, and brake values;
[0103] Step 6: Pre-train the BEV visual perception module according to the methods in Steps 2 to 5, and provide the obtained feature z to the DQN (Deep Q-Network) agent to let it make decisions for the continuous actions of the autonomous driving vehicle;
[0104] Step 7: The vehicle executes corresponding actions according to the formulated decision content, interacts with the autonomous driving simulator environment, and enables the training and optimization of the DQN model through the information collected by the experience buffer, thereby realizing autonomous driving.
[0105] In a specific implementation manner of this embodiment, Step 1 specifically includes the following steps:
[0106] Step 1-1: Arrange 6 RGB cameras on the autonomous driving vehicle of the autonomous driving simulator and select 25 training routes;
[0107] Step 1-2: According to the above settings, collect a large amount of diverse data sets in 25 routes, including vehicle state information (including position, steering, speed, acceleration), environmental information (waypoint information), and the ground truth of perception tasks and behavior cloning tasks.
[0108] In a specific implementation manner of this embodiment, Step 2 specifically includes the following steps, and the picture encoding method is as shown in the upper half of Figure 3 as follows:
[0109] Step 2-1: Use the backbone network ResNet and FPN to output feature maps of 6 panoramic images at different resolutions (1 / 4, 1 / 8, 1 / 16).
[0110] Step 2-2: Upsample the feature maps of different resolutions to 1 / 4 of the original image size, and then splice these feature maps to obtain the fused feature maps. This method is used to obtain feature maps for images at different angles.
[0111] In a specific implementation manner of this embodiment, the specific steps of step three are as follows. The BEV feature space construction method is as Figure 3 shown in the lower part of
[0112] Step 3-1: Based on the feature maps in step two, use the depth estimation algorithm of LSS to dimensionally elevate the image to construct a frustum and predict the image depth.
[0113] Step 3-2: Generate a three-dimensional point cloud using the image feature map, depth information map, and camera parameter information. The implementation formula is as follows:
[0114]
[0115] Among them, represents the image coordinates; represents the point cloud coordinates; are both matrices used to convert the reference system of the point cloud coordinates, is used to convert the reference system of the point cloud to the image, is used to correct the converted reference system; is the camera internal parameter, and the conversion from point cloud coordinates to image coordinates is achieved through
[0116] Step 3-3: Pool in the vertical direction to obtain the preliminary BEV space features.
[0117] Step 3-4: Based on the BEV space features in the previous step, use the ResNet network and FPN to fuse features of different scales, and finally obtain the optimized BEV space features.
[0118] In a specific implementation manner of this embodiment, the specific steps of step four are as follows. The high-dimensional feature extraction method is as Figure 2 shown in the structure diagram
[0119] Step 4-1: Flatten the optimized BEV space features and splice them with vehicle state information (position, heading, speed, acceleration) to form a new decision-making high-dimensional feature.
[0120] Step 4-2: Input the decision-making high-dimensional feature into the fully connected layer to form a high-dimensional representation for autonomous driving decision-making.
[0121] In a specific implementation manner of this embodiment, step five specifically includes the following steps:
[0122] Step 5-1: Construct an SSD detection head to perform object detection tasks from the BEV perspective;
[0123] Step 5-2: Construct an ASPP dilated convolution module to perform semantic segmentation tasks on the feature map from the BEV perspective;
[0124] Step 5-3: Construct a multi-layer perceptron to complete the regression tasks of steering, throttle, and brake;
[0125] Step 5-4: Calculate the losses of each branch. Use and to represent the two branches of the object detection task and the semantic segmentation task of the feature map from the BEV perspective respectively, and construct a cross-entropy loss to calculate the losses of the two tasks:
[0126]
[0127]
[0128] Among them, is the loss of the object detection task, is the loss of the semantic segmentation task of the image; and represent the ground truths of object detection and semantic segmentation of the image respectively; and represent the number of objects in the object detection task and the semantic segmentation task of the image respectively has no practical meaning and represents the sum starting from the first item, and the total number of items is and respectively; represents the input spatial features from the BEV perspective.
[0129] Use and and to represent the three branches of the regression task, namely the steering regression task, the throttle regression task, and the brake regression task, and construct a mean squared error loss (MSE) to calculate the loss of the regression task:
[0130]
[0131]
[0132]
[0133] Among them, and and Represent the losses of three branches: steering regression task, throttle regression task, and brake regression task. , , Represent the true values of steering, throttle, and brake, and z represents the input high-dimensional decision features.
[0134] The total loss combines the losses of the above five tasks, and trains the BEV visual perception model by performing multi-task learning and minimizing the total loss: To train the BEV visual perception model:
[0135]
[0136] Among them, , , , Are the weights of the object detection task loss, semantic segmentation task loss, steering regression task loss, throttle regression task loss, and brake regression task loss, respectively.
[0137] In a specific implementation manner of this embodiment, step six specifically includes the following steps:
[0138] Step 6-1, discretize the continuous action space , that is, steering, throttle, and brake.
[0139] For the steering wheel space, it is evenly discretized from -1 to 1 into 33 intervals, -1 corresponding to a full left turn and 1 corresponding to a full right turn; for the throttle and brake spaces, they are combined and discretized into 3 different actions: accelerating (throttle is 0.6, brake is 0), moving forward (throttle is 0, brake is 0), and decelerating (throttle is 0, brake is 1);
[0140] Step 6-2, shape the rewards. Reward formulation includes two aspects: 1) When a predetermined event is triggered, the agent obtains a sparse reward; 2) At each timestamp, the agent obtains a dense reward. For the sparse reward, three adverse events that give penalty rewards are defined: 1) Collision with a static object; 2) Collision with a vehicle or pedestrian; 3) Route deviation. In addition, a positive event that gives a reward is defined: successfully completing the route.
[0141] Among them, the dense reward consists of the deviation degree reward , the deviation distance reward , and the speed reward .
[0142] The calculation method of the deviation degree reward is as follows:
[0143]
[0144] Among them, represents the degree of route deviation, is the maximum threshold set to 90°.
[0145] Deviation distance reward is calculated as follows:
[0146]
[0147] Among them, represents the route deviation distance, represents the maximum threshold of the route deviation distance, preferably meters.
[0148] Speed reward is calculated as follows:
[0149]
[0150] Among them, when no dynamic objects (i.e., vehicles and pedestrians) are detected, according to traffic rules, set and as the minimum and maximum recommended speeds, is defined as the target speed, and its calculation formula is .
[0151] Step 6-3, Initialize the online Q-network and the target Q-network, create an experience replay buffer, and given the decision-making high-dimensional feature z from the BEV perception module at time, use it as the initial state of the DQN agent at this time and store it in the buffer;
[0152] Step 6-4 Select an action according to the policy, randomly select an action to explore with probability, and select the action with the largest Q value with , the formula is as follows:
[0153]
[0154] In a specific implementation manner in this embodiment, the step seven specifically includes the following steps:
[0155] Step 7-1, Based on the action instruction, the DQN agent executes and obtains a new state and reward, and at the same time stores the transition tuple (state, action, reward, new state) in the experience replay buffer;
[0156] Step 7-2, Randomly extract a batch of samples from the experience buffer to train the online Q-network and the target Q-network, and at the same time calculate the difference between the target Q value and the predicted Q value, and update the network parameters. The Q value calculation formula is:
[0157]
[0158] Among them, state the Q value of action a in the represents the reward, represents the discount factor, represents the next state, represents the next action.
[0159] Step 7-3, at a certain number of steps apart, copy the parameters of the online Q network to the target Q network to keep the target network stable;
[0160] Step 7-4, enter the action selection decision again and start iterating until the established goal (the vehicle reaches the set end point) is reached.
[0161] Embodiment 2
[0162] Based on the same inventive concept as Embodiment 1, this embodiment introduces an end-to-end autonomous driving system based on the BEV feature space, including:
[0163] An observation information acquisition module, configured to collect observation information of the driving route in real time, and the observation information includes panoramic images, vehicle state information, and environmental information;
[0164] A BEV visual perception module, configured to input the panoramic image into the pre-trained BEV visual perception module to obtain a feature map from the BEV perspective; splice the vehicle state information and the obtained feature map from the BEV perspective to obtain the feature z for autonomous driving decision-making;
[0165] A reinforcement learning module, configured to input the feature z into the pre-trained deep Q network algorithm to obtain the decision of the continuous actions of the autonomous driving vehicle; the autonomous driving vehicle executes corresponding actions according to the obtained decision, thereby realizing autonomous driving.
[0166] Embodiment 3
[0167] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned end-to-end autonomous driving method based on the BEV feature space are implemented.
[0168] Embodiment 4
[0169] Based on the same inventive concept as other embodiments, this embodiment introduces a computer device, including: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the above-mentioned end-to-end autonomous driving method based on the BEV feature space.
[0170] Example 5
[0171] Based on the same inventive concept as other embodiments, this embodiment introduces a computer program product, including a computer program, which when executed by a processor, implements the steps of the above-described end-to-end autonomous driving method based on the BEV feature space.
[0172] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0173] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0176] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms, and these all fall within the protection scope of the present invention.
Claims
1. An end-to-end autonomous driving method based on bev feature space, characterized in that: include: Collecting observation information of the driving route in real time, the observation information includes surround view images, vehicle status information and environmental information; Input the surround image into the pre-trained BEV visual perception module to obtain the feature map from the BEV perspective; The vehicle status information and the feature map obtained from the BEV perspective are combined to obtain the feature z used for autonomous driving decision-making; Input the feature z into a pre-trained deep Q-network algorithm to obtain a decision on the continuous action of the autonomous driving vehicle; The autonomous vehicle performs corresponding actions based on the decisions made, thereby achieving autonomous driving.
2. The end-to-end autonomous driving method based on bev feature space according to claim 1, characterized in that: The surround view image is input into the pre-trained BEV visual perception module to obtain a feature map under the BEV perspective, including: The BEV visual perception module includes an image encoding network and a BEV perception feature space; Input the surround view image into the image encoding network to obtain the image feature map; The information of the image feature map is used in combination with convolution, sampling and pooling operations to construct a BEV perception feature space, thereby obtaining a feature map from the BEV perspective.
3. The end-to-end autonomous driving method based on bev feature space according to claim 2, characterized in that: The training method of the BEV visual perception module includes: Collect observations of driving routes used for training; Inputting the surround view image for training into the image encoding network to obtain the image feature map for training; Using the information of the image feature map used for training, combined with convolution, sampling and pooling operations, a BEV perception feature space for training is constructed, and then a feature map from the BEV perspective for training is obtained; The vehicle state information used for training and the feature map from the perspective of the BEV used for training are concatenated to obtain the feature z of the autonomous driving decision used for training; Construct the perception task branch and the behavior cloning task branch, and obtain the total loss to train the BEV visual perception model.
4. The end-to-end autonomous driving method based on bev feature space according to claim 3, characterized in that: The perception task branch includes a target detection task and a semantic segmentation task; The feature map under the BEV perspective used for training is input into the SSD detection head to obtain the result of target detection; the feature map under the BEV perspective used for training is input into the ASPP hole convolution module to obtain the result of semantic segmentation.
5. The end-to-end autonomous driving method based on bev feature space according to claim 2, characterized in that: The information of the image features is used in combination with convolution, sampling and pooling operations to construct a perceptual BEV feature space, and then obtain a feature map from the BEV perspective, including: Based on the obtained image feature map, the LSS depth estimation algorithm is used to construct a visual cone by upgrading the dimension of the obtained image feature map to predict the image depth; The three-dimensional point cloud is generated using the image feature map, depth information map and camera parameter information. The implementation formula is as follows: ; in, represents the image coordinates; Represents the point cloud coordinates; They are all matrices used to transform the reference system of point cloud coordinates. Used to transform the reference frame of the point cloud to the image, Used to correct the transformed reference system; is the camera internal parameter, through Realize the conversion from point cloud coordinates to image coordinates; Pool the 3D point cloud in the vertical direction to obtain preliminary BEV spatial features; Based on the preliminary BEV spatial features, the ResNet network and FPN are used to fuse features of different scales, and finally the optimized BEV spatial features are obtained, that is, the feature map from the BEV perspective.
6. The end-to-end autonomous driving method based on bev feature space according to claim 4, characterized in that: Use a multi-layer perceptron to complete the regression value prediction task of steering, throttle, and brake; Calculate the loss of each branch; use , They represent the target detection task and the semantic segmentation task of the feature map from the BEV perspective, respectively, and construct the cross entropy loss to calculate the loss of the target detection task and the semantic segmentation task: ; ; in, is the loss of the target detection task, It is the loss of the semantic segmentation task of the feature map from the BEV perspective; , Represent the true value of semantic segmentation of feature maps from the perspective of target detection and BEV respectively; , Represents the number of targets in the target detection task and the image semantic segmentation task respectively. It has no practical meaning. It means to sum up from the first item. The total number of items is , ; Represents the spatial features of the input from the BEV perspective; use , , Represents three regression tasks, namely steering regression task, throttle regression task, and brake regression task. Construct the mean square error loss MSE to calculate the loss of each regression task: ; ; ; in, , , They represent the losses of the steering regression task, the throttle regression task, and the brake regression task respectively. , , Represent the true values of steering, accelerator, and brake respectively, and z represents the input high-dimensional decision feature; The total loss combines the losses of the above five tasks, performs multi-task learning and minimizes the total loss To train the BEV visual perception model: ; in, , , , They are the weights of the target detection task loss, semantic segmentation task loss, steering regression task loss, throttle regression task loss, and brake regression task loss.
7. An end-to-end autonomous driving system based on bev feature space, characterized in that: include: An observation information collection module is configured to collect observation information of the driving route in real time, wherein the observation information includes surround view images, vehicle status information, and environmental information; The BEV visual perception module is configured to input the surround view image into the pre-trained BEV visual perception module to obtain a feature map from the BEV perspective; splice the vehicle state information and the obtained feature map from the BEV perspective to obtain a feature z for autonomous driving decision-making; The reinforcement learning module is configured to input the feature z into a pre-trained deep Q network algorithm to obtain a decision on the continuous action of the autonomous driving vehicle; the autonomous driving vehicle performs corresponding actions according to the obtained decision, thereby realizing autonomous driving.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the end-to-end autonomous driving method based on the bev feature space described in any one of claims 1 to 6 are implemented.
9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the end-to-end autonomous driving method based on the bev feature space as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps of the end-to-end autonomous driving method based on the bev feature space described in any one of claims 1 to 6 are implemented.