Automatic driving motion planning method, device, equipment and medium
By using Transformer-based encoder and scene feature deconstruction and reorganization algorithm in the autonomous driving motion planner, the problem of insufficient generalization ability of the autonomous driving system in long-tail scenarios is solved, and the planning performance and safety of the system in different scenarios is improved.
Patent Information
- Application Number
- CN202510430422.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The autonomous driving motion planner based on imitation learning performs well in common scenarios, but has weak generalization ability in long-tail scenarios, resulting in the inability to make reasonable decisions in complex or dangerous scenarios, increasing the risk of collision and reducing the safety and robustness of the system.
An autonomous driving motion planning method is adopted to perform feature encoding through a Transformer-based encoder to deconstruct and reorganize scene features. The specific steps include: predicting the encoded features through the scene classifier, calculating the cross entropy loss to obtain the loss gradient corresponding to the encoded features; deconstructing the encoded features into scene-related features and scene-general features according to the correlation between the loss gradient corresponding to the encoded features of different feature dimensions and scene-generated features; for scene categories with a higher number of scenes than the distribution mean, the corresponding scene-related features are interpolated and summed with the scene-related features of other categories of scenes to generate recombinant features; input the recombinant features into the trajectory decoder, output the future motion trajectory of the bicycle and the surrounding agent prediction trajectory, and jointly optimize the planning model parameters by imitating the loss and the motion prediction loss of the surrounding agent.
It improves the consistency of the planning performance of the autonomous driving system in different types of scenarios, enhances the model's learning ability for long-tail scenarios, reduces the risk of collision, and improves the safety and robustness of the system.
Smart Images

Figure CN119940680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of driving motion planning, and in particular to an automatic driving motion planning method, device, equipment and medium. Background Art
[0002] With the rapid development of artificial intelligence and autonomous driving technology, motion planning methods based on imitation learning have been widely used in the field of autonomous driving. Such methods learn a large amount of real driving data to generate trajectories that conform to human driving behavior to improve the safety and driving efficiency of autonomous driving systems. However, in actual applications, the data distribution of autonomous driving scenarios is extremely uneven, showing a long-tail distribution phenomenon.
[0003] Specifically, most autonomous driving datasets are mainly composed of common driving scenarios (such as going straight, following a car, changing lanes, etc.), while complex or dangerous scenarios such as sudden avoidance, complex interactions at intersections, and atypical driving behaviors in congested environments only account for a small number of them. This imbalance in data distribution causes motion planners based on imitation learning to perform well in common scenarios, but have weak generalization capabilities in long-tail scenarios. For example, in a dataset that mainly consists of "stationary" or "going straight" scenarios, the model may only learn a strategy to maintain the current speed or stand still, but cannot make reasonable decisions when emergency avoidance or complex interactions are required, thereby increasing the risk of collision, which is bound to reduce the safety and robustness of the autonomous driving system.
[0004] Some existing studies balance the number of different scenarios by setting an upper threshold for the number of different scenarios, but this will result in the planner being unable to fully utilize the advantages of data-driven performance improvement for imitation learning algorithms. Some other studies have attempted to improve the adaptability and generalization ability of the model in complex scenarios through a variety of data enhancement techniques. Among them, the most common enhancement strategy is to optimize at the input data level, such as using data perturbation to introduce a certain amount of noise to the current state of the vehicle to expand the diversity of the data, or using regularization to align the coordinate system of each element of the scene to the vehicle as the center to ensure the consistency of the data level. Although these methods have alleviated the problem of uneven data distribution to a certain extent, they usually directly adjust the element attributes in the input space, thus lacking diversity at the feature level, resulting in existing methods still being difficult to effectively deal with the uneven distribution of autonomous driving data. Therefore, how to enhance the model's learning ability for long-tail scenarios and improve its adaptability in various driving scenarios is an urgent problem that needs to be solved by autonomous driving planning algorithms. Summary of the invention
[0005] In order to solve the above technical problems, the present invention provides an autonomous driving motion planning method, device, equipment and medium.
[0006] In order to solve the above technical problems, the present invention adopts the following technical solutions: In a first aspect, the present invention provides an autonomous driving motion planning method, wherein the training process of the adopted planning model includes: Receiving a perception result, the perception result including the vehicle state, surrounding agent history information and map elements; inputting the perception result into a Transformer-based encoder for feature encoding to obtain an encoded feature; The coding features are subjected to scene feature deconstruction and reorganization: the scene type is predicted for the coding features by a scene classifier, and the cross entropy loss is calculated to obtain the loss gradient corresponding to the coding features; the coding features are deconstructed into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories; for scene categories with a number of scenes higher than the distribution mean, the corresponding scene-related features are interpolated and summed with the scene-related features of other categories of scenes to generate reorganized features; The reconstructed features are input into the trajectory decoder, which outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents. The planning model parameters are jointly optimized by the imitation loss and the motion prediction loss of the surrounding agents.
[0007] In one embodiment, the deconstructing the coded features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coded features of different feature dimensions and the scene categories specifically includes: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; Indicates the correlation between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, , represents the operation of finding the mean, represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates the dimension, The operation of averaging is performed when the encoding feature is used, which means that the values of all elements in the encoding feature are averaged to obtain the contribution value of the feature dimension of each encoding feature; ; Indicates the proportion Get the threshold for deconstructing the encoded features, Represents the quantile function, which returns the value at that proportion in the sorted data set; Represents the total feature dimensions of the encoded features, represents the scale parameter of the quantile; Then according to the proportion Get the threshold value for deconstructing the coded features and convert the coded features Disassemble as follows: ; Represents scene-related features, Represents the common features of the scene, represents the indicator function, , , Represents element-wise multiplication.
[0008] In one embodiment, for a scene category with a number of scenes higher than the distribution mean, interpolating and summing the corresponding scene-related features with scene-related features of other categories of scenes to generate a recombined feature specifically includes: ; For the recombination characteristics, is the interpolation ratio, is the scene-related feature of the current scene, are scene-related features of other scenes, It is the scene common feature of the current scene.
[0009] In one embodiment, the jointly optimizing the planning model parameters by the imitation loss and the motion prediction loss of the surrounding agent specifically includes: Imitation loss for: ; represents the total planning time, represents the position of the estimated trajectory of the self-driving car at time t, represents the position of the true value trajectory of the self-vehicle planning at the tth time point, Indicates calculation of Manhattan distance; Motion prediction loss for surrounding agents for: ; represents the total number of predicted agents, represents the position of the estimated trajectory of the nth agent motion prediction at time t, represents the position of the true trajectory of the nth agent motion prediction at time t; Through the total loss Optimize planning model parameters.
[0010] In one of the embodiments, the vehicle state includes the vehicle's position, heading angle, speed and acceleration; the surrounding agent history information includes the position, heading angle, speed, acceleration, angular velocity and bounding box size of each agent; the map elements include the global route of the vehicle, as well as the traffic light detection results, crossroad perception results and lane line perception results of the current frame.
[0011] In a second aspect, the present invention provides an automatic driving motion planning device, comprising: Encoding module: receiving the perception results, which include the vehicle status, surrounding agent history information and map elements; inputting the perception results into the Transformer-based encoder for feature encoding to obtain encoding features; Feature deconstruction and reorganization module: deconstruct and reorganize the scene features of the coding features, predict the scene type of the coding features through the scene classifier, calculate the cross entropy loss to obtain the loss gradient corresponding to the coding features; deconstruct the coding features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories; for scene categories with a number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other categories of scenes to generate reorganized features; The trajectory decoding module inputs the reconstructed features into the trajectory decoder, outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents, and jointly optimizes the planning model parameters through the imitation loss and the motion prediction loss of the surrounding agents.
[0012] In one embodiment, in the feature deconstruction and reorganization module, the coding features are deconstructed into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories, specifically including: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; Indicates the correlation between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, represents the operation of finding the mean, represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates the dimension, The operation of averaging is performed when the encoding feature is used, which means that the values of all elements in the encoding feature are averaged to obtain the contribution value of the feature dimension of each encoding feature; ; Indicates the proportion Get the threshold for deconstructing the encoded features, Represents the quantile function, which returns the value at that proportion in the sorted data set; Represents the total feature dimensions of the encoded features, represents the scale parameter of the quantile; Then according to the proportion Get the threshold value for deconstructing the coded features and convert the coded features Disassemble as follows: ; Represents scene-related features, Represents the common features of the scene, represents the indicator function, , , Represents element-wise multiplication.
[0013] In one embodiment, in the feature deconstruction and reorganization module, for a scene category whose number of scenes is higher than the distribution mean, the corresponding scene-related features are interpolated and summed with scene-related features of other categories of scenes to generate reorganized features, specifically including: ; For the recombination characteristics, is the interpolation ratio, is the scene-related feature of the current scene, are scene-related features of other scenes, It is the scene common feature of the current scene.
[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of any one of the embodiments in the first aspect when executing the computer program.
[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method of any one of the embodiments in the first aspect.
[0016] Compared with the prior art, the beneficial technical effects of the present invention are: (1) The present invention proposes a motion planning network that considers balanced optimization of data features. This method takes into account the uneven distribution of different types of scenes. It deconstructs scene features during network processing and reconstructs features for more common scene types, so that the planner will not be restricted by the influence of long-tail distribution during training, improves the consistency of planning performance on different types of scenes, and improves the planning effect.
[0017] (2) The present invention designs a deconstruction and reconstruction algorithm for the encoded features of scene information. An additional scene classifier is introduced into the training architecture to determine the correlation between the encoded information and the scene category. The features are deconstructed into scene-related parts and scene-common features according to the weights of features of different dimensions and the gradient of the classifier prediction loss feedback. For scene categories with a number of scenes higher than the distribution mean, the scene-related partial features are interpolated and summed with the scene-related features of scenes of different categories to prevent the planner from overfitting to the features of these dominant scenes, thereby improving the planner performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of the model training method in an embodiment of the present invention.
[0019] Figure 2 It is a model framework diagram in an embodiment of the present invention.
[0020] Figure 3a , Figure 3b They are respectively diagrams of simulation results of a left-turn automatic driving task using the method of the present invention and the baseline method in embodiments of the present invention.
[0021] Figure 4a , Figure 4b They are respectively diagrams of simulation results of a straight-line autonomous driving task performed using the method of the present invention and the baseline method in an embodiment of the present invention.
[0022] Figure 5a , Figure 5b They are respectively diagrams of simulation results of pedestrian avoidance autonomous driving tasks using the method of the present invention and the baseline method in embodiments of the present invention.
[0023] Figure 6a , Figure 6b They are respectively diagrams of simulation results of a right turn automatic driving task using the method of the present invention and the baseline method in embodiments of the present invention. DETAILED DESCRIPTION
[0024] A preferred embodiment of the present invention is described in detail below with reference to the accompanying drawings.
[0025] The present invention proposes a motion planning network that considers balanced optimization of data features. This method takes into account the uneven distribution of different types of scenes. It will deconstruct the scene features during the network processing and reconstruct the features of more common scene types, so that the planner will not be limited by the influence of long-tail distribution during training, improve the consistency of planning performance on different types of scenes, and improve the planning effect.
[0026] The present invention designs a deconstruction and reconstruction algorithm for the encoding features of scene information, introduces an additional scene classifier in the training architecture to determine the correlation between the encoding information and the scene category, deconstructs the features into scene-related parts and scene-common features according to the weights of features in different dimensions and the gradient of the classifier prediction loss return, and for scene categories with a number of scenes higher than the distribution mean, interpolates and sums the scene-related partial features with the scene-related features of scenes of different categories to prevent the planner from overfitting to the features of these dominant scenes, thereby improving the planner performance.
[0027] like Figure 1 As shown, in an automatic driving motion planning method of the present invention, the training process of the planning model adopted includes the following steps: S1, receiving perception results, the perception results including the vehicle state, surrounding agent history information and map elements; inputting the perception results into a Transformer-based encoder for feature encoding to obtain encoded features.
[0028] S2, performing scene feature deconstruction and reorganization on the coding features: predicting the scene type of the coding features through a scene classifier, calculating the cross entropy loss to obtain the loss gradient corresponding to the coding features; deconstructing the coding features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories; for scene categories whose number of scenes is higher than the distribution mean, interpolating and summing the corresponding scene-related features with the scene-related features of other categories of scenes to generate reorganized features.
[0029] S3, inputs the reorganized features into the trajectory decoder, outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents, and jointly optimizes the planning model parameters through the imitation loss and the motion prediction loss of the surrounding agents.
[0030] In one of the embodiments, the vehicle state in step S1 includes the position, heading angle, speed and acceleration of the vehicle; the surrounding agent history information includes the position, heading angle, speed, acceleration, angular velocity and bounding box size of each agent; the map elements include the global route of the vehicle, as well as the traffic light detection results, crossroad perception results and lane line perception results of the current frame.
[0031] Assumptions is the token input to this step, This is the token for the vehicle status. There is only one token. is the token corresponding to the surrounding agents, including N tokens, where N is the number of input agent information; is the token corresponding to the map element, which contains L tokens, so Contains M=1+N+L tokens in total. Use Transformer to Encode to obtain the encoding feature F.
[0032] In one embodiment, in step S2, according to the correlation between the loss gradients corresponding to the coded features of different feature dimensions and the scene categories, the coded features are deconstructed into scene-related features and scene-general features, specifically including: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; ; Then according to the proportion Get the threshold value for deconstructing the coded features, and decompose the coded features as follows: ; Represents scene-related features, Represents common features of the scene.
[0033] In one embodiment, for a scene category with a number of scenes higher than the distribution mean, in step S2, the corresponding scene-related features are interpolated and summed with scene-related features of other categories of scenes to generate recombined features, specifically including: ; is the interpolation ratio, is the scene-related feature of the current scene, It is the scene-related features of other scenes.
[0034] Specifically, the encoded perception result is used as input and encoded using the Transformer network. The encoded features are used to classify the scene type and calculate the loss to obtain the loss gradient corresponding to the feature. The features are decomposed into scene category-related features and common features between different scenes according to the product of the feature value and the gradient at a fixed ratio. Then, for the common correlation features in the main scene, different common scene correlation features are used for interpolation. This method destroys the correlation features of the main scene to prevent the model from overfitting, and retains the common features, which is conducive to more efficient training and performance improvement of the planning network. The features are input into the classifier to predict the current scene type, and the cross entropy loss between the predicted value and the true value is calculated, and the loss gradient of the encoded features is obtained. Then, the correlation between the features of different dimensions and the scene category is calculated. Then, the threshold of feature deconstruction is obtained according to the ratio and the features are decomposed according to the following formula. The former represents the features that are highly correlated with the scene category, and the latter represents the features that are common between different scenes. According to the number and distribution mean of different types of scenes, it is judged whether the scene is in the main scene type (category with a higher number than the mean). For the features of the main scene, the scene-related features of another scene different from the current scene type are used for interpolation.
[0035] In one embodiment, the step S3 of jointly optimizing the planning model parameters by using the imitation loss and the motion prediction loss of the surrounding agents specifically includes: Imitation loss for: ; represents the total planning time, represents the position of the estimated trajectory of the self-driving car at time t, represents the position of the true value trajectory of the self-vehicle planning at the tth time point, Indicates calculation of Manhattan distance; Motion prediction loss for surrounding agents for: ; represents the total number of predicted agents, represents the position of the estimated trajectory of the nth agent motion prediction at time t, represents the position of the true trajectory of the nth agent motion prediction at time t; Through the total loss Optimize planning model parameters.
[0036] Specifically, the present invention uses the encoding result after feature interpolation as input, and the trajectory decoder structure outputs the motion trajectory of the current vehicle in the next 80 frames and the predicted trajectory of the surrounding agents in the next 80 frames. These two results are output only during network training. In the actual network reasoning process, only the motion trajectory of the vehicle in the next 80 frames is output. In the planning network training process, the main loss function is the imitation loss, and the auxiliary loss function is the motion prediction loss of the surrounding agents.
[0037] It should be understood that, although the steps in the flowcharts of the accompanying drawings of the specification are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts of the accompanying drawings of the specification may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.
[0038] The present invention also provides an automatic driving motion planning device. The device may be a system (including a distributed system), software (application), module, component, server, client, etc. that uses the method described in the embodiments of this specification and is combined with necessary implementation hardware. Based on the same innovative concept, the device in one or more embodiments provided in the embodiments of the present disclosure is as described in the following embodiments. Since the implementation scheme of the device to solve the problem is similar to the method, the implementation of the specific device in the embodiments of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "module" or "module group" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0039] Specifically, an automatic driving motion planning device includes: Encoding module: receiving the perception results, which include the vehicle status, surrounding agent history information and map elements; inputting the perception results into the Transformer-based encoder for feature encoding to obtain encoding features; Feature deconstruction and reorganization module: deconstruct and reorganize the scene features of the coding features, predict the scene type of the coding features through the scene classifier, calculate the cross entropy loss to obtain the loss gradient corresponding to the coding features; deconstruct the coding features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different dimensions and the scene categories; for scene categories with a number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other categories of scenes to generate reorganized features; The trajectory decoding module inputs the reconstructed features into the trajectory decoder, outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents, and jointly optimizes the planning model parameters through the imitation loss and the motion prediction loss of the surrounding agents.
[0040] In one embodiment, the feature deconstruction and reorganization module deconstructs the coded features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coded features of different feature dimensions and the scene categories, specifically including: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; Indicates the correlation between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, represents the operation of finding the mean, represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates the dimension, The operation of averaging is performed when the encoding feature is used, which means that the values of all elements in the encoding feature are averaged to obtain the contribution value of the feature dimension of each encoding feature; ; Indicates the proportion Get the threshold for deconstructing the encoded features, Represents the quantile function, which returns the value at that proportion in the sorted data set; Represents the total feature dimensions of the encoded features, represents the scale parameter of the quantile; Then according to the proportion Get the threshold value for deconstructing the coded features and convert the coded features Disassemble as follows: ; Represents scene-related features, Represents the common features of the scene, represents the indicator function, , Represents element-wise multiplication.
[0041] In one embodiment, in the feature deconstruction and reorganization module, for a scene category whose number of scenes is higher than the distribution mean, the corresponding scene-related features are interpolated and summed with scene-related features of other categories of scenes to generate reorganized features, specifically including: ; For the recombination characteristics, is the interpolation ratio, is the scene-related feature of the current scene, are scene-related features of other scenes, It is the scene common feature of the current scene.
[0042] In one embodiment, the present invention also provides a computer device. The computer device includes a processor, a memory and a network interface connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data used in the above method. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0043] In an exemplary embodiment, the present invention also provides a computer-readable storage medium including instructions, such as a memory including instructions, and the above instructions can be executed by a processor to perform the above method. The storage medium can be a computer-readable storage medium, for example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0044] In one embodiment, if Figure 3a and Figure 3b As shown, the present invention also simulates the left-turn autonomous driving task, using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention respectively. Figure 3a As shown, the automatic driving motion planning method of the present invention is used to successfully complete the automatic driving task of the left turn situation; Figure 3b As shown, the baseline method suffers from overfitting the long-tail distribution, resulting in failure to make positive decisions (no move).
[0045] In one embodiment, if Figure 4aand Figure 4b As shown, the present invention also simulates a straight-line autonomous driving task, using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention respectively. Figure 4a As shown, the automatic driving motion planning method of the present invention is used to successfully complete the automatic driving task in a straight line situation; Figure 4b As shown, the baseline method crashes in the straight-line autonomous driving task.
[0046] In one embodiment, if Figure 5a and Figure 5b As shown, the present invention also simulates the pedestrian avoidance autonomous driving task, using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention respectively. Figure 5a As shown, the autonomous driving motion planning method of the present invention is used to successfully complete the pedestrian avoidance autonomous driving task, actively avoid pedestrians, and show positive decision-making after the pedestrians leave; Figure 5b As shown, the vehicle in the baseline method collides with the pedestrian.
[0047] In one embodiment, if Figure 6a and Figure 6b As shown, the present invention also simulates the right turn automatic driving task, respectively using the automatic driving motion planning method and the baseline method (PLUTO) in the present invention. Figure 6a As shown, the automatic driving motion planning method of the present invention is used to successfully complete the automatic driving task of the right turn situation; Figure 6b As shown, the baseline method suffers from overfitting the long-tail distribution, resulting in failure to make positive decisions (no move).
[0048] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.
[0049] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
Claims
1. A method for autonomous driving motion planning, characterized in that: The training process of the adopted planning model includes: Receiving a perception result, the perception result including the vehicle state, surrounding agent history information and map elements; inputting the perception result into a Transformer-based encoder for feature encoding to obtain an encoded feature; The coding features are subjected to scene feature deconstruction and reorganization: the scene type is predicted for the coding features by a scene classifier, and the cross entropy loss is calculated to obtain the loss gradient corresponding to the coding features; the coding features are deconstructed into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories; for scene categories with a number of scenes higher than the distribution mean, the corresponding scene-related features are interpolated and summed with the scene-related features of other categories of scenes to generate reorganized features; The reconstructed features are input into the trajectory decoder, which outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents. The planning model parameters are jointly optimized by the imitation loss and the motion prediction loss of the surrounding agents.
2. The automatic driving motion planning method according to claim 1, characterized in that: According to the correlation between the loss gradients corresponding to the coded features of different feature dimensions and the scene categories, the coded features are deconstructed into scene-related features and scene-general features, specifically including: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; Indicates the correlation between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, , represents the operation of finding the mean, represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates the dimension, The operation of averaging is performed when the encoding feature is used, which means that the values of all elements in the encoding feature are averaged to obtain the contribution value of the feature dimension of each encoding feature; ; Indicates the proportion Get the threshold for deconstructing the encoded features, Represents the quantile function, which returns the value at that proportion in the sorted data set; Represents the total feature dimensions of the encoded features, represents the scale parameter of the quantile; Then according to the proportion Get the threshold value for deconstructing the coded features and convert the coded features Disassemble as follows: ; Represents scene-related features, Represents the common features of the scene, represents the indicator function, , , Represents element-wise multiplication.
3. The automatic driving motion planning method according to claim 1, characterized in that: For a scene category whose number of scenes is higher than the distribution mean, the corresponding scene-related features are interpolated and summed with scene-related features of other categories of scenes to generate recombined features, specifically including: ; For the recombination characteristics, is the interpolation ratio, is the scene-related feature of the current scene, are scene-related features of other scenes, It is the scene common feature of the current scene.
4. The automatic driving motion planning method according to claim 1, characterized in that: The method of jointly optimizing the planning model parameters by imitating the loss and the motion prediction loss of the surrounding agent specifically includes: Imitation loss for: ; represents the total planning time, represents the position of the estimated trajectory of the self-driving car at time t, represents the position of the true value trajectory of the self-vehicle planning at the tth time point, Indicates calculation of Manhattan distance; Motion prediction loss for surrounding agents for: ; represents the total number of predicted agents, represents the position of the estimated trajectory of the nth agent motion prediction at time t, represents the position of the true trajectory of the nth agent motion prediction at time t; Through the total loss Optimize planning model parameters.
5. The automatic driving motion planning method according to claim 1, characterized in that: The vehicle state includes the vehicle's position, heading angle, speed and acceleration; the surrounding agent history information includes the position, heading angle, speed, acceleration, angular velocity and bounding box size of each agent; the map elements include the global route of the vehicle, as well as the traffic light detection results, crossroads perception results and lane line perception results of the current frame.
6. An automatic driving motion planning device, characterized in that: include: Encoding module: receiving perception results, which include vehicle status, surrounding agent history information and map elements; Inputting the perception result into a Transformer-based encoder for feature encoding to obtain encoded features; Feature deconstruction and reorganization module: deconstruct and reorganize the scene features of the coding features, predict the scene type of the coding features through the scene classifier, calculate the cross entropy loss to obtain the loss gradient corresponding to the coding features; deconstruct the coding features into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coding features of different feature dimensions and the scene categories; for scene categories with a number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other categories of scenes to generate reorganized features; The trajectory decoding module inputs the reconstructed features into the trajectory decoder, outputs the future motion trajectory of the ego vehicle and the predicted trajectory of the surrounding agents, and jointly optimizes the planning model parameters through the imitation loss and the motion prediction loss of the surrounding agents.
7. The automatic driving motion planning device according to claim 6, characterized in that: In the feature deconstruction and reorganization module, the coded features are deconstructed into scene-related features and scene-general features according to the correlation between the loss gradients corresponding to the coded features of different feature dimensions and the scene categories, specifically including: Calculate the correlation between the loss gradient corresponding to the encoding features of different feature dimensions and the scene category: ; Indicates the correlation between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, represents the operation of finding the mean, represents the encoding feature of the i-th feature dimension, Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates the dimension, The operation of averaging is performed when the encoding feature is used, which means that the values of all elements in the encoding feature are averaged to obtain the contribution value of the feature dimension of each encoding feature; ; Indicates the proportion Get the threshold for deconstructing the encoded features, Represents the quantile function, which returns the value at that proportion in the sorted data set; Represents the total feature dimensions of the encoded features, represents the scale parameter of the quantile; Then according to the proportion Get the threshold value for deconstructing the coded features and convert the coded features Disassemble as follows: ; Represents scene-related features, Represents the common features of the scene, represents the indicator function, , , Represents element-wise multiplication.
8. The automatic driving motion planning device according to claim 6, characterized in that: In the feature deconstruction and reorganization module, for the scene category whose number of scenes is higher than the distribution mean, the corresponding scene-related features are interpolated and summed with the scene-related features of other categories of scenes to generate reorganized features, specifically including: ; For the recombination characteristics, is the interpolation ratio, is the scene-related feature of the current scene, are scene-related features of other scenes, It is the scene common feature of the current scene.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Automatic iteration method of trajectory prediction model, electronic equipment and storage medium
CN114880842A
Automatic driving motion planning and model training method, device and equipment
CN118734923A
Long-tail pedestrian trajectory prediction method based on subclass balance contrast learning
CN119048556A
Long-tail scene recognition method and device based on vehicle supervision data, medium and product
CN119339456A
Automatic driving track prediction method and system based on track decoupling and Mama
CN119348649A