Autonomous Driving Motion Planning Method, Device, Equipment and Medium
By deconstructing and recombinant features based on Transformer's encoder and scene classifier, the problem of weak generalization ability of long-tail scenarios in the autonomous driving system under uneven data distribution is solved, and more efficient planning effects and safety are achieved.
Patent Information
- Application Number
- CN202510430422.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing autonomous driving motion planning methods are weak in the face of uneven data distribution, especially in long-tail scenarios, resulting in reduced safety and robustness. It is difficult for existing methods to effectively deal with uneven data distribution.
Feature encoding is used for Transformer-based encoder, and features are deconstructed and reorganized by scene classifiers, and the planning model parameters are optimized using cross-entropy loss and imitation loss, feature reconstruction is carried out for common scenarios, and features higher than the mean scenario category are interpolated to improve the model's adaptability in different scenarios.
It improves the planning performance and consistency of the autonomous driving system in different types of scenarios, reduces the overfitting phenomenon, enhances the model's learning ability for long-tail scenarios, and improves the safety and robustness of autonomous driving.
Smart Images

Figure CN119940680B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of driving motion planning, and particularly to an autonomous driving motion planning method, device, equipment and medium. Background Art
[0002] With the rapid development of artificial intelligence and autonomous driving technologies, motion planning methods based on imitation learning have been widely applied in the field of autonomous driving. These methods generate trajectories that conform to human driving behaviors by learning a large amount of real driving data, so as to improve the safety and driving efficiency of autonomous driving systems. However, in practical applications, the data distribution in autonomous driving scenarios is extremely unbalanced, showing a long-tail distribution phenomenon.
[0003] Specifically, most autonomous driving data sets mainly consist of common driving scenarios (such as going straight, following a vehicle, changing lanes, etc.), while complex or dangerous scenarios such as sudden avoidance, complex interactions at intersections, and atypical driving behaviors in congested environments only account for a small proportion in the data set. This imbalance in data distribution results in motion planners based on imitation learning performing well in common scenarios but having weak generalization ability in long-tail scenarios. For example, in a data set mainly composed of "stationary" or "going straight" scenarios, the model may only learn strategies to maintain the current speed or stay stationary, but be unable to make reasonable decisions in situations that require emergency avoidance or complex interactions, thus increasing the collision risk, which will inevitably reduce the safety and robustness of the autonomous driving system.
[0004] Some existing studies try to balance the number of different scenarios by setting upper threshold values for the number of different scenarios, but this will cause the planner to not fully utilize the advantages of data-driven improvement of the performance of imitation learning algorithms. There are also some studies that attempt to improve the adaptability and generalization ability of the model in complex scenarios through various data augmentation techniques. Among them, the most common augmentation strategy is to optimize at the input data level. For example, a certain amount of noise is introduced into the current state of the ego vehicle in a data perturbation manner to expand the diversity of the data, or a regularization method is used to unify the coordinate systems of the various elements of the scenario to the ego vehicle as the center to ensure the consistency at the data level. Although these methods alleviate the problem of unbalanced data distribution to a certain extent, they usually directly adjust the element attributes in the input space, lacking diversity at the feature level, resulting in the existing methods still being difficult to effectively cope with the unbalanced data distribution in autonomous driving. Therefore, how to enhance the model's learning ability for long-tail scenarios and improve its adaptability in various driving situations is an urgent problem to be solved in autonomous driving planning algorithms. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides an autonomous driving motion planning method, device, equipment and medium.
[0006] To solve the above technical problems, the present invention adopts the following technical solutions:
[0007] In a first aspect, the present invention provides an autonomous driving motion planning method, and the training process of the adopted planning model includes:
[0008] Receiving perception results, where the perception results include the self-vehicle state, surrounding agent historical information, and map elements; inputting the perception results into an encoder based on Transformer for feature encoding to obtain encoded features;
[0009] Performing scene feature deconstruction and recombination on the encoded features: predicting the scene type of the encoded features through a scene classifier, calculating the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories, deconstructing the encoded features into scene-related features and scene-general features; for the scene categories with the number of scenes higher than the distribution mean, interpolating and summing the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features;
[0010] Inputting the recombined features into a trajectory decoder, outputting the future motion trajectory of the self-vehicle and the predicted trajectories of surrounding agents, and jointly optimizing the parameters of the planning model through imitation loss and the motion prediction loss of surrounding agents.
[0011] In one embodiment, the deconstruction of the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories specifically includes:
[0012] Calculating the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories:
[0013] ;
[0014] represents the relevance between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, , represents the operation of taking the mean, represents the encoded feature of the i-th feature dimension, represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; represents the dimension, and the operation of taking the mean is performed when , indicating that the values of all elements in the encoded features are averaged to obtain the contribution value of each feature dimension of each encoded feature;
[0015] ;
[0016] Indicates according to a ratio Obtain the threshold for deconstructing the encoded features Indicates the quantile function, and returns the value at that ratio position in the sorted dataset Indicates the total feature dimension of the encoded features Indicates the ratio parameter of the quantile
[0017] After that, according to the ratio Obtain the threshold for deconstructing the encoded features, and deconstruct the encoded features Decompose according to the following formula
[0018] ;
[0019] Represents scene-related features Represents scene-general features Indicates the indicator function , , Indicates element-wise multiplication
[0020] In one embodiment, for the scene category with the number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other category scenes to generate a recombined feature, specifically including
[0021] ;
[0022] Is the recombined feature Is the interpolation ratio Is the scene-related feature of the current scene Is the scene-related feature of other scenes Is the scene-general feature of the current scene
[0023] In one embodiment, the planning model parameters are jointly optimized by the imitation loss and the motion prediction loss of surrounding agents, specifically including
[0024] Imitation loss Is
[0025] ;
[0026] Indicates the total number of planning time steps Indicates the position of the ego vehicle's planned estimated trajectory at the t-th time step Indicates the position of the ego vehicle's planned ground truth trajectory at the t-th time step Indicates calculating the Manhattan distance
[0027] Motion prediction loss of surrounding agents is:
[0028] ;
[0029] represents the total number of predicted agents, represents the position of the estimated trajectory of the nth agent's motion prediction at the t-th time point, represents the position of the true trajectory of the nth agent's motion prediction at the t-th time point;
[0030] The total loss is used to optimize the parameters of the planning model.
[0031] In one embodiment, the ego-vehicle state includes the position, orientation angle, speed, and acceleration of the ego-vehicle; the surrounding agent historical information includes the position, orientation angle, speed, acceleration, angular velocity, and bounding box size of each agent; the map elements include the global route on which the vehicle travels, as well as the traffic light detection result, crossroad condition perception result, and lane line perception result of the current frame.
[0032] In a second aspect, the present invention provides an autonomous driving motion planning device, including:
[0033] An encoding module: receiving the perception results, where the perception results include the ego-vehicle state, surrounding agent historical information, and map elements; inputting the perception results into an encoder based on Transformer for feature encoding to obtain encoded features;
[0034] A feature deconstruction and recombination module: deconstructing and recombining the scene features of the encoded features, predicting the scene type of the encoded features through a scene classifier, calculating the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; deconstructing the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories; for the scene categories with the number of scenes higher than the distribution mean, interpolating and summing the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features;
[0035] A trajectory decoding module, inputting the recombined features into a trajectory decoder, outputting the future motion trajectory of the ego-vehicle and the predicted trajectories of surrounding agents, and jointly optimizing the parameters of the planning model through the imitation loss and the motion prediction loss of surrounding agents.
[0036] In one embodiment, in the feature deconstruction and recombination module, the deconstructing the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories specifically includes:
[0037] Calculating the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories:
[0038] ;
[0039] Represents the relevance between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category. Represents the operation of taking the mean. Represents the encoded feature of the i-th feature dimension. Represents the loss gradient corresponding to the encoded feature of the i-th feature dimension. Represents the dimension. At perform the operation of taking the mean, which means averaging the values of all elements in the encoded feature to obtain the contribution value of each feature dimension of each encoded feature.
[0040] ;
[0041] Represents obtaining the threshold for decomposing the encoded feature according to the ratio The quantile function returns the value at that ratio position in the sorted dataset. Represents the total number of feature dimensions of the encoded feature. Represents the ratio parameter of the quantile. After that, obtain the threshold for decomposing the encoded feature according to the ratio
[0042] and decompose the encoded feature according to the following formula:
[0043] ;
[0044] Represents the scene-related feature. Represents the scene-general feature. Represents the indicator function. , , Represents element-wise multiplication.
[0045] In one embodiment, in the feature decomposition and recombination module, for the scene category with the number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other category scenes to generate a recombined feature, which specifically includes:
[0046] ;
[0047] is the recombined feature. is the interpolation ratio. is the scene-related feature of the current scene. are scene-related features of other scenarios, which are scene-general features of the current scenario.
[0048] In a third aspect, the present invention provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the method in any one of the embodiments of the first aspect are implemented.
[0049] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in any one of the embodiments of the first aspect are implemented.
[0050] Compared with the prior art, the beneficial technical effects of the present invention are:
[0051] (1) The present invention proposes a motion planning network considering data feature balance optimization. This method takes into account the uneven distribution of different types of scenarios, deconstructs the scene features during the network processing, and reconstructs the features for the more common scene types to prevent the planner from being limited by the long-tail distribution during training, improve the consistency of the planning performance on different types of scenarios, and improve the planning effect.
[0052] (2) The present invention designs an algorithm for deconstructing and reorganizing the encoded features of scene information, introduces an additional scene classifier in the training architecture to judge the relevance between the encoded information and the scene categories, and deconstructs the features into scene-related parts and scene-general features according to the weights of different-dimensional features and the gradients backpropagated by the classifier prediction loss. For the scene categories with the number of scenes higher than the distribution mean, interpolate and sum the scene-related partial features of these scene categories with the scene-related features of different-category scenes to prevent the planner from overfitting to the features of these dominant scenes, thereby improving the performance of the planner. Description of the Drawings
[0053] Figure 1 is a flowchart of the model training method in the embodiment of the present invention.
[0054] Figure 2 is a model framework diagram in the embodiment of the present invention.
[0055] Figure 3a 、 Figure 3b are respectively simulation result diagrams of the left-turn automatic driving task using the method of the present invention and the baseline method in the embodiment of the present invention.
[0056] Figure 4a 、 Figure 4b are respectively simulation result diagrams of the straight-line automatic driving task using the method of the present invention and the baseline method in the embodiment of the present invention.
[0057] Figure 5a and Figure 5b are respectively the simulation result diagrams of the pedestrian avoidance for the autonomous driving task using the method of the present invention and the baseline method in the embodiments of the present invention.
[0058] Figure 6a and Figure 6b are respectively the simulation result diagrams of the right - turn autonomous driving task using the method of the present invention and the baseline method in the embodiments of the present invention. Detailed Embodiment
[0059] A preferred embodiment of the present invention will be described in detail below with reference to the accompanying drawings.
[0060] The present invention proposes a motion planning network considering the optimization of data - feature balance. This method takes into account the uneven distribution of different types of scenarios. During the network processing, it will deconstruct the scenario features and reconstruct the features for the more common scenario types to prevent the planner from being limited by the long - tail distribution during training, improve the consistency of the planning performance on different types of scenarios, and enhance the planning effect.
[0061] The present invention designs an algorithm for deconstructing and reorganizing the encoded features of scenario information. An additional scenario classifier is introduced into the training architecture to judge the correlation between the encoded information and the scenario categories. According to the weights of different - dimensional features and the gradients back - propagated from the classifier prediction loss, the features are deconstructed into scenario - related parts and scenario - general features. For the scenario categories with the number of scenarios higher than the distribution mean, the scenario - related partial features of it are interpolated and summed with the scenario - related features of different - category scenarios to prevent the planner from over - fitting to the features of these dominant scenarios, thereby improving the performance of the planner.
[0062] As Figure 1 shown, the training process of the planning model adopted by an autonomous driving motion planning method in the present invention includes the following steps:
[0063] S1. Receive the perception results, where the perception results include the vehicle's own state, the historical information of surrounding agents, and map elements; input the perception results into an encoder based on Transformer for feature encoding to obtain encoded features.
[0064] S2. Decompose and reorganize the encoded features according to scene features: Predict the scene type of the encoded features through a scene classifier, calculate the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; According to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories, decompose the encoded features into scene-related features and scene-general features; For the scene categories with the number of scenes higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other category scenes to generate reorganized features.
[0065] S3. Input the reorganized features into the trajectory decoder, output the future motion trajectory of the ego vehicle and the predicted trajectories of surrounding agents, and jointly optimize the parameters of the planning model through the imitation loss and the motion prediction loss of the surrounding agents.
[0066] In one embodiment, the ego vehicle state in step S1 includes the position, orientation angle, speed, and acceleration of the ego vehicle; the historical information of surrounding agents includes the position, orientation angle, speed, acceleration, angular velocity, and bounding box size of each agent; the map elements include the global route of vehicle driving, as well as the traffic light detection results, crossroad condition perception results, and lane line perception results of the current frame.
[0067] Suppose is the token input to this step, is the token of the ego vehicle state, and there is only one token; is the token corresponding to the surrounding agents, including N tokens, where N is the number of input agent information; is the token corresponding to the map elements, including L tokens, so altogether contains M = 1 + N + L tokens. Use Transformer to encode to obtain the encoded feature F.
[0068] In one embodiment, according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories in step S2, decomposing the encoded features into scene-related features and scene-general features specifically includes:
[0069] Calculate the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories:
[0070] ;
[0071] represents the encoded feature of the i-th feature dimension, represents the loss gradient corresponding to the encoded feature of the i-th feature dimension;
[0072] ;
[0073] After that, according to the ratio obtain the threshold for deconstructing the encoded features, and decompose the encoded features according to the following formula:
[0074] ;
[0075] represents scene-related features, represents scene-general features.
[0076] In one embodiment, for the scene categories where the number of scenes in step S2 is higher than the distribution mean, interpolate and sum the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features, specifically including:
[0077] ;
[0078] is the interpolation ratio, is the scene-related feature of the current scene, is the scene-related feature of other scenes.
[0079] Specifically, use the Transformer network to encode with the encoded perception result as the input, classify the scene type using the encoded features and calculate the loss to obtain the loss gradient corresponding to the features. Decompose the features into features related to scene categories and features common among different scenes according to the product of the feature value and the gradient at a fixed set ratio. Then, for the common correlation features in the main scenes, use different common scene correlation features for interpolation. This method destroys the correlation features of the main scenes to prevent the model from overfitting, and at the same time retains the common features, which is beneficial to the more efficient training of the planning network and improves the performance. Input the features into the classifier to predict the current scene type, calculate the cross-entropy loss between the predicted value and the true value, and obtain the loss gradient of the encoded features. Then calculate the correlation between features of different dimensions and scene categories. After that, obtain the threshold for feature deconstruction according to the ratio and decompose the features according to the following formula. The former term represents the features highly correlated with the scene category, and the latter term represents the features common among different scenes. Judge whether the scene is in the main scene type (the category with a number higher than the mean) according to the number and distribution mean of different types of scenes. For the features of the main scene, use the scene-related features of another scene different from the current scene type for interpolation.
[0080] In one embodiment, in step S3, jointly optimize the planning model parameters through the imitation loss and the motion prediction loss of surrounding agents, specifically including:
[0081] The imitation loss is: ;
[0082] represents the total planned time represents the position of the ego vehicle's planned estimated trajectory at the t-th moment represents the position of the ego vehicle's true planned trajectory at the t-th moment represents calculating the Manhattan distance;
[0083] Motion prediction loss of surrounding agents is: ;
[0084] represents the total number of predicted agents represents the position of the estimated trajectory of the n-th agent's motion prediction at the t-th moment represents the position of the true trajectory of the n-th agent's motion prediction at the t-th moment;
[0085] Optimize the parameters of the planning model through the total loss
[0086] Specifically, the present invention takes the encoded result after feature interpolation as input, and the decoder structure of the trajectory outputs the future 80-frame motion trajectory of the ego vehicle at the current moment and the predicted trajectories of surrounding agents in the future 80 frames. These two results are only output during network training. During the actual inference process of the network, only the future 80-frame motion trajectory of the ego vehicle is output. During the training process of the planning network, the main loss function is the imitation loss, and the auxiliary loss function is the motion prediction loss of surrounding agents.
[0087] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown in sequence according to the indication of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0088] The present invention also provides an autonomous driving motion planning device. The device may be a system (including a distributed system), software (application), module, component, server, client, etc. that uses the method described in the embodiments of this specification and combines the necessary implementation hardware. Based on the same innovative concept, the devices in one or more embodiments provided by the embodiments of the present disclosure are as described in the following embodiments. Since the implementation solutions for the device to solve problems are similar to those of the method, the implementation of the specific device in the embodiments of this specification can refer to the implementation of the foregoing method, and the repeated parts will not be elaborated. As used hereinafter, the term "module" or "modular" may be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0089] Specifically, an autonomous driving motion planning device includes:
[0090] An encoding module: receives perception results, where the perception results include the vehicle's own state, surrounding agent historical information, and map elements; inputs the perception results into an encoder based on Transformer for feature encoding to obtain encoded features;
[0091] A feature deconstruction and recombination module: deconstructs and recombines the scene features of the encoded features, predicts the scene type of the encoded features through a scene classifier, calculates the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; deconstructs the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different dimensions and the scene categories; for the scene categories with the number of scenes higher than the distribution mean, interpolates and sums the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features;
[0092] A trajectory decoding module inputs the recombined features into a trajectory decoder, outputs the future motion trajectory of the vehicle itself and the predicted trajectories of surrounding agents, and jointly optimizes the parameters of the planning model through imitation loss and the motion prediction loss of surrounding agents.
[0093] In one embodiment, deconstructing the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories in the feature deconstruction and recombination module specifically includes:
[0094] Calculating the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories:
[0095] ;
[0096] represents the relevance between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, represents the operation of calculating the mean value, represents the encoded feature of the i-th feature dimension, represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; represents the dimension, at when the operation of calculating the mean value is performed, it means averaging the values of all elements in the encoded feature to obtain the contribution value of each feature dimension of the encoded feature;
[0097] ;
[0098] represents according to the ratio obtains the threshold for decomposing the encoded feature structure, represents the quantile function, and the value returned is the value at that proportion position in the sorted dataset; represents the total number of feature dimensions of the encoded feature, represents the proportional parameter of the quantile;
[0099] After that, according to the ratio obtains the threshold for decomposing the encoded feature structure, and decomposes the encoded feature as follows:
[0100] ;
[0101] represents the scene-related features, represents the scene-general features, represents the indicator function, , represents the element-wise multiplication.
[0102] In one embodiment, in the feature decomposition and recombination module, for the scene categories with the number of scenes higher than the distribution mean, the corresponding scene-related features are interpolated and summed with the scene-related features of other category scenes to generate the recombined features, specifically including:
[0103] ;
[0104] is the recombined feature, is the interpolation ratio, is the scene-related feature of the current scene, is the scene-related feature of other scenes, is the scene-general feature of the current scene.
[0105] In one embodiment, the present invention further provides a computer device. The computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store the data used in the above method. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the above method.
[0106] In an exemplary embodiment, the present invention further provides a computer-readable storage medium including instructions, such as a memory including instructions, and the above instructions can be executed by a processor to complete the above method. The storage medium may be a computer-readable storage medium. For example, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0107] In one of the embodiments, as Figure 3a and Figure 3b shown, the present invention also conducts simulations of left-turn autonomous driving tasks, respectively using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention. As Figure 3a shown, by using the autonomous driving motion planning method in the present invention, the autonomous driving task in the left-turn situation is successfully completed; as Figure 3b shown, the baseline method is affected by overfitting of the long-tail distribution, resulting in an inability to make a positive decision (not moving).
[0108] In one of the embodiments, as Figure 4a and Figure 4b shown, the present invention also conducts simulations of straight-line autonomous driving tasks, respectively using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention. As Figure 4a shown, by using the autonomous driving motion planning method in the present invention, the autonomous driving task in the straight-line situation is successfully completed; as Figure 4b shown, the baseline method crashes during the straight-line autonomous driving task.
[0109] In one of the embodiments, as Figure 5a and Figure 5b shown, the present invention also conducts simulations of pedestrian avoidance autonomous driving tasks, respectively using the autonomous driving motion planning method and the baseline method (PLUTO) in the present invention. As Figure 5aAs shown, by using the autonomous driving motion planning method in the present invention, the autonomous driving task of pedestrian avoidance is successfully completed, pedestrians are actively avoided, and positive decisions are made after the pedestrians leave; as Figure 5b shown, the vehicle in the baseline method collides with the pedestrian.
[0110] In one embodiment, as Figure 6a and Figure 6b shown, the present invention also conducts simulations of right-turn autonomous driving tasks, using the autonomous driving motion planning method in the present invention and the baseline method (PLUTO) respectively. As Figure 6a shown, by using the autonomous driving motion planning method in the present invention, the autonomous driving task in the right-turn situation is successfully completed; as Figure 6b shown, affected by overfitting of the long-tail distribution, the baseline method fails to make a positive decision (staying still).
[0111] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any regard, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.
[0112] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. An autonomous driving motion planning method, characterized in that, The training process of the adopted planning model includes: Receiving the perception results, which include the ego-vehicle state, the historical information of surrounding agents, and map elements; inputting the perception results into a Transformer-based encoder for feature encoding to obtain encoded features; Performing scene feature deconstruction and recombination on the encoded features: predicting the scene type of the encoded features through a scene classifier, calculating the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories, deconstructing the encoded features into scene-related features and scene-general features, specifically including: Calculating the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories: ; Represents the relevance between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, , represents the operation of taking the mean, represents the encoded feature of the i-th feature dimension, represents the loss gradient corresponding to the encoded feature of the i-th feature dimension; represents the dimension, and the operation of taking the mean is performed at which means averaging the values of all elements in the encoded feature to obtain the contribution value of each feature dimension of the encoded feature; ; Indicates in proportion Obtain the threshold for encoding feature deconstruction, Indicates the quantile function, which returns the value at that proportion position in the sorted dataset; Indicates the total feature dimension number of the encoded feature, Indicates the proportional parameter of the quantile; After that, according to the ratio Obtain the threshold for encoding feature deconstruction, and deconstruct the encoding feature Decompose it according to the following formula: ; Represent scene-related features, represent general scene features, denote the indicator function, , , denote element-wise multiplication; For the scene categories with the number of scenes higher than the distribution mean, interpolating and summing the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features, specifically including: ; is a recombination feature, is the ratio of interpolation, is a scene-related feature of the current scene, is a scene-related feature of other scenes, is a scene-general feature of the current scene; Inputting the recombined features into a trajectory decoder, outputting the future motion trajectory of the ego-vehicle and the predicted trajectories of surrounding agents, and jointly optimizing the parameters of the planning model through the imitation loss and the motion prediction loss of surrounding agents.
2. The autonomous driving motion planning method according to claim 1, characterized in that The joint optimization of the parameters of the planning model through the imitation loss and the motion prediction loss of surrounding agents specifically includes: Imitation loss is as follows: ; Represents the total planned time, Represents the position of the ego vehicle's planned estimated trajectory at the t-th moment, Represents the position of the ego vehicle's planned true trajectory at the t-th moment, Represents the calculation of the Manhattan distance; Motion prediction loss of surrounding agents is as follows: ; Indicates the total number of predicted agents, Indicates the position of the estimated trajectory of the nth agent's motion prediction at the t-th time point, Indicates the position of the true trajectory of the nth agent's motion prediction at the t-th time point; Through the total loss Optimize the parameters of the planning model.
3. The method for autonomous driving motion planning according to claim 1, wherein The ego-vehicle state includes the position, heading angle, speed, and acceleration of the ego-vehicle; the historical information of surrounding agents includes the position, heading angle, speed, acceleration, angular velocity, and bounding box size of each agent; the map elements include the global route of vehicle travel, as well as the traffic light detection results, cross-road condition perception results, and lane line perception results of the current frame.
4. An automatic driving motion planning device, characterized in that, Including: An encoding module: receiving the perception results, which include the ego-vehicle state, the historical information of surrounding agents, and map elements; Inputting the perception results into a Transformer-based encoder for feature encoding to obtain encoded features; A feature deconstruction and recombination module: performing scene feature deconstruction and recombination on the encoded features, predicting the scene type of the encoded features through a scene classifier, calculating the cross-entropy loss to obtain the loss gradient corresponding to the encoded features; according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories, deconstructing the encoded features into scene-related features and scene-general features; for the scene categories with the number of scenes higher than the distribution mean, interpolating and summing the corresponding scene-related features with the scene-related features of other category scenes to generate recombined features; A trajectory decoding module, inputting the recombined features into a trajectory decoder, outputting the future motion trajectory of the ego-vehicle and the predicted trajectories of surrounding agents, and jointly optimizing the parameters of the planning model through the imitation loss and the motion prediction loss of surrounding agents; In the feature deconstruction and recombination module, the deconstruction of the encoded features into scene-related features and scene-general features according to the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories specifically includes: Calculating the relevance between the loss gradients corresponding to the encoded features of different feature dimensions and the scene categories: ; Indicates the relevance between the loss gradient corresponding to the encoded feature of the i-th feature dimension and the scene category, Indicates the operation of taking the mean, Indicates the encoded feature of the i-th feature dimension, Indicates the loss gradient corresponding to the encoded feature of the i-th feature dimension; Indicates dimension, at When performing the operation of taking the mean, it means averaging the values of all elements in the encoded feature to obtain the contribution value of the feature dimension of each encoded feature; ; Indicates in proportion Obtain the threshold for encoding feature deconstruction, Indicates the quantile function, which returns the value at that proportion position in the sorted dataset; Indicates the total feature dimension of the encoded feature, Indicates the proportion parameter of the quantile After that, according to the ratio obtain the threshold for deconstructing the encoded features, and deconstruct the encoded features by disassembling according to the following formula: ; represent scene-related features, represent general scene features, denote an indicator function, , , denote element-wise multiplication; In the feature deconstruction and recombination module, for the scene categories with the number of scenes higher than the distribution mean, the scene-related features corresponding thereto are interpolated and summed with the scene-related features of other category scenes to generate recombination features, which specifically include: ; is a recombination feature, is the ratio of interpolation, is a scene-related feature of the current scene, is a scene-related feature of other scenes, is a scene-general feature of the current scene.
5. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 3 are implemented.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Automatic iteration method of trajectory prediction model, electronic equipment and storage medium
CN114880842A
Automatic driving motion planning and model training method, device and equipment
CN118734923A
Long-tail pedestrian trajectory prediction method based on subclass balance contrast learning
CN119048556A