Track planning method and device, electronic equipment and automatic driving vehicle
By analyzing the differences between the historical expected motion trajectory of obstacles and the actual motion trajectory, correcting the target model parameters of the autonomous driving system, and optimizing the trajectory planning results, the problem of insufficient accuracy and adaptability of trajectory planning in the existing technology is solved, and the resilience and coordinated driving capabilities of the autonomous driving system are improved.
Patent Information
- Application Number
- CN202510307607.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-06
AI Technical Summary
When faced with complex and dynamic traffic environments, existing autonomous driving systems are difficult to ensure the accuracy and adaptability of trajectory planning, and rely on complex mathematical models or large amounts of data processing, resulting in high computing requirements, increased system complexity and cost.
By analyzing the differences between the historical expected motion trajectory of obstacles and the actual motion trajectory, correct the parameters of the target model, optimize the trajectory planning results, and realize posterior analysis and parameter self-correction of autonomous driving decisions and planning.
The adaptability of the autonomous driving system in complex interactive scenarios is improved, ensuring that self-driving vehicles make reasonable and accurate action outputs, so that autonomous driving vehicles can better adapt to complex traffic scenarios and drive in coordination with other vehicles.
Smart Images

Figure CN120096623A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of autonomous driving technology, and specifically to trajectory planning methods, devices, electronic equipment, and autonomous driving vehicles. Background Art
[0002] With the continuous breakthroughs in autonomous driving technology, autonomous driving systems have gradually become an important development direction in the future transportation field. At present, autonomous driving systems mainly rely on the close cooperation of modules such as perception, prediction, decision-making and planning control. They collect environmental information through various sensors, make decisions and plans after a series of algorithm processing, and control the vehicle to drive safely in a complex and changeable traffic environment. Summary of the invention
[0003] The present disclosure provides a trajectory planning method, device, electronic device and autonomous driving vehicle.
[0004] According to one aspect of the present disclosure, a trajectory planning method is provided, including: during the driving of a target vehicle, obtaining driving environment information of a current period, a first expected motion trajectory and an actual motion trajectory of an obstacle in a historical period; wherein the driving trajectory of the target vehicle in the current period is generated based on the first expected motion trajectory; updating parameters of a target model based on a trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that the trajectory deviation between a second expected motion trajectory output by the updated target model and the actual motion trajectory is smaller than the trajectory deviation between the first expected motion trajectory and the actual motion trajectory; and processing the driving environment information using the updated target model to generate a target trajectory, wherein the target trajectory is used to control the target vehicle to travel in a target period.
[0005] According to another aspect of the present disclosure, a trajectory planning device is provided, including: an acquisition module, an update module and a generation module.
[0006] The acquisition module is used to acquire the driving environment information of the current period, the first expected motion trajectory and the actual motion trajectory of the obstacle in the historical period during the driving of the target vehicle; wherein the target trajectory of the target vehicle in the current period is generated based on the first expected motion trajectory.
[0007] An updating module is used to update the parameters of the target model based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that the trajectory deviation between the second expected motion trajectory output by the updated target model and the actual motion trajectory is less than the trajectory deviation between the first expected motion trajectory and the actual motion trajectory.
[0008] The generation module is used to process the driving environment information using the updated target model to generate a target trajectory, and the target trajectory is used to control the target vehicle to travel in a target period.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.
[0010] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising the electronic device described above.
[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.
[0012] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0013] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0015] Figure 1 An exemplary system architecture to which the trajectory planning method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown;
[0016] Figure 2 A flow chart of a trajectory planning method according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 3 A schematic diagram of a framework for implementing a trajectory planning method according to an embodiment of the present disclosure is schematically shown;
[0018] Figure 4A A schematic diagram schematically shows a method of determining a tracing period according to a change in trajectory deviation between a historical expected motion trajectory and an actual motion trajectory of an obstacle according to an embodiment of the present disclosure;
[0019] Figure 4B A schematic diagram schematically shows a method of determining a tracing period according to a change in trajectory deviation between a historical expected motion trajectory and an actual motion trajectory of an obstacle according to another embodiment of the present disclosure;
[0020] Figure 5A schematic diagram schematically shows parameter update of a target model according to an embodiment of the present disclosure;
[0021] Figure 6 A block diagram of a trajectory planning device according to an embodiment of the present disclosure is schematically shown; and
[0022] Figure 7 A block diagram of an electronic device suitable for implementing a trajectory planning method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0023] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] The problem that continues to be solved in the field of autonomous driving technology is to provide autonomous vehicles with safe driving trajectories in complex and dynamically changing traffic scenarios.
[0025] In related examples, trajectory planning is mainly achieved through model-driven, data-driven, and model-data-driven hybrid methods.
[0026] The method of trajectory planning based on mathematical models usually relies on a relatively fixed simulation environment. As the environment changes, the performance of the original model also decreases. In addition, this method has high computational requirements, especially when simulating large-scale environments or performing real-time optimization, which requires powerful computing power, increasing the complexity and cost of the system.
[0027] Although data-driven methods have high flexibility and adaptability, they rely on a large amount of high-quality data for training. For autonomous driving systems, this means that a large amount of sensor and driving decision data needs to be collected and processed. In the actual environment, there may be problems such as incomplete data and noise interference, which will affect the accuracy and reliability of the data, and thus affect the learning effect of the model. Secondly, this method usually lacks interpretability. Since it relies on black box models such as deep learning and reinforcement learning, it is difficult to trace the root cause of the problem when the system fails or does not behave as expected, and there is a lack of effective explanation and feedback mechanisms. In addition, the training of data-driven methods usually requires a lot of time and computing resources. Especially when environmental conditions change dramatically, the training process may require frequent updates, which increases the operation and maintenance costs of the system.
[0028] The hybrid method based on model and data needs to find a balance between model and data in order to provide a good starting point for the data-driven part. The design and adjustment of this balance are very complex. Over-reliance on the model may lead to the lack of sufficient adaptability of the system in the face of dynamic environments, while over-reliance on data may lead to reduced interpretability and controllability of the model. Secondly, this method requires a large amount of data to supplement the deficiencies of the model, especially in rare situations, a large amount of real-time data is required for optimization and updating. Therefore, the system needs to continuously collect and process a large amount of sensor data, which will increase the burden of data processing and storage, resulting in increased resource consumption of the system. Finally, although the hybrid method can improve the robustness of the system to a certain extent, its implementation is complex and the input information of different modules (models and data) may conflict. How to design an effective fusion mechanism has become a difficulty.
[0029] In view of this, the embodiments of the present disclosure provide a trajectory planning method, which corrects the parameters of the model based on the difference between the expected motion trajectory and the actual motion trajectory of the obstacle in the historical period, realizes the posteriori analysis of the trajectory planning results, optimizes the accuracy of autonomous driving decisions and planning, and completes the parameter self-correction of the autonomous driving system, thereby improving the adaptability of the autonomous driving system in complex interactive scenarios, ensuring that the self-driving vehicle makes reasonable and accurate action outputs, and enabling the autonomous driving vehicle to better adapt to complex traffic scenarios and travel in coordination with other vehicles.
[0030] Compared with the trajectory planning method based on mathematical models in related examples, the embodiments of the present disclosure can effectively reduce the dependence on complex mathematical models. By analyzing historical data, it can learn from previous decisions and feedback and optimize the current decision plan without relying entirely on large-scale environmental models calculated in real time, thereby avoiding the challenges of model obsolescence and the need for frequent updates in traditional model-based methods.
[0031] Compared with the data-driven trajectory planning method in the related examples, the disclosed embodiment has stronger interpretability. By backtracking and analyzing historical decision frames, the system can clearly track the reasons and evolution process behind the decision, facilitate problem troubleshooting, and thus improve the system's adjustability and maintenance efficiency.
[0032] Compared with the trajectory planning method based on the combination of mathematical models and data-driven in related examples, the embodiment of the present disclosure provides a more natural way to adjust the decision-making and planning process. Historical data optimizes decision-making planning through continuous feedback, ensuring that the system can more efficiently cope with complex and dynamic driving environments, while reducing the burden of data processing and storage.
[0033] Figure 1 An exemplary system architecture to which the trajectory planning method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.
[0034] It should be noted that Figure 1 The examples shown are only examples of system architectures to which the embodiments of the present disclosure can be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, an exemplary system architecture to which the trajectory planning method and apparatus can be applied may include a terminal device, but the terminal device may implement the trajectory planning method and apparatus provided by the embodiments of the present disclosure without interacting with the server.
[0035] like Figure 1 As shown, the system architecture 100 according to this embodiment may include an autonomous driving vehicle 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the autonomous driving vehicle 101 and the server 103. The network 102 may include a wireless communication link, etc.
[0036] A variety of sensors may be configured on the autonomous driving vehicle 101 to collect driving environment information. The collected driving environment information is then sent to the server 103 via the network 102. The server 103 may generate a target trajectory by executing the trajectory planning method of the embodiment of the present disclosure, and send the target trajectory to the autonomous driving vehicle 101 to control the autonomous driving vehicle to travel along the target trajectory.
[0037] The trajectory planning method provided in the embodiment of the present disclosure may also be generally executed by a processor configured on the automatic driving device 101. Accordingly, the trajectory planning device provided in the embodiment of the present disclosure may also be provided in the automatic driving vehicle 101.
[0038] Several relatively complex driving scenarios are listed below, all of which can be applied to the trajectory planning method provided in the embodiments of the present disclosure.
[0039] In scenario A, the autonomous driving vehicle 101 is traveling in a straight line on the road, and there are vehicles in front and behind the autonomous driving vehicle 101 that are preparing to enter the lane where the autonomous driving vehicle 101 is located.
[0040] In scenario B, the autonomous driving vehicle 101 is preparing to change lanes. In the lane where the autonomous driving vehicle 101 is expected to travel, there are vehicles traveling both in front of and behind the autonomous driving vehicle 101.
[0041] In scenario C, the autonomous driving vehicle 101 drives to the intersection area, and the autonomous driving vehicle 101 intends to go straight through the intersection. In the opposite lane, there is a vehicle turning left and intends to pass the lane where the autonomous driving vehicle 101 is located, and there is also a vehicle intending to turn right and enter the lane where the autonomous driving vehicle 101 is located.
[0042] In scenario D, the autonomous driving vehicle 101 drives to the intersection area and plans to turn right. There are vehicles in both the opposite lane and the left lane that are planning to enter the same lane as the autonomous driving vehicle 101.
[0043] In the above four driving scenarios, any deviation in the target trajectory provided for the autonomous driving vehicle may lead to the risk of collision with other vehicles. Since the trajectory planning method provided in the embodiment of the present disclosure is generated by accurately determining the expected driving trajectories of other vehicles that affect the safe driving of the autonomous driving vehicle 101, the autonomous driving vehicle 101 can safely cooperate with other vehicles on the road according to the target trajectory generated by the trajectory planning method provided in the embodiment of the present disclosure.
[0044] It should be understood that Figure 1 The number of autonomous driving vehicles, networks, and servers in the embodiment is only for illustration purposes. Any number of autonomous driving vehicles, networks, and servers may be provided as required.
[0045] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0046] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0047] Figure 2 The flowchart of the trajectory planning method according to the embodiment of the present disclosure is schematically shown.
[0048] like Figure 2 As shown, the method 200 includes operations S210 to S230.
[0049] In operation S210, during the driving of the target vehicle, driving environment information of the current period, and a first expected motion trajectory and an actual motion trajectory of an obstacle in a historical period are obtained.
[0050] In operation S220, the parameters of the target model are updated based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that the trajectory deviation between the second expected motion trajectory output by the updated target model and the actual motion trajectory is smaller than the trajectory deviation between the first expected motion trajectory and the actual motion trajectory.
[0051] In operation S230, the driving environment information is processed using the updated target model to generate a target trajectory, which is used to control the target vehicle to travel in a target period.
[0052] According to an embodiment of the present disclosure, the target vehicle may be an autonomous driving vehicle. The driving environment information represents information that affects the driving decision of the autonomous driving vehicle, such as lane line information, location information of static obstacles, road state information, current driving state information of the autonomous driving vehicle, driving state information of dynamic obstacles, etc. Static obstacles include but are not limited to various forms of roadblocks, green belts, etc. Dynamic obstacles may be various vehicles traveling on the road, including but not limited to various vehicles.
[0053] The obstacles in the disclosed embodiments represent dynamic obstacles that affect the driving decision of the autonomous vehicle, for example, vehicles in the lane where the autonomous vehicle is located, in adjacent lanes, or vehicles that may enter or leave the lane where the autonomous vehicle is located by changing lanes, or vehicles that are currently driving in a lane that the autonomous vehicle may enter or leave by changing lanes.
[0054] According to an embodiment of the present disclosure, the first expected motion trajectory may be generated using a target model based on the motion information of the obstacle collected by the sensor configured on the autonomous driving vehicle. The actual motion trajectory may be the real motion information of the obstacle in the historical period. The driving trajectory of the target vehicle in the current period is planned based on the first expected motion trajectory.
[0055] In some embodiments, the target model may be a data processing module configured on the autonomous driving system for generating trajectories. The disclosed embodiments do not specifically limit the model.
[0056] In some embodiments, while the target vehicle is driving, the tracing time can be configured to calculate the deviation between the expected motion trajectory and the actual motion trajectory of the obstacle in the historical period in real time, and the parameters of the target model can be updated based on the calculation results to reduce the deviation between the expected motion trajectory and the actual motion trajectory of the obstacle output by the updated target model.
[0057] For example, the tracing time can be 10 seconds. At the 20th second, the target vehicle can trace back the first expected motion trajectory and the actual motion trajectory of the obstacle between the 10th and 20th seconds, and calculate the trajectory deviation. For example, the trajectory deviation can be 0.2. Then, update the parameters of the target model. Since the first expected motion trajectory of the obstacle between the 10th and 20th seconds is generated based on the driving state of the obstacle between the 1st and 5th seconds, the driving state of the obstacle between the 1st and 5th seconds can be input into the updated target model, the second expected motion trajectory can be output, and the trajectory deviation of 0.02 can be calculated. Similarly, the parameters of the target model can be updated multiple times until the trajectory deviation converges.
[0058] Then, the driving environment information is processed using the updated target model to generate a target trajectory for controlling the target vehicle to travel in a target period of time.
[0059] According to an embodiment of the present disclosure, the target period may be the next time period relative to the current period. The length of the time period may be configured according to actual needs, for example, it may be 5 seconds.
[0060] For example, the motion state of the obstacle in the current period (20s to 25s) can be input into the updated model, and the expected motion trajectory of the obstacle in the 25s to 30s can be output. Based on the expected motion trajectory and the driving state of the target vehicle in the current period, the target trajectory of the target vehicle in the 25s to 30s is generated to control the target vehicle to drive along the target trajectory in the 25s to 30s, thus achieving safe coordinated driving with the obstacle.
[0061] According to the embodiments of the present disclosure, the parameters of the model are corrected based on the difference between the expected motion trajectory and the actual motion trajectory of the obstacle in the historical period, so as to realize the a posteriori analysis of the trajectory planning results, optimize the accuracy of the autonomous driving decision and planning, and complete the parameter self-correction of the autonomous driving system, thereby improving the adaptability of the autonomous driving system in complex interactive scenarios, ensuring that the self-driving vehicle makes reasonable and accurate action outputs, and enabling the autonomous driving vehicle to better adapt to complex traffic scenarios and travel in coordination with other vehicles.
[0062] This is only an exemplary embodiment, but is not limited thereto and may also include other trajectory planning methods known in the art, as long as it can implement a posteriori analysis based on historical data of obstacles to automatically correct the model parameters of the autonomous driving system.
[0063] Reference below Figure 3~Figure 5 , combined with specific embodiments Figure 2 The method shown is further explained.
[0064] Figure 3 A schematic diagram of a framework for implementing a trajectory planning method according to an embodiment of the present disclosure is schematically shown.
[0065] like Figure 3 As shown, the framework 300 may include: a historical data storage and analysis module 310 , a decision planning result posterior module 320 , a parameter updating module 330 and a user-defined interface 340 .
[0066] The historical data storage and analysis module 310 is used to continuously collect and store the driving environment information 301, vehicle driving information 302, obstacle movement information 303 and decision planning results 304 at each moment during the driving of the target vehicle.
[0067] For example: The target vehicle driving information can be represents, including the actual driving position of the autonomous vehicle ( )、Actual speed( )、actual acceleration( ) and other information, as well as the planned position output by the autonomous driving system ( )、Planning speed( ) and planning acceleration ( )wait.
[0068] Obstacle driving information is provided by Indicates that it covers the actual movement trajectory of the obstacle ( )、Actual speed( )、actual acceleration( ) and other data, as well as the predicted trajectory of the obstacle's future motion state during the global decision-making and planning process of the autonomous driving system ( ), prediction speed ( )、Predicted acceleration( ) and other data.
[0069] As the target vehicle continues to interact with surrounding obstacles, data can be continuously updated and stored. Each decision-making planning result can be stored in real time and associated with the driving status information at the previous moment to form time series data. These time series data not only provide accurate references for subsequent decision-making processes, but also provide necessary historical data support for the autonomous driving system to perform posterior analysis, anomaly detection, parameter self-correction and other functions.
[0070] As time goes by, the system will gradually accumulate a large amount of historical frame data to form a rich historical data pool 311, as shown in formula (1).
[0071] (1)
[0072] In order to further improve the data processing efficiency in the subsequent parameter update process, the historical data processing unit 312 can be used to perform data analysis and preprocessing on the historical data in the historical data pool 101. The preprocessing process includes but is not limited to unified data format, data filtering and smoothing, data normalization and feature engineering.
[0073] There are differences in the data formats collected and output by different sensors and system modules. In order to integrate heterogeneous data into a standard format that the system can process, the original historical data can be uniformly converted. Taking obstacle motion information as an example, when collecting the motion information of each obstacle, the motion trajectory obtained by the SL (longitudinal axis-lateral axis) coordinate system may be established with the obstacle as a reference. In order to ensure that the motion information of all traffic participants can be compared and analyzed in the same framework, all autonomous driving vehicles and obstacle motion information involved in autonomous driving decision-making planning need to be converted to the same coordinate system.
[0074] Data filtering can use methods such as mean filtering and Kalman filtering to remove instantaneous noise caused by defects in sensors and autonomous driving systems, thereby providing more accurate data collection results. The smoothing process further reduces discontinuities in the data through methods such as low-pass filtering and Bayesian smoothing, making the data more coherent, which not only helps to improve data stability, but also avoids interference of unrealistic mutations on the autonomous driving system.
[0075] Data normalization is to map data of different scales and units into a unified standardized numerical range, thereby eliminating dimensional differences and ensuring that each feature is processed at the same scale to facilitate subsequent data analysis and parameter tuning. In the posterior analysis of autonomous driving decision planning, normalized data can provide more stable and consistent input, which is helpful for subsequent self-correction of model parameters.
[0076] Feature engineering is a crucial step in data preprocessing. Its purpose is to extract features that are beneficial to decision planning from raw data, thereby improving the performance of the autonomous driving system. Feature engineering involves not only the processing of raw sensor data, but also the analysis and conversion of historical frame time series data. First, according to the driving scenario, screen the most critical features for the target task, eliminate redundant and irrelevant features, and reduce computational complexity. Secondly, according to the nature of the task, fuse different features to achieve new feature construction, providing richer information for subsequent decision posteriors and parameter self-correction. Finally, feature encoding is performed to meet the decision planning needs of special scenarios, further enhancing the environmental understanding ability of the autonomous driving system.
[0077] In some embodiments, by recording the key information such as the perception data, vehicle status information, and decision execution output collected by the vehicle in real time during driving, a comprehensive historical data repository is built to ensure that each frame of data collected can be accurately recorded and archived. Data processing is completed based on data smoothing, feature engineering, etc., providing a reference basis for subsequent decision planning and optimization, ensuring that the decision system can be reasonably optimized based on real historical behavior.
[0078] The output result of the historical data storage and analysis module 310 is the filtered and smoothed historical decision planning and real data, which can be directly used as the input of the decision planning result posterior module 320.
[0079] The decision planning result posterior module 320 is used to perform post-analysis on the decision planning results generated by the autonomous driving system. Its function is to trace back the decision planning path based on historical data and driving status to complete the evaluation of the rationality of the decision planning.
[0080] The decision planning result posterior module 320 may include: a historical data extraction unit 321, a decision planning effect quantitative analysis unit 322 and a decision planning rationality evaluation unit 323.
[0081] The historical data extraction unit 321 can extract historical data of a predetermined dimension from the historical data according to a predetermined tracing time range, for example, 4 to 10 seconds. The predetermined dimension can be determined based on the relevance to the decision planning result.
[0082] The decision planning effect quantitative analysis unit 322 can be used to perform quantitative analysis on the decision planning effect by calculating the deviation between the expected motion trajectory and the actual motion trajectory of the obstacle.
[0083] The decision planning rationality evaluation unit 323 can be used to combine historical data and various quantitative analysis results obtained by calculation to comprehensively analyze the decision planning results of the autonomous driving system, and judge the rationality and effectiveness of the decision planning results in practical applications. Rationality evaluation not only focuses on whether the decision planning meets the expected goals, but also takes into account the complexity of the environment, emergencies, and the performance of the system in real scenarios. For example: thresholds for driving safety and smoothness can be set according to specific driving environments, and safety and rationality can be judged to determine whether there is a risk of collision or uneven operations such as transitional braking.
[0084] In some embodiments, the decision planning execution and posterior analysis links combine real-time perception and historical data to generate and execute driving decisions based on the input historical data and current sensor feedback. During the decision execution process, the posterior analysis module is responsible for evaluating the effect of real-time decisions, and comparing the current execution results with historical data to determine the rationality and feasibility of the strategy. If deviations occur, the system will provide immediate feedback and adjust the strategy to make the decision process more accurate and flexible.
[0085] The parameter update module 330 is used to automatically correct the system's parameter deviations by analyzing historical data, decision-making planning posterior results and actual execution conditions, so as to improve the stability and adaptability of the autonomous driving vehicle.
[0086] The parameter updating module 330 may include: a loss / reward calculation unit 331 , a gradient back propagation unit 332 , and a parameter updating unit 333 .
[0087] For example, the loss / reward calculation unit 331 can be used to determine the loss value or reward value based on the deviation between the expected motion trajectory and the actual motion trajectory of the obstacle in the historical period. The gradient back propagation unit 332 can be used to adjust the model parameters by gradient ascent until the loss value or reward value converges. The parameter updating unit 333 can be used to update the model parameters to the parameters when the loss value or reward value converges.
[0088] In some embodiments, the parameter self-correction link relies on the combination of historical data and real-time feedback. By analyzing the deviations in the decision-making and planning process, the system's control parameters are automatically adjusted to ensure that the optimal performance of the autonomous driving system is always maintained. In order to avoid the accumulation of deviations caused by dynamic changes in the environment, the self-correction mechanism continuously monitors the execution status and feedback data of the autonomous driving system, uses adaptive algorithms to identify potential problems, and makes corrections based on historical data.
[0089] The user-defined interface module 340 is used to meet the application requirements of different scenarios, provide flexible customized function configuration and analysis tools, and allow developers to adjust the decision-making planning system parameters and set evaluation criteria according to specific intelligent driving scenarios, thereby achieving highly personalized autonomous driving decision-making planning and analysis. The information that can be customized by users includes but is not limited to the time window length 341, evaluation indicators and thresholds 342, and loss / reward functions.
[0090] In some embodiments, an intuitive and easy-to-use graphical operation interface can be provided for users, so that developers and users can operate and configure the decision planning system in a graphical manner. The interface usually includes drag-and-drop decision planning and evaluation module components, a visual control panel, and real-time feedback functions, so that developers and debuggers can intuitively set decision planning evaluation parameters, safety and smoothness thresholds, etc. through visual operations.
[0091] In some embodiments, a programmatic interface can also be provided for developers, allowing them to customize and expand the system through programming. The API (Application Programming Interface) can be used to integrate with other systems or adjust the parameters of the decision-making planning module through scripting. Through the open API, developers can flexibly call functional modules in the system, perform advanced operations, obtain real-time feedback and historical data, and perform automated testing.
[0092] In some embodiments, developers can also use flexible configuration files to finely control the system's settings and parameters in text or code. These configuration files are usually in a standard format, such as JSON, XML, or YAML. R&D personnel can define the operating parameters, evaluation criteria, algorithm tuning, and other details of the decision-making planning system in the file.
[0093] Since the driving environment of the target vehicle is complex and changeable during driving, in order to adapt to the changes in the driving environment and adjust the driving trajectory in time, the predetermined time can be configured through the above-mentioned user-defined interface module 340 to reduce the amount of data processing during the parameter update process.
[0094] In some embodiments, before updating the parameters of the target model based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, the method also includes the following operations: determining a target historical period from the historical period based on a predetermined duration; and extracting a target expected motion trajectory and a target actual motion trajectory for the target historical period from the first expected motion trajectory and the actual motion trajectory, respectively.
[0095] According to an embodiment of the present disclosure, the predetermined duration may include a lower duration limit and an upper duration limit, for example, 4 to 10 seconds, where 4 seconds is the lower duration limit and 10 seconds is the upper duration limit.
[0096] When the duration that an obstacle participates in the driving decision of the autonomous vehicle is less than the lower limit of the duration, the retrospective duration can be dynamically increased forward according to the duration difference to satisfy the requirement of using historical time series data of at least the length of the lower limit of the duration for parameter update.
[0097] For example, if you start tracing back from the 5th second, but the obstacle appears in the driving environment of the autonomous vehicle from the 3rd second, then you can continue tracing back 1 second, that is, the target historical period can be from the 1st second to the 5th second.
[0098] When the obstacle is involved in the driving decision of the autonomous vehicle for a period longer than the upper limit, only the time corresponding to the upper limit can be traced back to meet the requirement of using historical time series data of at most the length of the upper limit for parameter update.
[0099] For example, if you trace back from the 20th second, and the obstacle appears in the driving environment of the autonomous vehicle from the 1st second, then you can only trace back 10 seconds, that is, the target historical period can be from the 10th second to the 20th second.
[0100] Then, the target expected motion trajectory from the 10th second to the 20th second is extracted from the first expected motion trajectory, and the target actual motion trajectory from the 10th second to the 20th second is extracted from the actual motion trajectory.
[0101] According to the embodiments of the present disclosure, by configuring the traceability period of historical data, redundant data in the parameter update process is reduced, the data processing efficiency is further improved, the timeliness of real-time correction of parameters during the operation of the target vehicle is improved, and the accuracy of trajectory planning is improved.
[0102] In some embodiments, in addition to the predetermined duration, the target historical period may be determined by comprehensively considering the change between the expected movement trajectory and the actual movement trajectory of the obstacle within the historical period.
[0103] For example, a historical period may include T moments, where T is an integer greater than 1. Determining a target historical period from the historical period based on a predetermined duration may include the following operations: determining the Tth moment as the start moment of tracing back; determining the end moment of tracing back based on the trajectory deviation of each moment and the predetermined duration; and determining the period between the start moment of tracing back and the end moment of tracing back as the target historical period.
[0104] According to an embodiment of the present disclosure, the trajectory deviation at each moment characterizes the trajectory deviation between the first expected motion trajectory and the actual motion trajectory. The trajectory deviation between the first motion trajectory and the actual motion trajectory includes these multiple trajectory deviations at each moment. For example: the historical period may include T moments, the first expected motion trajectory includes T expected trajectory points corresponding to the T moments, and the actual motion trajectory includes T actual trajectory points corresponding to the T moments. The trajectory deviation at the tth moment can be expressed as the absolute value of the difference between the expected trajectory point at the tth moment and the actual trajectory point at the tth moment, or it can be expressed as the ratio of the absolute value of the difference between the expected trajectory point at the tth moment and the actual trajectory point at the tth moment to the expected trajectory point at the tth moment.
[0105] In some embodiments, determining the tracing termination moment based on the trajectory deviation at each moment and the predetermined duration may include the following operations: in response to determining that the trajectory deviation at the t-th moment is greater than or equal to a predetermined threshold, and the time difference between the t-th moment and the T-th moment is greater than or equal to a predetermined duration, determining the moment whose time difference with the T-th moment is the predetermined duration as the tracing termination moment, t=1,…,T.
[0106] Figure 4A A schematic diagram of determining a tracing period according to a change in trajectory deviation between a historical expected motion trajectory and an actual motion trajectory of an obstacle according to an embodiment of the present disclosure is schematically shown.
[0107] like Figure 4A As shown in the curve 400A showing the change of trajectory deviation over time, since the trajectory deviation at the tjth moment is equal to the predetermined threshold, the trajectory deviation at the tjth moment is traced back to the tjth moment for a predetermined period of time. 1 At time t, determine 1 The time is the end time of retrospective.1 The period between time t and time T is the target historical period.
[0108] In some embodiments, determining the tracing termination moment based on the trajectory deviation at each moment and a predetermined duration may include the following operations: in response to determining that the trajectory deviation at the tth moment is greater than or equal to a predetermined threshold, and the time difference between the tth moment and the Tth moment is less than a predetermined duration, determining the tth moment as the tracing termination moment.
[0109] Figure 4B A schematic diagram of determining a tracing period according to a change in trajectory deviation between a historical expected motion trajectory and an actual motion trajectory of an obstacle according to another embodiment of the present disclosure is schematically shown.
[0110] like Figure 4B As shown in the curve 400B showing the change of trajectory deviation over time, the t 1 The time between time t and time t is the predetermined time. i The trajectory deviation at time t is greater than the predetermined threshold, so it can be determined that i The time is the end time of retrospection. i The period from moment t to moment T is determined as the target historical period.
[0111] By combining the deviation between the expected motion trajectory of the obstacle and the time motion trajectory and the predetermined duration to determine the target historical period, on the one hand, the redundant data in the parameter update process can be reduced; on the other hand, the influence of data with large trajectory deviation on the convergence speed or convergence quality in the parameter update process can be reduced, further improving the efficiency of parameter update.
[0112] According to an embodiment of the present disclosure, updating the parameters of the target model based on the first expected motion trajectory and the actual motion trajectory may include the following operations: generating a trajectory deviation based on the objective function and the first expected motion trajectory and the actual motion trajectory; and obtaining an updated target model by updating the parameters of the target model based on the trajectory deviation.
[0113] According to the embodiments of the present disclosure, the objective function includes but is not limited to a cost function, an error function, etc. Based on the trajectory deviation, the deviation between the output expected motion trajectory and the actual motion trajectory can be made smaller and smaller by continuously updating the parameters of the target model until convergence, thereby obtaining an updated target model.
[0114] According to the embodiments of the present disclosure, by comparing the expected motion trajectory of the obstacle with the actual motion trajectory and adjusting the model parameters, it is possible to accurately identify the potential problems of the target model adaptive algorithm, and perform corrections and updates in a timely manner, thereby further improving the accuracy of the decision-making planning results for the target trajectory in the next time period.
[0115] In some embodiments, when there are expected movement trajectories and actual movement trajectories of obstacles for each historical moment in the historical period, an expected position for each historical moment can be obtained from the first expected movement trajectory; an actual position for each historical moment can be obtained from the actual movement trajectory; and based on the objective function, a trajectory deviation can be generated by processing the difference between each expected position and each actual position.
[0116] For example, the trajectory deviation can be calculated using formula (2).
[0117] (2)
[0118] Among them, La represents the trajectory deviation; T represents the historical duration; represents the horizontal coordinate of the actual position of the obstacle at time t, The ordinate represents the actual position of the obstacle at time t; represents the horizontal coordinate of the expected position of the obstacle at time t, The ordinate represents the expected position of the obstacle at time t.
[0119] In some embodiments, when there are expected motion trajectories and actual motion trajectories of obstacles at some historical moments in the historical period, for example, the target historical period is from the 5th to the 10th second, but the obstacle participates in the driving decision from the 7th second. The expected end position can be obtained from the first expected motion trajectory; the actual end position can be obtained from the actual motion trajectory; and based on the objective function, the trajectory deviation can be generated by processing the difference between the expected end position and the actual end position.
[0120] For example, the trajectory deviation can be calculated using formula (3).
[0121] (3)
[0122] Wherein, Lb represents the trajectory deviation; The horizontal coordinate of the actual position of the obstacle at the end time of the target historical period, The ordinate of the actual position of the obstacle at the end time of the target historical period; The horizontal coordinate of the expected obstacle position at the end time of the target historical period, The ordinate of the expected position of the obstacle at the end of the target historical period.
[0123] According to an embodiment of the present disclosure, a trajectory deviation calculation method is adaptively selected based on the status of historical data of obstacles within a target historical period, so that the calculation result of the trajectory deviation can truly reflect the deviation of the output result of the current autonomous driving system, further improving the adaptability and flexibility to environmental changes.
[0124] In actual application scenarios, there may be multiple obstacles that affect the driving decisions of autonomous vehicles. In order to accurately output the expected motion trajectory of each obstacle, the model parameters can be independently adjusted based on the historical data of each obstacle, further reducing the impact of the accumulated deviations caused by dynamic changes in the environment on the accuracy of trajectory planning.
[0125] According to an embodiment of the present disclosure, there are N obstacles, where N is an integer greater than 1; based on the trajectory deviation, by updating the parameters of the target model, an updated target model is obtained, which may include the following operations: based on the trajectory deviation for the nth obstacle, by updating the parameters of the target model, an intermediate model is obtained, n=1,...N-1; and based on the trajectory deviation for the Nth obstacle, by updating the parameters of the intermediate model, an updated model is obtained.
[0126] Figure 5 A schematic diagram schematically shows parameter updating of a target model according to an embodiment of the present disclosure.
[0127] like Figure 5 As shown, in the embodiment 500, the T1 period is the historical period, and the T2 period is the current period. T1a 501a input parameter is ω 0 Model T510 outputs the expected motion trajectory Ta of obstacle Oa 1 502a.
[0128] Then, based on the expected motion trajectory Ta of the obstacle Oa 1 502a and the actual movement trajectory Aa of the obstacle Oa 1 503a Calculate the trajectory deviation L 1 411, and based on the trajectory deviation L 1 411 Adjust the parameters of model T510 until L 1 Convergence, at this time, the parameters of model T510 are updated to ω 1 .
[0129] Next, the obstacle Ob motion state M T1b 501b input parameter is ω 1 Model T520 outputs the expected motion trajectory Tb of obstacle Oa 1 502b.
[0130] Then, based on the expected motion trajectory Tb of the obstacle Oa1 502b and the actual movement trajectory Ab of the obstacle Ob 1 503b Calculate trajectory deviation L 2 421, and based on the trajectory deviation L 2 421 Adjust the parameters of model T520 until L 2 Convergence, at this time, the parameters of model T520 are updated to ω 2 .
[0131] In some embodiments, the trajectory deviation can be used as the loss value, and the loss value can be reduced by adjusting the model parameters until the loss value converges. The parameter adjustment process is shown in formula (4):
[0132] (4)
[0133] in, represents the model parameters; represents the loss function; Indicates the predetermined coefficient.
[0134] In some embodiments, a reward function can be constructed with trajectory deviation as an independent variable, and the reward value can be increased by adjusting the model parameters until the reward value converges. The parameter adjustment process is shown in formula (5):
[0135] (5)
[0136] in, represents the model parameters; represents the reward function; Indicates the predetermined coefficient.
[0137] According to an embodiment of the present disclosure, the model parameters are independently adjusted based on the historical data of each obstacle, further reducing the impact of the accumulated deviations caused by dynamic changes in the environment on the accuracy of trajectory planning.
[0138] In some embodiments, the driving environment information may include the driving status of N obstacles and the driving status of the target vehicle; processing the driving environment information using the updated target model to generate a target trajectory may include the following operations: using the intermediate model to process the driving status of the nth obstacle to generate an expected driving trajectory of the nth obstacle in the target period; using the updated target model to process the driving status of the Nth obstacle to generate an expected driving trajectory of the Nth obstacle in the target period; and using the updated target model to process the N expected driving trajectories of the N obstacles in the target period and the driving status of the target vehicle to generate a target trajectory.
[0139] like Figure 5 As shown, we can use the parameter ω 1The model T520 handles T 2 Time period obstacle Oa motion state M T2a 504a outputs the expected motion trajectory Ta of the obstacle Oa in the next period 2 505a. Using the parameter ω 2 The model T530 handles T 2 Time period obstacle Ob motion state M T2b 504b outputs the expected motion trajectory Tb of the obstacle Ob in the next period of time 2 505b.
[0140] Finally, using the parameter ω 2 The expected motion trajectory Ta of the model T530 to handle the obstacle Oa 2 505a, expected motion trajectory Tb of obstacle Ob 2 505b and the driving state M of the target vehicle in the current period T2 501, output the target trajectory 502 of the target vehicle in the next period.
[0141] According to an embodiment of the present disclosure, a model with independently updated parameters is used to output the expected motion trajectory of the next time period for each obstacle, which further improves the accuracy of the expected movement trend of the obstacle, thereby improving the accuracy of the target trajectory planning of the autonomous driving vehicle.
[0142] Figure 6 A block diagram of a trajectory planning device according to an embodiment of the present disclosure is schematically shown.
[0143] like Figure 6 As shown, the device 600 may include an acquisition module 610 , an update module 620 and a generation module 630 .
[0144] The acquisition module 610 is used to acquire the driving environment information of the current period, the first expected motion trajectory and the actual motion trajectory of the obstacle in the historical period during the driving of the target vehicle; wherein the target trajectory of the target vehicle in the current period is planned based on the first expected motion trajectory.
[0145] The updating module 620 is used to update the parameters of the target model based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that the trajectory deviation between the second expected motion trajectory output by the updated target model and the actual motion trajectory is smaller than the trajectory deviation between the first fishing area motion trajectory and the actual motion trajectory.
[0146] The generation module 630 is used to process the driving environment information using the updated target model to generate a target trajectory, and the target trajectory is used to control the target vehicle to travel in a target period.
[0147] According to an embodiment of the present disclosure, the above-mentioned device also includes: a determination module and an extraction module.
[0148] The determination module is used to determine a target historical period from the historical periods based on a predetermined duration.
[0149] The extraction module is used to extract the target expected motion trajectory and the target actual motion trajectory for the target historical period from the first expected motion trajectory and the actual motion trajectory respectively.
[0150] According to an embodiment of the present disclosure, the historical period includes T moments, where T is an integer greater than 1; the determination module includes: a first determination submodule, a second determination submodule, and a third determination submodule.
[0151] The first determining submodule is used to determine the Tth moment as the tracing start moment.
[0152] The second determination submodule is used to determine the tracing end time based on the trajectory deviation at each moment and the predetermined time length. The trajectory deviation at each moment represents the deviation between the first expected motion trajectory and the actual motion trajectory.
[0153] The third determination submodule is used to determine the period between the tracing start time and the tracing end time as the target historical period.
[0154] According to an embodiment of the present disclosure, the second determining submodule includes a first determining unit and a second determining unit.
[0155] The first determining unit is used to determine the tth moment as the tracing termination moment in response to determining that the trajectory deviation at the tth moment is greater than or equal to a predetermined threshold and the time difference between the tth moment and the Tth moment is less than a predetermined time length, t=1,…,T.
[0156] The second determining unit is used to determine the moment whose time difference with the Tth moment is the predetermined time length as the tracing termination moment in response to determining that the trajectory deviation at the tth moment is greater than or equal to a predetermined threshold and the time difference between the tth moment and the Tth moment is greater than or equal to a predetermined time length.
[0157] According to an embodiment of the present disclosure, the update module may include a first generation submodule and an update submodule.
[0158] The first generating submodule is used to generate a trajectory deviation based on the objective function according to the first expected motion trajectory and the actual motion trajectory.
[0159] The updating submodule is used to obtain an updated target model by updating the parameters of the target model based on the trajectory deviation.
[0160] According to an embodiment of the present disclosure, the first generating submodule includes a first acquiring unit, a second acquiring unit and a first generating unit.
[0161] The first acquisition unit is used to acquire the expected position for each historical moment from the first expected motion trajectory.
[0162] The second acquisition unit is used to acquire the actual position at each historical moment from the actual motion trajectory.
[0163] The first generating unit is used to generate a trajectory deviation by processing a difference between each expected position and each actual position based on the target function.
[0164] According to an embodiment of the present disclosure, the first generating submodule includes a third acquiring unit, a fourth acquiring unit and a second generating unit.
[0165] The third acquisition unit is used to acquire the expected terminal position from the first expected motion trajectory.
[0166] The fourth acquisition unit is used to acquire the actual end position from the actual motion trajectory.
[0167] The second generating unit is used to generate a trajectory deviation by processing the difference between the expected terminal position and the actual terminal position based on the objective function.
[0168] According to an embodiment of the present disclosure, the updating submodule includes: a first updating unit and a second updating unit.
[0169] The first updating unit is used to obtain an intermediate model by updating parameters of the target model based on the trajectory deviation for the nth obstacle, n=1,...N-1.
[0170] The second updating unit is used to obtain an updated model by updating parameters of the intermediate model based on the trajectory deviation of the Nth obstacle.
[0171] According to an embodiment of the present disclosure, the driving environment information includes the driving states of N obstacles and the driving state of the target vehicle. The generation module includes: a second generation submodule, a third generation submodule and a fourth generation submodule.
[0172] The second generating submodule is used to process the driving state of the n-th obstacle by using the intermediate model to generate the expected driving trajectory of the n-th obstacle in the target period.
[0173] The third generating submodule is used to process the driving state of the Nth obstacle by using the updated target model to generate the expected driving trajectory of the Nth obstacle in the target period.
[0174] The fourth generation submodule is used to process the N expected driving trajectories of the N obstacles in the target time period and the driving state of the target vehicle by using the updated target model to generate a target trajectory.
[0175] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, an autonomous driving vehicle, a readable storage medium, and a computer program product.
[0176] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.
[0177] According to an embodiment of the present disclosure, an autonomous driving vehicle includes the electronic device described above.
[0178] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0179] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method as described above.
[0180] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0181] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0182] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0183] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as the trajectory planning method. For example, in some embodiments, the trajectory planning method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the trajectory planning method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the trajectory planning method in any other appropriate manner (e.g., by means of firmware).
[0184] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0185] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0186] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0188] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0189] A computer system may include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises through computer programs running on respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server combined with a blockchain.
[0190] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0191] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A trajectory planning method, comprising: During the driving of the target vehicle, the driving environment information of the current period, the first expected motion trajectory and the actual motion trajectory of the obstacle in the historical period are obtained; wherein the driving trajectory of the target vehicle in the current period is generated based on the first expected motion trajectory; updating the parameters of the target model based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that the trajectory deviation between the second expected motion trajectory output by the updated target model and the actual motion trajectory is less than the trajectory deviation between the first expected motion trajectory and the actual motion trajectory; and The driving environment information is processed using the updated target model to generate a target trajectory, and the target trajectory is used to control the target vehicle to travel in the target time period.
2. The method according to claim 1, before updating the parameters of the target model based on the trajectory deviation between the first expected motion trajectory and the actual motion trajectory, the method further comprises: Determining a target historical period from the historical periods based on a predetermined duration; as well as The target expected motion trajectory and the target actual motion trajectory for the target historical period are extracted from the first expected motion trajectory and the actual motion trajectory respectively.
3. The method according to claim 2, wherein: The historical period includes T moments, where T is an integer greater than 1; The determining the target historical period from the historical periods based on the predetermined duration includes: Determine the Tth moment as the starting moment of tracing back; Determine a traceback termination time based on the trajectory deviation at each moment and the predetermined time length; the trajectory deviation at each moment represents the trajectory deviation between the first expected motion trajectory and the actual motion trajectory; and The period between the tracing back start time and the tracing back end time is determined as the target historical period.
4. The method according to claim 3, wherein: The step of determining the tracing termination time based on the trajectory deviation at each moment and the predetermined time duration includes: In response to determining that the trajectory deviation at the tth moment is greater than or equal to a predetermined threshold, and the time difference between the tth moment and the Tth moment is less than the predetermined duration, determining the tth moment as the tracing termination moment, t=1, ..., T; and In response to determining that the trajectory deviation at the tth moment is greater than or equal to the predetermined threshold, and the time difference between the tth moment and the Tth moment is greater than or equal to the predetermined duration, a moment whose time difference with the Tth moment is the predetermined duration is determined as the retrospective termination moment.
5. The method according to claim 1, wherein: The updating of the parameters of the target model based on the first expected motion trajectory and the actual motion trajectory includes: Based on the objective function, generating a trajectory deviation according to the first expected motion trajectory and the actual motion trajectory; Based on the trajectory deviation, the updated target model is obtained by updating the parameters of the target model.
6. The method according to claim 5, wherein: The generating a trajectory deviation based on the objective function and according to the first expected motion trajectory and the actual motion trajectory includes: Obtaining an expected position for each historical moment from the first expected motion trajectory; Obtaining the actual position for each of the historical moments from the actual motion trajectory; and The trajectory deviation is generated by processing the difference between each of the expected positions and each of the actual positions based on the objective function.
7. The method according to claim 5, wherein generating the trajectory deviation according to the first expected motion trajectory and the actual motion trajectory comprises: Acquire an expected terminal position from the first expected motion trajectory; Obtaining an actual end position from the actual motion trajectory; as well as The trajectory deviation is generated by processing a difference between the expected end position and the actual end position based on the objective function.
8. The method according to claim 5, wherein: The number of obstacles is N, where N is an integer greater than 1; The step of obtaining the updated target model by updating the parameters of the target model based on the trajectory deviation includes: Based on the trajectory deviation for the nth obstacle, an intermediate model is obtained by updating the parameters of the target model, n=1,...N-1; as well as Based on the trajectory deviation for the Nth obstacle, the updated model is obtained by updating the parameters of the intermediate model.
9. The method according to claim 8, wherein: The driving environment information includes the driving status of N obstacles and the driving status of the target vehicle; The step of processing the driving environment information using the updated target model to generate a target trajectory includes: Processing the driving state of the nth obstacle by using the intermediate model to generate an expected driving trajectory of the nth obstacle in the target time period; Processing the driving state of the Nth obstacle using the updated target model to generate an expected driving trajectory of the Nth obstacle in the target period; and The updated target model is used to process the N expected driving trajectories of the N obstacles in the target time period and the driving state of the target vehicle to generate the target trajectory.
10. A trajectory planning device, comprising: An acquisition module, used to acquire driving environment information of a current period, a first expected motion trajectory and an actual motion trajectory of an obstacle in a historical period during the driving of the target vehicle; wherein the target trajectory of the target vehicle in the current period is generated based on the first expected motion trajectory; an updating module, configured to update parameters of a target model based on a trajectory deviation between the first expected motion trajectory and the actual motion trajectory, so that a trajectory deviation between a second expected motion trajectory output by the updated target model and the actual motion trajectory is smaller than a trajectory deviation between the first expected motion trajectory and the actual motion trajectory; and A generation module is used to process the driving environment information using the updated target model to generate a target trajectory, and the target trajectory is used to control the target vehicle to travel in the target time period.
11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. An autonomous driving vehicle comprising the electronic device described in claim 11.
13. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
14. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.