Vehicle driving path planning method and device
Through the scene trajectory generator combined with unsupervised training and driving preference model to optimize path planning, the problem of path planning deviation in the existing technology is solved, and the path planning is more accurate and in line with driver's preferences is achieved, and the ride experience is improved.
Patent Information
- Application Number
- CN202510282428.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, in vehicle path planning, the driving scenario data is limited and the driver's preferences are not taken into account, resulting in path planning deviations, and the passenger's riding experience is poor.
The scene trajectory generator is used to combine unsupervised training models and reward models related to driving preferences to generate path driving trajectories, and virtual scenes are optimized through iterative training to improve path planning accuracy and driver/occupant experience.
Improve the accuracy of path planning, match the generated path driving trajectory with the driving scene and meet driving preferences, and improve the driver/occupant's riding experience.
Smart Images

Figure CN120333479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of transportation, and particularly to a method and device for vehicle driving path planning. Background Art
[0002] Whether it is full-automatic vehicle driving, assisted automatic driving assisted by a driver, or driver-operated driving, the in-vehicle system can plan a driving path for the vehicle so that the vehicle can drive smoothly under relatively poor road conditions. Currently, although a large amount of data such as driving scenario data and vehicle driving trajectory data is used to train a model for planning a driving path for the vehicle, due to the limited driving scenario data and vehicle driving trajectory data, and the training process does not involve the preferences of drivers, there are deviations in path planning, resulting in a poor riding experience for passengers. Summary of the Invention
[0003] In view of this, the present invention provides a method, device and vehicle for vehicle driving path planning, which can improve the accuracy of path planning for different scenarios, and can give a path driving trajectory according to the driving preferences of the driver, so as to improve the riding experience of the driver / passenger.
[0004] To solve the above technical problems, the present invention provides the following technical solutions:
[0005] In a first aspect, the present invention provides a method for vehicle driving path planning, including:
[0006] Determine the driving scenario in which the vehicle is located;
[0007] Use a scenario trajectory generator to generate a path driving trajectory for the driving scenario, where the scenario trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on the real scenario and the virtual scenario extended by the scenario trajectory generator, and in combination with a reward model related to driving preferences;
[0008] Control the vehicle to drive according to the path driving trajectory.
[0009] Optionally, the above method for vehicle driving path planning further includes:
[0010] Iteratively execute the following steps N1 to N3 until an iteration stop condition is met to obtain a scenario trajectory generator:
[0011] Step N1: Use the current scenario data to train an unsupervised model, and use the trained model to generate a virtual scenario. Generate multiple sets of predicted driving trajectories for the real scenario and the virtual scenario respectively. For the first training cycle, the current scenario data is the real scenario. For non-first training cycles, the current scenario data is the real scenario and the corrected virtual scenario obtained in the previous training cycle;
[0012] Step N2: Use a scenario trajectory evaluator to evaluate the virtual scenario, multiple sets of predicted driving trajectories corresponding to the real scenario, and multiple sets of predicted driving trajectories corresponding to the virtual scenario;
[0013] Step N3: According to the evaluation results, correct the virtual scenario and adjust the unsupervised training model.
[0014] Optionally, generating multiple sets of predicted driving trajectories for the real scenario and the virtual scenario in Step N1 includes:
[0015] Generate a predicted driving trajectory according to the current execution action of the vehicle, the average value of the vehicle state before the current execution action, the state noise variance corresponding to the vehicle state, the cumulative state of the noise, and the noise variance of the current execution action.
[0016] Optionally, the generating of the predicted driving trajectory includes:
[0017] Generate a predicted driving trajectory using the following first calculation formula;
[0018] First calculation formula:
[0019] p[a(k)∣a(k - 1)] = N[a(k); μ(k), σ(k)^2×I]
[0020] Where p[a(k)∣a(k - 1)] represents the change of the vehicle's action from the (k - 1)-th step to the k-th step; a(k) represents the vehicle's action at the k-th step; a(k - 1) represents the vehicle's action at the (k - 1)-th step; μ(k) represents the average value of the vehicle state before the k-th step; σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; N[] represents the Gaussian distribution formula; I represents an adjustable parameter related to the driving trajectory.
[0021] Optionally, the above vehicle driving path planning method further includes:
[0022] Use the following second calculation formula to calculate the average value of the vehicle state before the k-th step;
[0023] μ(k) = alpha(k - 1)×belta(k)×a(0)+alpha(k - 1)×a(k)
[0024] Among them, μ(k) represents the average value of the vehicle state before the k-th step and the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; a(0) represents the starting action of the vehicle; a(k) represents the action of the vehicle at the k-th step.
[0025] Optionally, the above vehicle driving path planning method further includes: calculating the state noise variance of the vehicle state at the k-th step by using the following third calculation formula;
[0026] σ(k)^2 = [1 - alpha(k - 1)] × belta(k) / [1 - alpha(k)]
[0027] σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; alpha(k) represents the cumulative noise state before the k-th step and the k-th step.
[0028] Optionally, the above vehicle driving path planning method further includes:
[0029] By using the vehicle driving scenario, the vehicle driving trajectory corresponding to the vehicle driving scenario, and the driving preference corresponding to the vehicle driving trajectory, a scene trajectory evaluator is generated by supervising and iteratively training a reward model related to the driving preference.
[0030] Optionally, in the above step N2, evaluating the multiple sets of predicted driving trajectories corresponding to the real scenario and the multiple sets of predicted driving trajectories corresponding to the virtual scenario includes:
[0031] Evaluating the multiple sets of predicted driving trajectories corresponding to the real scenario and the multiple sets of predicted driving trajectories corresponding to the virtual scenario by using the following fourth calculation formula;
[0032] Fourth calculation formula:
[0033]
[0034] P(R(s i ) > R(s j )) represents the probability that the i-th predicted driving trajectory in the scenario S (real scenario or virtual scenario) is better than the j-th predicted driving trajectory; R(s i ) represents the score of the i-th predicted driving trajectory of the scenario S by the scene trajectory evaluator; R(s j ) represents the score of the j-th predicted driving trajectory of the scenario S by the scene trajectory evaluator.
[0035] Optionally, correcting the virtual scene in N3 includes: eliminating virtual scenes that do not meet a preset evaluation threshold.
[0036] Optionally, the above vehicle driving path planning method further includes:
[0037] Based on reinforcement learning and the evaluation result of the scenario trajectory evaluator on the scenario trajectory generator, adjusting the scenario trajectory generator.
[0038] Optionally, the above vehicle driving path planning method further includes:
[0039] Deploying the scenario trajectory generator to the in-vehicle system;
[0040] Based on the in-vehicle system, performing the step of using the scenario trajectory generator to generate a path driving trajectory for the driving scenario.
[0041] Optionally, the above vehicle driving path planning method further includes:
[0042] Deploying the scenario trajectory evaluator to the in-vehicle system;
[0043] In response to optimizing the path planning, using driving-related data including driver preferences generated by the vehicle to adjust the scenario trajectory evaluator deployed in the in-vehicle system;
[0044] Using the adjusted scenario trajectory evaluator to evaluate at least one set of predicted driving trajectories generated by the scenario trajectory generator;
[0045] According to the evaluation result, adjusting the scenario trajectory generator deployed in the in-vehicle system.
[0046] In a second aspect, an embodiment of the present invention provides a vehicle driving path planning device, including: a scenario determination module, a trajectory generation module, and a control module, where
[0047] The scenario determination module is configured to determine the driving scenario in which the vehicle is located;
[0048] The trajectory generation module is configured to use a scenario trajectory generator to generate a path driving trajectory for the driving scenario, where the scenario trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on the real scenario and the virtual scenario extended by the scenario trajectory generator, and in combination with a reward model related to driving preferences;
[0049] The control module is configured to control the vehicle to drive according to the path driving trajectory.
[0050] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0051] One or more processors;
[0052] A storage device for storing one or more programs,
[0053] When the one or more programs are executed by the one or more processors, the one or more processors implement the vehicle driving path planning method provided in the embodiment of the first aspect as described above.
[0054] In a fourth aspect, an embodiment of the present invention provides a computer-readable medium, on which a computer program for a vehicle driving path planning method is stored.
[0055] When the computer program is executed by an in-vehicle processor, it implements the vehicle driving path planning method provided in the embodiment of the first aspect as described above.
[0056] In a fifth aspect, an embodiment of the present invention provides a vehicle that can implement the vehicle driving path planning method provided in the embodiment of the first aspect or includes the vehicle driving path planning device provided in the embodiment of the second aspect as described above.
[0057] The technical solution of the above invention has the following advantages or beneficial effects:
[0058] The technical solution provided by the embodiment of the present invention generates a corresponding path driving trajectory for the driving scenario in which the vehicle is located based on a scenario trajectory generator. Since this scenario trajectory generator is based on real scenarios and extended virtual scenarios through an unsupervised training model (i.e., more and wider scenario data is introduced in the training process), and the scenario trajectory evaluator related to driving preferences is used to adjust the virtual scenario and the unsupervised training model, that is, this scenario trajectory generator is obtained based on more scenario training, improving the accuracy of path planning for different scenarios, and introducing driving preferences to adjust it during the training process, so that the path driving trajectory generated by this scenario trajectory generator not only matches the driving scenario, but also can meet the driving preferences, improving the driving experience of the driver / occupant. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is an exemplary system architecture diagram to which an embodiment of the present invention can be applied;
[0060] Figure 2 is a main flow schematic diagram of the vehicle driving path planning method provided by an embodiment of the present invention;
[0061] Figure 3 is a main flow schematic diagram of iteratively training a scenario trajectory generator provided by an embodiment of the present invention;
[0062] Figure 4It is a partial structural schematic diagram of a vehicle driving path planning device of the first structure provided by an embodiment of the present invention;
[0063] Figure 5 It is a partial structural schematic diagram of a vehicle driving path planning device of the second structure provided by an embodiment of the present invention;
[0064] Figure 6 It is a structural schematic diagram of a computer system suitable for implementing the vehicle driving path planning of an embodiment of the present invention. Detailed implementation manners
[0065] The technical solution provided by the embodiment of the present invention for various driving scenarios of the vehicle can be driving scenarios from different dimensions. Exemplarily, the driving scenario can be different scenarios under the weather dimension, such as heavy rain weather, strong wind weather, heavy snow weather, foggy weather, cold weather, hot weather, etc., or it can be scenarios under the road dimension, such as potholed roads, icy roads, rural roads, highway sections, single-lane roads, multi-lane roads, roads located in residential areas, roads passing by schools, scenic area roads, etc., or it can also be scenarios under the traffic dimension, such as traffic congestion, smooth traffic, traffic light intersections, highway exits, crowded sections of people, etc. In addition, the driving scenario of the vehicle can also be a combination of the above-mentioned driving scenarios under multiple dimensions.
[0066] Due to different driving scenarios, the vehicle is required to have different driving paths, such as avoiding potholed areas of the road, passing through traffic congestion areas with green waves, avoiding the influence of heavy rain, heavy snow, etc. In order to be able to plan more reasonable driving paths for the above various different driving scenarios and improve the accuracy of path planning for different scenarios, the embodiment of the present invention provides a vehicle driving path planning method and device.
[0067] It should be noted that the technical solution provided by the embodiment of the present invention can be applied to the driving path planning of other intelligent agents (such as floor cleaning robots, drones, etc.) in addition to the vehicle driving path planning.
[0068] The following describes the exemplary embodiments of the present invention with reference to the accompanying drawings, including various details of the embodiments of the present invention to help understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted below.
[0069] It should be pointed out that, without conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0070] Based on the system architecture on which the technical solution provided by the embodiments of the present invention depends, the technical scenarios applicable to the technical solution provided by the embodiments of the present invention will be described below.
[0071] Figure 1 An exemplary system architecture 100 on which the vehicle driving path planning method or the vehicle driving path planning device according to the embodiments of the present invention can be applied is shown.
[0072] As Figure 1 shown, the system architecture 100 may include a backend service system 101, a network 102, and a vehicle system 103. Information interaction is carried out between the backend service system 101 and the vehicle system 103 through the network 102. Specifically, the backend service system 101 can implement some or all of the technical solutions in the vehicle driving path planning method or the vehicle driving path planning device provided by the embodiments of the present invention by being mounted on a server, a service cluster, or a cloud server. Exemplarily, the backend service system 101 can train an unsupervised training model based on a real scenario and a virtual scenario extended by a scenario trajectory generator, and adjust the virtual scenario and the unsupervised training model by combining a scenario trajectory evaluator trained with a reward model related to driving preferences to generate a scenario trajectory generator, and provide the scenario trajectory generator for the vehicle system 103 through the network 102. Correspondingly, the vehicle system 103 determines the driving scenario in which the vehicle is located; uses the scenario trajectory generator to generate a path driving trajectory for the driving scenario, and controls the vehicle to drive according to the path driving trajectory. Additionally, the process of determining the driving scenario in which the vehicle is located; using the scenario trajectory generator to generate a path driving trajectory for the driving scenario can also be completed by the backend service system 101. Correspondingly, the backend service system 101 sends a control instruction to the vehicle system 103 through the network 102 to control the vehicle to drive according to the path driving trajectory.
[0073] Furthermore, the vehicle system 103 may include various systems, such as a driving control system 1031, a power system 1032, a sensor system 1033, a control system 1034, an assisted lane change system 1035, one or more peripheral devices 1036, a power supply 1037, a computer system 1038, and a user interface 1039. Among them, the vehicle driving path planning method provided by the embodiments of the present invention can be realized by controlling the interaction between the above-mentioned various systems. Optionally, the vehicle system 103 may include more or fewer systems, and each system may include multiple components. In addition, each system and component of the vehicle system 103 can be interconnected by wire or wirelessly.
[0074] Among them, the vehicle system 103 includes a driving control system 1031, and the driving control system 1031 can be in a full or partial autonomous driving mode. For example, the driving control system 1031 can automatically control the vehicle to travel along a specific trajectory without interacting with a person, according to the driving trajectory signal or driving trajectory instruction provided by the control system 1034.
[0075] The power system 1032 may include components that provide power movement for the vehicle. For example, the power system 1032 may include an engine, an energy source, a transmission, wheels, tires, etc. Among them, the engine can be an internal combustion engine, an electric motor, an air compression engine, or other types of engine combinations, such as a hybrid engine composed of a gasoline engine and an electric motor, or a hybrid engine composed of an internal combustion engine and an air compression engine. The engine converts the energy source into mechanical energy and provides it to the transmission. Examples of the energy source can include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other power sources. The energy source can also provide energy for other systems of the vehicle. In addition, the transmission may include a gearbox, a differential, a drive shaft, a clutch, etc.
[0076] The sensor system 1033 may include sensors that sense the surrounding environment of the vehicle (such as sensors that sense whether there are obstacles around, etc.) and pressure sensors that sense whether there are passengers on the seat. For example, a positioning system (the positioning system can be a global positioning system (GPS) system), or it can be a Beidou system or other positioning systems), radar, a laser rangefinder, an inertial measurement unit (IMU), and a camera, etc. The positioning system can be used to locate the geographical position of the vehicle. The IMU is used to sense the position and orientation changes of the vehicle based on inertial acceleration. In one embodiment, the IMU can be a combination of an accelerometer and a gyroscope. Radar can use radio signals to sense objects within the surrounding environment of the vehicle. In some embodiments, in addition to sensing objects, radar can also be used to sense the speed and / or forward direction of the objects, etc.
[0077] Among them, in order to detect environmental information, objects, etc. outside the vehicle, a camera, etc. can be configured at an appropriate position outside the vehicle. For example, in order to obtain an environmental image of the side of the vehicle, the camera can be on the side mirror of the vehicle. The camera can be a static or video camera.
[0078] The control system 1034 may include a software system for implementing vehicle driving path planning or vehicle driving control. For example, a system for analyzing the vehicle's surrounding environment, a system for pre-tightening seat belts, a route planning system, an obstacle avoidance system, a vision system for image analysis, etc. The control system 1034 may also include hardware systems such as the throttle, steering wheel system, seat belt system, airbag system, and peripheral devices (such as projection devices, displays, etc.). Additionally, the control system 1034 may additionally or alternatively include components other than those shown and described. Or some of the above-described components may be reduced.
[0079] Further, as described above, in the case where the background service system 101 generates a virtual scene through an unsupervised training model based on the real scene and the virtual scene extended by the scene trajectory generator, and adjusts the virtual scene and the unsupervised training model by combining a reward model related to driving preferences to generate a scene trajectory generator, the control system 1034 may further implement a part of the above-described vehicle driving path planning method, determine the driving scene in which the vehicle is located; use the scene trajectory generator to generate a path driving trajectory for the driving scene; and control the vehicle to drive according to the path driving trajectory. Additionally, during the process of controlling the vehicle to drive along the planned driving path through the control system 1034, the drive assist lane change system 1035 outputs a lane change signal or a lane change command to the driving control system 1031, and the driving control system 1031 automatically controls the vehicle to change lanes according to the lane change signal or lane change command output by the assist lane change system 1035.
[0080] Additionally, the control system 1034 may also interact with external sensors, other autonomous driving devices, other computer systems, or users through the peripheral device 1036. The peripheral device 1036 may include a wireless communication system, an in-vehicle computer, a microphone and / or a speaker, a camera, and a projector, etc.
[0081] In some embodiments, the peripheral device 1036 provides a means for the user of the control system 1034 to interact with the user interface. For example, the in-vehicle computer may provide information to the user of the vehicle. The user interface may also operate the in-vehicle computer to receive user input. The in-vehicle computer may be operated through a touch screen. In other cases, the peripheral device may provide a means for communicating with other devices located inside the vehicle. For example, the microphone may receive audio (such as voice commands or other audio inputs) from the user of the control system. Similarly, the speaker may output audio to the user of the control system.
[0082] A wireless communication system can communicate wirelessly with one or more devices directly or via a communication network. For example, the wireless communication system can communicate using networks such as cellular networks, WiFi, and wireless local area network (WLAN), or can communicate directly with devices using infrared links, Bluetooth, or ZigBee. Other wireless protocols, such as various autonomous driving communication systems, etc.
[0083] The power supply 1037 can supply power to various components of the vehicle. The power supply 1037 can be a rechargeable lithium-ion or lead-acid battery.
[0084] In the case where some technical means for implementing the vehicle driving path planning method are realized in the vehicle system 103, part or all of the functions for realizing the vehicle driving path planning are controlled by the computer system 1038. The computer system 1038 can include at least one processor, and the processor executes instructions stored in a non-transitory computer-readable medium such as a memory. The computer system 1038 provides part of the execution code for the above control system to realize the vehicle driving path planning.
[0085] The processor can be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, the processor can be a special device such as an application specific integrated circuits (ASIC) or other hardware-based processors. Those of ordinary skill in the art should understand that the processor, computer, or memory can actually include multiple processors, computers, or memories that may or may not be stored in the same physical housing. For example, the memory can be a hard disk drive or other storage media located in a housing different from the computer. Therefore, the reference to a processor or computer will be understood to include a reference to a collection of processors or computers or memories that may or may not operate in parallel. Different from using a single processor to execute the steps described herein, some components such as the steering component and the deceleration component can each have its own processor, and the processor only executes determinations related to functions specific to the component.
[0086] The user interface 1039 is used to provide information to or receive information from the user of the vehicle. Optionally, the user interface 1039 can include one or more input / output devices within the set of peripheral devices 1036, such as a wireless communication system, an in-vehicle computer, a microphone, and a speaker.
[0087] It should be understood that the above components are only examples. In actual applications, the components in the above-mentioned modules or systems may be added or deleted according to actual needs. Figure 1It should not be construed as a limitation on the embodiments of the present application.
[0088] Figure 2 It is a schematic diagram of the main steps of a vehicle driving path planning method according to an embodiment of the present invention. Specifically, as Figure 2 shown, the vehicle driving path planning method may further include the following steps:
[0089] Step S201: Determine the scenario: Determine the driving scenario in which the vehicle is located.
[0090] This step can analyze the driving scenario in which the vehicle is located through the road information where the vehicle is located, such as snow-covered sections, waterlogged sections or potholed roads, the weather information of the environment where the vehicle is located, such as heavy rain, heavy snow, low temperature or high temperature, and the traffic state where the vehicle is located, such as traffic light intersections, road congestion or highway exits, etc. The driving scenario can be obtained by combining the road information, weather information and traffic state where the vehicle is located.
[0091] Step S202: Generate a driving trajectory: Use a scenario trajectory generator to generate a path driving trajectory for the driving scenario. Among them, the scenario trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on the real scenario and the virtual scenario extended by the scenario trajectory generator, and combining a reward model related to driving preferences.
[0092] Step S203: Control the vehicle: Control the vehicle to drive according to the path driving trajectory.
[0093] It should be noted that the above path driving trajectory can be a driving trajectory planning for a short-distance section, such as how to drive on a potholed section to avoid potholes on the potholed section, such as for a waterlogged section or a snow-covered section, on the waterlogged section or the snow-covered section, how to avoid waterlogged positions or snow-covered positions or how to control the vehicle speed, the rotation of the vehicle steering wheel, etc., so that the vehicle drives smoothly on the waterlogged section or the snow-covered section. In addition, the above path driving trajectory can also be a path planning for the vehicle from the starting position to the target position, such as which road to choose among multiple roads.
[0094] In addition, the above vehicle driving path planning method can meet the needs of autonomous driving and can also meet the needs of drivers participating in driving. Specifically, for autonomous driving, controlling the vehicle to drive according to the path driving trajectory means that the interaction between the vehicle's various systems completes the vehicle's driving according to the path driving trajectory; for the driver participating in driving, the vehicle gives the path driving trajectory to drive, so that the driver drives and operates the vehicle according to the path driving trajectory.
[0095] For Figure 2The provided technical solution generates a corresponding path driving trajectory for the driving scenario where the vehicle is located based on a scenario trajectory generator. Since this scenario trajectory generator is based on real scenarios and extended virtual scenarios through an unsupervised training model (i.e., more and wider scenario data is introduced during the training process), and combines a reward model related to driving preferences to train a scenario trajectory evaluator to adjust the virtual scenario and the unsupervised training model, that is, this scenario trajectory generator is trained based on more scenarios, improving the accuracy of path planning for different scenarios, and introducing driving preferences to adjust it during the training process, so that the path driving trajectory generated by this scenario trajectory generator not only matches the driving scenario, but also can meet the driving preferences, enhancing the driving experience of the driver / occupant.
[0096] Furthermore, the above vehicle driving path planning method further includes a process of training the scenario trajectory generator. This process of training the scenario trajectory generator is generally completed on the background service system 101. Specifically, the process of training the scenario trajectory generator included in the vehicle driving path planning method may include the following steps:
[0097] Iteratively execute Figure 3 the steps S301 to S303 shown until the iteration stop condition is met, to obtain the scenario trajectory generator:
[0098] Step S301: Generate a virtual scenario and multiple groups of predicted driving trajectories corresponding to the real scenario and the virtual scenario respectively: Using the current scenario data, through an unsupervised training model, and using the trained model to generate a virtual scenario, and generate multiple groups of predicted driving trajectories for the real scenario and the virtual scenario respectively. Among them, for the first training cycle, the current scenario data is the real scenario, and for non-first training cycles, the current scenario data is the real scenario and the corrected virtual scenario obtained in the previous training cycle.
[0099] Among them, the unsupervised training model refers to training a neural network model such as a convolutional neural network or a recurrent neural network and other models in an unsupervised manner.
[0100] It is worth noting that the current scenario data is marked with scenario identifiers. The virtual scenario can be a driving scenario or a combination of multiple driving scenarios.
[0101] In this step, generating multiple sets of predicted driving trajectories for the real scenario and the virtual scenario respectively means generating multiple sets of predicted driving trajectories for the real scenario and generating multiple sets of predicted driving trajectories for the virtual scenario. Exemplarily, for the real scenario scene_a, this step generates two sets of predicted driving trajectories for the real scenario scene_a, namely traj_1(t) and traj_2(t), that is, [scene_a, traj_1(t)] and [scene_a, traj_2(t)]. For the virtual scenario scene_a', this step generates two sets of predicted driving trajectories for the real scenario scene_a', namely traj_1'(t) and traj_2'(t), that is, [scene_a', traj_1'(t)] and [scene_a', traj_2'(t)].
[0102] Specifically, generating multiple sets of predicted driving trajectories for the real scenario and the virtual scenario respectively in this step may include: generating a predicted driving trajectory based on the current execution action of the vehicle, the average value of the vehicle state before the current execution action, the state noise variance corresponding to the vehicle state, the state accumulation of the noise, and the noise variance of the current execution action. By combining the current execution action of the vehicle, the average value of the vehicle state, and the noises related to the vehicle such as the state noise variance, the state accumulation of the noise, and the noise variance of the current execution action, etc., the accuracy of the predicted driving trajectory can be effectively improved. Among them, the current execution action may include the regulation of vehicle parameters such as the driving speed and the steering angle of the steering wheel of the vehicle. The vehicle state refers to the pose where the vehicle is currently located.
[0103] It is worth noting that in the process of generating multiple sets of predicted driving trajectories for the real scenario and the virtual scenario in this step, map data can be introduced to improve the accuracy of the generated multiple sets of predicted driving trajectories.
[0104] More specifically, generating a predicted driving trajectory may include: using the following calculation formula (1) to generate a predicted driving trajectory;
[0105] p[a(k)∣a(k - 1)] = N[a(k); μ(k), σ(k)^2×I] (1)
[0106] Where, p[a(k)∣a(k - 1)] represents the change of the vehicle's action from the (k - 1)-th step to the k-th step; a(k) represents the vehicle's action at the k-th step; a(k - 1) represents the vehicle's action at the (k - 1)-th step; μ(k) represents the average value of the vehicle state before the k-th step; σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; N[] represents the Gaussian distribution formula; I represents a parameter related to the driving trajectory that can be regulated.
[0107] Among them, the action of the vehicle at the k-th step refers to the attitude or pose that the vehicle needs to achieve by controlling the vehicle parameters during the adjustment at the k-th step.
[0108] This Gaussian distribution formula is an existing Gaussian distribution formula. Exemplarily, Among them, σ(k) represents the square root value of the state noise variance of the vehicle state at the k-th step.
[0109] Among them, the noise of the vehicle state refers to the deviation between the actual state of the vehicle and the theoretical state corresponding to the vehicle parameters when the vehicle is adjusted to the state corresponding to the vehicle parameters by controlling the vehicle parameters.
[0110] Among them, using the following calculation formula (2), calculate the average value of the vehicle state before the k-th step;
[0111] μ(k) = alpha(k - 1)×belta(k)×a(0) + alpha(k - 1)×a(k) (2)
[0112] Among them, μ(k) represents the average value of the vehicle state at the k-th step and before the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; a(0) represents the starting action of the vehicle; a(k) represents the action of the vehicle at the k-th step.
[0113] Furthermore, using the following calculation formula (3), calculate the state noise variance of the vehicle state at the k-th step;
[0114] σ(k)^2 = [1 - alpha(k - 1)] × belta(k) / [1 - alpha(k)] (3)
[0115] σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; alpha(k) represents the cumulative noise state at the k-th step and before the k-th step. The cumulative noise state before the k-th step refers to the sum of the state noises of each step from the 1st step to the k - 1th step; the noise variance at the k-th step refers to the variance between the actual action of the vehicle generated when the vehicle parameters are adjusted to the k-th step and the theoretical action corresponding to the vehicle parameters. The cumulative noise state at the k-th step and before the k-th step refers to the sum of the state noises of each step from the 1st step to the k-th step.
[0116] Step S302: Scene and trajectory evaluation: Use the scene trajectory evaluator to evaluate the virtual scene, multiple groups of predicted driving trajectories corresponding to the real scene, and multiple groups of predicted driving trajectories corresponding to the virtual scene.
[0117] Step S303: Modify the scenario and the model: According to the evaluation results, modify the virtual scenario and adjust the unsupervised training model.
[0118] During the process of training the scenario trajectory generator, by introducing the average vehicle state, state noise variance, etc., the influence of vehicle state noise and vehicle action noise on the driving trajectory is considered during the training process, so as to improve the accuracy of the path driving trajectory generated by the scenario trajectory generator for the driving scenario in the future.
[0119] Furthermore, the above process evaluates the virtual scenario, multiple groups of predicted driving trajectories corresponding to the real scenario, and multiple groups of predicted driving trajectories corresponding to the virtual scenario through the scenario trajectory evaluator, so that the entire evaluation process does not require manual participation. While improving the accuracy of the generated scenario trajectory generator, the efficiency of the generated scenario trajectory generator can be improved.
[0120] Specifically, modifying the virtual scenario may include: eliminating the virtual scenarios that do not meet the preset evaluation threshold. Among them, the preset evaluation threshold is generally a threshold that can be adjusted by the user according to vehicle models, vehicle brands, etc.
[0121] Furthermore, the above vehicle driving path planning method further includes: using the vehicle driving scenario, the vehicle driving trajectory corresponding to the vehicle driving scenario, and the driving preference corresponding to the vehicle driving trajectory, and generating a scenario trajectory evaluator by supervised iterative training of the reward model related to the driving preference. That is to say, the scenario trajectory evaluator generated in the embodiment of the present invention is obtained by training the reward model. The driving preference corresponding to the vehicle driving trajectory is introduced during the training process, so that the driving preference is combined during the evaluation process of the scenario trajectory evaluator, so that the path driving trajectory generated by the scenario trajectory generator can also meet the driving preference, so as to improve the riding experience of the driver / occupant.
[0122] Furthermore, in step S302, the specific implementation manner of evaluating multiple groups of predicted driving trajectories corresponding to the real scenario and multiple groups of predicted driving trajectories corresponding to the virtual scenario may include:
[0123] Using the following calculation formula (4) to evaluate multiple groups of predicted driving trajectories corresponding to the real scenario and multiple groups of predicted driving trajectories corresponding to the virtual scenario;
[0124]
[0125] P(R(s i )>R(s j )) represents the probability that the i-th predicted driving trajectory in the scenario S (real scenario or virtual scenario) is better than the j-th predicted driving trajectory; R(s i) represents the score of the \(i\)-th predicted driving trajectory of the scene trajectory evaluator for scene \(S\); \(R(s j ) represents the score of the \(j\)-th predicted driving trajectory of the scene trajectory evaluator for scene \(S\).
[0126] Through the above process, by comprehensively considering the scores of different predicted driving trajectories of scene \(S\) by the scene trajectory evaluator, it is possible to better evaluate the advantages and disadvantages of different predicted driving trajectories of scene \(S\), so as to better guide the training and adjustment of the scene trajectory evaluator.
[0127] Furthermore, the above vehicle driving path planning method further includes: adjusting the scene trajectory generator based on the evaluation results of the reinforcement learning and the scene trajectory evaluator for the scene trajectory generator. It should be noted that introducing reinforcement learning to adjust the scene trajectory generator can be after adjusting the scene trajectory generator based on the above comprehensive scene trajectory evaluator, that is, after generating the scene trajectory generator, then introducing reinforcement learning to fine-tune the scene trajectory generator, improving the accuracy of the scene trajectory generator while reducing the computational resource overhead of generating the scene trajectory generator.
[0128] Specifically, the scene trajectory generator can be obtained through the iterative Figure 3 steps S301 to S303 given, and directly use this scene trajectory generator to execute the path driving trajectory generation for the driving scene. In this iterative process, a scene trajectory evaluator trained with a reward model related to driving preferences is introduced.
[0129] Alternatively, after obtaining the scene trajectory generator through the iterative Figure 3 steps S301 to S303 given or for the scene trajectory generator currently used by the vehicle, based on the evaluation results of the reinforcement learning and the scene trajectory evaluator for the scene trajectory generator, adjust the scene trajectory generator. This adjustment process is essentially a fine-tuning of the scene trajectory generator, and this fine-tuning process has a relatively small computational resource overhead, which can be completed in the background service system or on the vehicle system.
[0130] Furthermore, the above vehicle driving path planning method further includes: deploying the scene trajectory generator to the in-vehicle system; based on the in-vehicle system, execute the step of using the scene trajectory generator to generate the path driving trajectory for the driving scene. Through this process, the efficiency of generating the path driving trajectory for the driving scene can be improved, so as to timely feedback the required path driving trajectory to the vehicle, making the path driving trajectory match the vehicle driving and ensuring the stable driving of the vehicle.
[0131] In addition, the above scene trajectory generator can also be deployed to the background service system.
[0132] Further, the above vehicle driving path planning method further includes: deploying a scenario trajectory evaluator in the in-vehicle system; in response to optimizing the path planning, using driving-related data generated by the vehicle and including driver preferences to adjust the scenario trajectory evaluator deployed in the in-vehicle system; using the adjusted scenario trajectory evaluator to evaluate at least one set of predicted driving trajectories generated by the scenario trajectory generator; and adjusting the scenario trajectory generator deployed in the in-vehicle system according to the evaluation results. Through this process, the scenario trajectory evaluator deployed in the in-vehicle system can be adjusted based on the driver preferences of the vehicle, so that the scenario trajectory evaluator has personalized adjustments for different vehicles, better conforming to the driving preferences of the vehicle it is in. Subsequently, optimizing the scenario trajectory generator based on the scenario trajectory evaluator can better meet the driver preferences of the vehicle it is in, so that the path driving trajectory can be personalized according to the driver preferences of different vehicles, better meeting the personalized driving needs of different vehicles and enhancing the driving / riding experience of the driver / occupant of the vehicle.
[0133] Further, an embodiment of the present invention provides a vehicle driving path planning device. As Figure 4 and Figure 5 shown, the vehicle driving path planning device may include: a scenario determination module 401, a trajectory generation module 402, and a control module 403, where
[0134] The scenario determination module 401 is configured to determine the driving scenario in which the vehicle is located;
[0135] The trajectory generation module 402 is configured to use a scenario trajectory generator to generate a path driving trajectory for the driving scenario, where the scenario trajectory generator is obtained by adjusting a virtual scenario and an unsupervised training model through an unsupervised training model based on a real scenario and a virtual scenario extended by the scenario trajectory generator, and in combination with a reward model related to driving preferences;
[0136] The control module 403 is configured to control the vehicle to travel according to the path driving trajectory.
[0137] Among them, the trajectory generation module 402 may be located in the background service system 101 or in the vehicle system 103. The scenario determination module 401 and the control module 403 may also be located in the background service system 101 or the vehicle system 103. Exemplarily, as Figure 4 shown, the scenario determination module 401, the trajectory generation module 402, and the control module 403 are all located in the vehicle system 103, as Figure 5 shown, the scenario determination module 401 and the control module 403 are located in the vehicle system 103, and the trajectory generation module 402 is located in the background service system 101. Preferably, as Figure 4 shown, the scenario determination module 401, the trajectory generation module 402, and the control module 403 are all located in the vehicle system 103.
[0138] Further, as Figure 4 and Figure 5 shown, the above vehicle driving path planning device further includes: a training module 404, which is generally located in the background service system 103. Among them,
[0139] The training module 404 is used to iteratively execute the following steps N1 to N3 until the iteration stop condition is met, and obtain a scenario trajectory generator:
[0140] Step N1: Use the current scenario data, pass through an unsupervised training model, and use the trained model to generate a virtual scenario, and generate multiple sets of predicted driving trajectories for the real scenario and the virtual scenario respectively. Among them, for the first training cycle, the current scenario data is the real scenario, and for non-first training cycles, the current scenario data is the real scenario and the corrected virtual scenario obtained in the previous training cycle;
[0141] Step N2: Use the scenario trajectory evaluator to evaluate the virtual scenario, multiple sets of predicted driving trajectories corresponding to the real scenario, and multiple sets of predicted driving trajectories corresponding to the virtual scenario;
[0142] Step N3: According to the evaluation results, correct the virtual scenario and adjust the unsupervised training model.
[0143] Further, the training module 404 is further used to generate a predicted driving trajectory according to the current execution action of the vehicle, the average value of the vehicle state before the current execution action of the vehicle, the state noise variance corresponding to the vehicle state, the state accumulation of the noise, and the noise variance of the current execution action.
[0144] Further, the training module 404 is further used to generate a predicted driving trajectory by using the following calculation formula (1);
[0145] p[a(k)∣a(k - 1)] = N[a(k); μ(k), σ(k)^2×I] (1)
[0146] Among them, p[a(k)∣a(k - 1)] represents the change of the vehicle's action from the (k - 1)-th step to the k-th step; a(k) represents the vehicle's action at the k-th step; a(k - 1) represents the vehicle's action at the (k - 1)-th step; μ(k) represents the average value of the vehicle state before the k-th step; σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; N[] represents the Gaussian distribution formula; I represents an adjustable parameter related to the driving trajectory.
[0147] Further, the training module 404 is further used to calculate the average value of the vehicle state before the k-th step by using the following calculation formula (2);
[0148] μ(k) = alpha(k - 1)×belta(k)×a(0) + alpha(k - 1)×a(k) (2)
[0149] Among them, μ(k) represents the average value of the vehicle state at the k-th step and before the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; a(0) represents the starting action of the vehicle; a(k) represents the action of the vehicle at the k-th step.
[0150] Furthermore, the training module 404 is further configured to calculate the state noise variance of the vehicle state at the k-th step by using the following calculation formula (3);
[0151] σ(k)^2 = [1 - alpha(k - 1)] × belta(k) / [1 - alpha(k)] (3)
[0152] σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; alpha(k) represents the cumulative noise state at the k-th step and before the k-th step.
[0153] Furthermore, the training module 404 is further configured to generate a scenario trajectory evaluator by using the vehicle driving scenario, the vehicle driving trajectory corresponding to the vehicle driving scenario, and the driving preference corresponding to the vehicle driving trajectory, and performing supervised iterative training on the reward model related to the driving preference.
[0154] Furthermore, as Figure 4 and Figure 5 shown, the above vehicle driving path planning device further includes: an evaluation module 405, which can be deployed in the vehicle system 103 as Figure 4 shown, or can be deployed in the background service system 101 as Figure 5 shown. The evaluation module 405 is used to deploy a scenario trajectory evaluator, and the scenario trajectory evaluator uses the following calculation formula (4) to evaluate multiple sets of predicted driving trajectories corresponding to the real scenario and multiple sets of predicted driving trajectories corresponding to the virtual scenario;
[0155]
[0156] P(R(s i )>R(s j )) represents the probability that the i-th predicted driving trajectory in the scenario S (real scenario or virtual scenario) is better than the j-th predicted driving trajectory; R(s i) represents the score of the i-th predicted driving trajectory of the scene S by the scene trajectory evaluator; R(s j ) represents the score of the j-th predicted driving trajectory of the scene S by the scene trajectory evaluator.
[0157] Furthermore, the training module 404 is further configured to eliminate virtual scenes that do not meet the preset evaluation threshold.
[0158] Furthermore, the training module 404 is further configured to adjust the scene trajectory generator based on reinforcement learning and the evaluation result of the scene trajectory generator by the scene trajectory evaluator.
[0159] Furthermore, as Figure 4 and Figure 5 shown, the above vehicle driving path planning device further includes: an interactive deployment module 406, and the interactive deployment module 406 is deployed in the background service system 101. Among them,
[0160] The interactive deployment module 406 is configured to deploy the scene trajectory generator to the in-vehicle system, so as to execute the step of generating a path driving trajectory for the driving scene by using the scene trajectory generator based on the in-vehicle system.
[0161] Furthermore, the interactive deployment module 406 is further configured to deploy the scene trajectory evaluator to the in-vehicle system;
[0162] The evaluation module 405 deployed on the vehicle system 103 responds to the optimized path planning, uses the driving-related data including the driver's preferences generated by the vehicle to adjust the scene trajectory evaluator deployed on the in-vehicle system; uses the adjusted scene trajectory evaluator to evaluate at least one set of predicted driving trajectories generated by the scene trajectory generator; and adjusts the scene trajectory generator deployed on the in-vehicle system according to the evaluation result.
[0163] Furthermore, an embodiment of the present invention further provides an electronic device. The electronic device may include:
[0164] One or more processors;
[0165] A storage device for storing one or more programs,
[0166] When the one or more programs are executed by the one or more processors, the one or more processors implement the vehicle driving path planning method provided in the above embodiment.
[0167] Furthermore, an embodiment of the present invention further provides a computer-readable medium, on which a computer program for implementing the vehicle driving path planning method is stored,
[0168] When the computer program is executed by an in-vehicle processor, it implements the vehicle driving path planning method provided in the first aspect embodiment above.
[0169] Further, as Figure 4 and Figure 5 shown, an embodiment of the present invention further provides a vehicle driving path planning system, which includes a vehicle system 103 and a background service system 101.
[0170] Further, an embodiment of the present invention may further provide a vehicle, which can implement the above vehicle driving path planning method or vehicle driving path planning device.
[0171] Next, refer to Figure 6 , which shows a schematic structural diagram of a computer system 600 suitable for implementing the vehicle driving path planning method according to an embodiment of the present invention. Figure 6 The shown computer system is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0172] As Figure 6 shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage section 608 into the random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the system 600 are also stored. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The input / output (I / O) interface 605 is also connected to the bus 604.
[0173] The following components are connected to the I / O interface 605: including an input section 606; including output sections 607 such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; including a storage section 608 such as a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as required. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as required, so that the computer program read from it can be installed into the storage section 608 as required.
[0174] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through the communication section 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the system of the present invention are executed.
[0175] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, and the above-mentioned module, segment of a program, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0177] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes the above-mentioned scenario determination module, trajectory generation module, and control module. Among them, the names of these modules or units do not, in some cases, constitute a limitation on the module or unit itself. For example, the scenario determination module can also be described as "the module or unit that determines the driving scenario in which the vehicle is located".
[0178] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist separately and not be assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: determining the driving scenario in which the vehicle is located; using a scenario trajectory generator to generate a path driving trajectory for the driving scenario, where the scenario trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on the real scenario and the virtual scenario extended by the scenario trajectory generator, and in combination with a reward model related to driving preferences; and controlling the vehicle to drive according to the path driving trajectory.
[0179] According to the technical solution of the embodiment of the present invention, the scenario trajectory generator generates a corresponding path driving trajectory for the driving scenario in which the vehicle is located. Since the scenario trajectory generator is based on the real scenario and the extended virtual scenario through an unsupervised training model (that is, more and wider scenario data is introduced in the training process), and the scenario trajectory evaluator trained in combination with the reward model related to driving preferences adjusts the virtual scenario and the unsupervised training model. That is, the scenario trajectory generator is obtained by training based on more scenarios, which improves the accuracy of path planning for different scenarios, and the driving preference is introduced in the training process to adjust it, so that the path driving trajectory generated by the scenario trajectory generator not only matches the driving scenario, but also can meet the driving preference, improving the driving experience of the driver / occupant.
[0180] The above specific implementation manners do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A vehicle driving path planning method, characterized in that, Including: Determine the driving scenario in which the vehicle is located; Generate a path driving trajectory for the driving scenario by using a scenario trajectory generator, wherein the scenario trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on a real scenario and a virtual scenario extended by the scenario trajectory generator, and in combination with a reward model related to driving preferences; Control the vehicle to drive according to the path driving trajectory.
2. The vehicle driving path planning method according to claim 1, wherein It also includes: Iteratively execute the following steps N1 to N3 until an iteration stop condition is met to obtain a scenario trajectory generator: Step N1: Use the current scenario data, through an unsupervised training model, and use the trained model to generate a virtual scenario, and generate multiple sets of predicted driving trajectories for the real scenario and the virtual scenario respectively. For the first training cycle, the current scenario data is the real scenario, and for non-first training cycles, the current scenario data is the real scenario and the corrected virtual scenario obtained in the previous training cycle; Step N2: Use a scenario trajectory evaluator to evaluate the virtual scenario, multiple sets of predicted driving trajectories corresponding to the real scenario, and multiple sets of predicted driving trajectories corresponding to the virtual scenario; Step N3: According to the evaluation results, correct the virtual scenario and adjust the unsupervised training model.
3. The vehicle driving path planning method according to claim 2, wherein, The step of generating multiple sets of predicted driving trajectories for the real scenario and the virtual scenario in Step N1 includes: Generate a predicted driving trajectory according to the current execution action of the vehicle, the average value of the vehicle state before the current execution action of the vehicle, the state noise variance corresponding to the vehicle state, the cumulative state of the noise, and the noise variance of the current execution action.
4. The vehicle driving path planning method according to claim 3, wherein, The generation of the predicted driving trajectory includes: Use the following first calculation formula to generate a predicted driving trajectory; First calculation formula: p[a(k)∣a(k - 1)] = N[a(k); μ(k), σ(k)^2 × I] Wherein, p[a(k)∣a(k - 1)] represents the change of the vehicle's action from the (k - 1)-th action to the k-th action; a(k) represents the vehicle's action at the k-th step; a(k - 1) represents the vehicle's action at the (k - 1)-th step; μ(k) represents the average value of the vehicle state before the k-th step; σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; N[] represents the Gaussian distribution formula; I represents a controllable parameter related to the driving trajectory.
5. The vehicle driving path planning method according to claim 4, wherein It also includes: Use the following second calculation formula to calculate the average value of the vehicle state before the k-th step; Second calculation formula: μ(k) = alpha(k - 1) × belta(k) × a(0) + alpha(k - 1) × a(k) Wherein, μ(k) represents the average value of the vehicle state at the k-th step and before the k-th step; alpha(k - 1) represents the cumulative state of the noise before the k-th step; belta(k) represents the noise variance at the k-th step; a(0) represents the starting action of the vehicle; a(k) represents the vehicle's action at the k-th step; And / or The vehicle driving path planning method further includes: calculating the state noise variance of the vehicle state at the k-th step by using the following third calculation formula; Third calculation formula: σ(k)^2 = [1 - alpha(k - 1)] × belta(k) / [1 - alpha(k)] σ(k)^2 represents the state noise variance of the vehicle state at the k-th step; alpha(k - 1) represents the cumulative noise state before the k-th step; belta(k) represents the noise variance at the k-th step; alpha(k) represents the cumulative noise state at the k-th step and before the k-th step.
6. The vehicle driving path planning method according to claim 1 or 2, characterized in that It further includes: By using the vehicle driving scenario, the vehicle driving trajectory corresponding to the vehicle driving scenario, and the driving preference corresponding to the vehicle driving trajectory, a scene trajectory evaluator is generated by supervising and iteratively training the reward model related to the driving preference.
7. The vehicle driving path planning method according to any one of claims 2 to 5, characterized in that In step N2, the evaluation of multiple groups of predicted driving trajectories corresponding to the real scenario and multiple groups of predicted driving trajectories corresponding to the virtual scenario includes: Evaluating multiple groups of predicted driving trajectories corresponding to the real scenario and multiple groups of predicted driving trajectories corresponding to the virtual scenario by using the following fourth calculation formula; Fourth calculation formula: P(R(s i )) > R(s j )) represents the probability that the i-th predicted driving trajectory in scenario S (real scenario or virtual scenario) is better than the j-th predicted driving trajectory; R(s i ) represents the score of the i-th predicted driving trajectory of scenario S by the scenario trajectory evaluator; R(s j ) represents the score of the j-th predicted driving trajectory of scenario S by the scenario trajectory evaluator.
8. The vehicle driving path planning method according to any one of claims 2 to 5, characterized in that The modification of the virtual scenario in N3 includes: Eliminating the virtual scenarios that do not meet the preset evaluation threshold.
9. The vehicle driving path planning method according to any one of claims 2 to 5, characterized in that It further includes: Based on the reinforcement learning and the evaluation result of the scene trajectory evaluator on the scene trajectory generator, the scene trajectory generator is adjusted.
10. The vehicle driving path planning method according to claim 1 or 2, characterized in that, It further includes: Deploying the scene trajectory generator to the in-vehicle system; Based on the in-vehicle system, performing the step of generating a path driving trajectory for the driving scenario by using the scene trajectory generator.
11. The vehicle driving path planning method according to claim 10, wherein It further includes: Deploying the scene trajectory evaluator to the in-vehicle system; In response to optimizing the path planning, using the driving-related data generated by the vehicle and including the driver's preference to adjust the scene trajectory evaluator deployed in the in-vehicle system; Using the adjusted scene trajectory evaluator to evaluate at least one group of predicted driving trajectories generated by the scene trajectory generator; According to the evaluation result, adjusting the scene trajectory generator deployed in the in-vehicle system.
12. A vehicle driving path planning device, characterized in that, It includes: A scene determination module, a trajectory generation module, and a control module, where The scene determination module is used to determine the driving scenario where the vehicle is located; The trajectory generation module is used to generate a path driving trajectory for the driving scenario by using a scene trajectory generator, where the scene trajectory generator is obtained by adjusting the virtual scenario and the unsupervised training model through an unsupervised training model based on the real scenario and the virtual scenario extended by the scene trajectory generator, and combining a reward model related to the driving preference; The control module is used to control the vehicle to drive according to the path driving trajectory.