Vehicle path planning method and device, vehicle and storage medium
By combining a large language model and a path planning model, interactive text is generated using vehicle location, obstacle and traffic information to optimize path planning, solving the problem of rigid path planning in existing technologies and achieving safer and more efficient autonomous driving.
Patent Information
- Application Number
- CN202411003072.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-24
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2044-07-24
AI Technical Summary
Existing vehicle path planning algorithms are relatively rigid and cannot generate optimal driving trajectories, affecting the performance of autonomous driving, especially in complex scenarios such as schools and construction zones where they cannot flexibly handle obstacles.
This method combines a large language model and a path planning model. By acquiring information on vehicle location, obstacles, maps, and traffic participants, it generates interactive text and plans paths. The Transformer language model and the Actor-Critic algorithm are then used for path optimization.
It improves the safety and efficiency of vehicle operation, enables flexible handling of complex scenarios, generates more suitable driving strategies and paths, and enhances the performance of autonomous driving.
Smart Images

Figure CN119197550B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle technology, and in particular to a vehicle routing method, apparatus, vehicle, and storage medium. Background Technology
[0002] In the field of autonomous driving, vehicles acquire information about their surrounding environment through environmental perception modules, and combine this information with their own position and motion status to make behavioral decisions, plan paths, and control motion, generating corresponding control commands. These control commands are then sent to the vehicle's chassis actuators to achieve autonomous driving.
[0003] In related technologies, driving trajectories are generally generated using search algorithms. However, this generation method is rather rigid, resulting in driving trajectories that are often not optimal, which greatly affects the performance of autonomous driving. Summary of the Invention
[0004] This application proposes a vehicle path planning method, apparatus, vehicle, and storage medium to better realize vehicle path planning.
[0005] In a first aspect, embodiments of this application provide a vehicle path planning method, the method comprising: acquiring driving-related parameters of the vehicle at the current moment, the driving-related parameters including at least the vehicle's positioning information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle; inputting the driving-related parameters into a pre-trained large language model to obtain interactive text output by the large language model, the interactive text being used to characterize the vehicle's driving strategy after the current moment; inputting the map information, the navigation information, and the interactive text into a pre-trained path planning model to obtain a planned path output by the path planning model; and controlling the vehicle to drive according to the planned path.
[0006] Secondly, embodiments of this application provide a vehicle path planning device, comprising: a parameter acquisition module, a strategy generation module, a path planning module, and a control module. The parameter acquisition module acquires driving-related parameters of the vehicle at the current moment, including at least the vehicle's location information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle. The strategy generation module inputs the driving-related parameters into a pre-trained large language model to obtain interactive text output by the large language model, the interactive text representing the vehicle's driving strategy after the current moment. The path planning module inputs the map information, the navigation information, and the interactive text into a pre-trained path planning model to obtain a planned path output by the path planning model. The control module controls the vehicle to travel according to the planned path.
[0007] Thirdly, embodiments of this application provide a vehicle, including: one or more processors; a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to perform the methods described above.
[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code that can be invoked by a processor to execute the methods described above.
[0009] The solution provided in this application obtains the vehicle's current driving-related parameters, which include at least the vehicle's location information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle. These driving-related parameters are then input into a pre-trained large language model to obtain interactive text output by the model. This interactive text represents the vehicle's driving strategy after the current moment. The map information, navigation information, and interactive text are then input into a pre-trained path planning model to obtain a planned path output by the model. The vehicle is then controlled to drive along the planned path. In this way, the pre-trained large language model can understand the vehicle's current driving scenario based on its driving-related parameters, flexibly deduce a more suitable driving strategy for that scenario, and output it. The pre-trained path planning model then generates a more suitable driving path based on the more suitable driving strategy, map information, and navigation information, and controls the vehicle to drive along the more suitable path, thereby improving the vehicle's driving safety and efficiency. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A schematic flowchart of a vehicle routing method provided in an embodiment of this application is shown.
[0012] Figure 2 A flowchart illustrating a vehicle routing method provided in another embodiment of this application is shown.
[0013] Figure 3 It shows Figure 2 A flowchart illustrating a sub-step of step S230 in one embodiment.
[0014] Figure 4 The diagram shows a model architecture diagram of a path planning model provided in an embodiment of this application.
[0015] Figure 5 It shows Figure 2 A flowchart illustrating a sub-step of step S240 in one embodiment.
[0016] Figure 6 A schematic flowchart of a model training method provided in an embodiment of this application is shown.
[0017] Figure 7 This paper illustrates a model architecture diagram of an initial language model provided in an embodiment of this application.
[0018] Figure 8 This is a block diagram of a vehicle routing device according to an embodiment of this application.
[0019] Figure 9 This is a block diagram of a vehicle used to perform the vehicle routing method according to an embodiment of this application.
[0020] Figure 10 This is a storage unit in this application embodiment for storing or carrying program code that implements the vehicle routing method according to this application embodiment. Detailed Implementation
[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0022] It should be noted that some processes described in the specification, claims, and accompanying drawings of this application include multiple operations that appear in a specific order. These operations may not be performed in the order they appear herein, or they may be performed in parallel. Operation numbers such as S110, S120, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be performed sequentially or in parallel. Also, the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover non-exclusive inclusion, such that a process, method, system, product, or server that includes a series of steps or sub-modules is not necessarily limited to those steps or sub-modules that are explicitly listed, but may include other steps or sub-modules that are not explicitly listed or that are inherent to such process, method, product, or device.
[0023] In related technologies, driving trajectories are generally generated using search algorithms. These algorithms primarily involve scattering data points, connecting them into multiple lines, and then optimizing each line. However, this approach is rigid and inflexible, failing to address scenarios requiring deceleration, such as those near schools or construction sites. Furthermore, this method lacks spatiotemporal separation for interactive processing; it cannot simultaneously generate trajectory / control sequences involving lateral and longitudinal movements, making it impossible to generate a driving trajectory like "first slowing down to yield to pedestrian A, then accelerating to overtake vehicle B on the right."
[0024] To address the aforementioned problems, the inventors have proposed a vehicle routing method, apparatus, vehicle, and storage medium. The vehicle routing method provided in this application is described in detail below.
[0025] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a vehicle routing method according to an embodiment of this application. The following will be combined with... Figure 1 The vehicle routing method provided in this application embodiment is described in detail. The vehicle routing method may include the following steps:
[0026] Step S110: Obtain the driving-related parameters of the vehicle at the current moment. The driving-related parameters include at least the vehicle's positioning information, obstacle information, map information, navigation information, and the predicted trajectory information of traffic participants around the vehicle.
[0027] In this embodiment, the vehicle's positioning information can be obtained through a Global Positioning System (GPS). Obstacle information may include the obstacle's coordinates, speed, orientation, and type, and can be determined based on sensor data collected by the vehicle's cameras, lidar, and other sensors. Map information may include static information such as lane lines, road topology, traffic rules, and road attributes. Traffic participants around the vehicle may include other vehicles and other living beings (e.g., pedestrians or animals); the predicted trajectory information of traffic participants around the vehicle can be obtained through predictive analysis based on their current coordinates, speed, and orientation.
[0028] Step S120: Input the driving-related parameters into a pre-trained large language model to obtain the interactive text output by the large language model. The interactive text is used to characterize the driving strategy of the vehicle after the current moment.
[0029] Optionally, the large language model can be trained based on the Transformer language model. That is, through extensive training, the large language model can understand the current driving scenario of the vehicle by inputting driving-related parameters, and deduce the driving strategy of the vehicle after the current moment based on the understood driving scenario, that is, output the aforementioned interactive text.
[0030] For example, based on the input driving-related parameters, the large language model understands the current driving scenario of the vehicle as "unprotected left turn, pedestrian A is crossing the opposite sidewalk on the left, and vehicle B on the right is traveling from right to left relative to the vehicle". Based on this, the large language model can deduce a driving strategy such as "the vehicle first slows down to let pedestrian A pass, and then accelerates to overtake vehicle B on the right", and output it as the above interactive text.
[0031] Optionally, the large language model can deduce an optimal driving strategy, that is, output an optimal interactive text; of course, the large language model can also deduce multiple driving strategies, that is, output multiple interactive texts, and this embodiment does not limit this.
[0032] Step S130: Input the map information, the navigation information, and the interactive text into the pre-trained path planning model to obtain the planned path output by the path planning model.
[0033] The path planning model can be obtained by training an initial planning model based on the Actor-Critic Algorithm using a large amount of sample data.
[0034] For example, the interactive text is "The vehicle first slows down to yield to pedestrian A, then accelerates to overtake vehicle B on the right".
[0035] At this point, the planned path output by the path planning model can be "generating a driving trajectory behind pedestrian A after pedestrian A crosses the pedestrian crossing, and generating a driving trajectory in front of the right vehicle B".
[0036] Step S140: Control the vehicle to travel along the planned path.
[0037] In this embodiment, the path planning model outputs not only the planned path for selecting the vehicle based on the semantics of the interactive text, but also a control sequence for controlling the vehicle to travel along the planned path. This control sequence includes at least the vehicle's acceleration parameters and steering wheel angle parameters. Based on these parameters, the acceleration and steering wheel angle parameters can be transmitted to the vehicle's actuator control module via the CAN bus. The actuator control module then controls the vehicle to travel along the planned path according to these acceleration and steering wheel angle parameters.
[0038] Specifically, the above-mentioned planning path and control sequence can be represented by the following two formulas:
[0039] traj={(x,y)|x=f(t),y=g(t),t∈T}
[0040] ControlSeq={(a,SteerAngle)|a=h(t),SteerAngle=m(t),t∈T}
[0041] In this context, `traj` represents the planned path, (x, y) are the coordinates of the trajectory points within the planned path, `t` represents each prediction time, `T` represents the prediction time domain (e.g., 3 seconds), the function `f(t)` represents the change in the x-coordinate value of the trajectory point with each prediction time, and the function `g(t)` represents the change in the y-coordinate value of the trajectory point with each prediction time. `ControlSeq` represents the control sequence, `a` is the acceleration parameter, `SteerAngle` is the steering wheel angle parameter, the function `h(t)` represents the value of the vehicle's acceleration parameter with each prediction time, and the function `m(t)` represents the value of the vehicle's steering wheel angle parameter with each prediction time.
[0042] In this embodiment, the pre-trained large language model can understand the current driving scenario of the vehicle based on the vehicle's driving-related parameters, flexibly deduce a more suitable driving strategy for the driving scenario and output it. Then, the pre-trained path planning model generates a suitable driving path according to the appropriate driving strategy, map information and navigation information, and controls the vehicle to drive along the appropriate driving path, thereby improving the vehicle's driving safety and driving efficiency.
[0043] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating a vehicle routing method according to another embodiment of this application. The following will be combined with... Figure 2 The vehicle routing method provided in this application embodiment is described in detail. The vehicle routing method may include the following steps:
[0044] Step S210: Obtain the driving-related parameters at the current moment. The driving-related parameters include at least the vehicle's positioning information, obstacle information, map information, navigation information, and the predicted trajectory information of traffic participants around the vehicle.
[0045] In this embodiment, the specific implementation of step S210 can be found in the content of the foregoing embodiments, and will not be repeated here.
[0046] Step S220: Input the driving-related parameters into a pre-trained large language model to obtain multiple interactive texts output by the large language model. The interactive texts are used to characterize the driving strategy of the vehicle after the current moment.
[0047] In this embodiment, after understanding the current driving scenario of the vehicle based on the vehicle's driving-related parameters, the large language model can deduce a variety of feasible driving strategies for the driving scenario, that is, it can output multiple interactive texts, with different interactive texts representing different driving strategies.
[0048] For example, based on the input driving-related parameters, the large language model understands the current driving scenario of the vehicle as "unprotected left turn, pedestrian A is crossing towards the vehicle on the left sidewalk, and vehicle B is traveling from right to left relative to the vehicle on the right." Based on this, the large language model can deduce several different interactive texts as shown below:
[0049] Interactive text 1: The car accelerates past pedestrian A and then accelerates past car B on the right;
[0050] Interactive text 2: Provided that right vehicle B has already passed pedestrian A, the vehicle accelerates to pass pedestrian A and yields to right vehicle B;
[0051] Interactive text 3: The car slows down to let pedestrian A pass, then accelerates to overtake the car on the right, B;
[0052] Interactive text 4: The vehicle slows down to yield to pedestrian A and also slows down to yield to vehicle B on the right.
[0053] Step S230: Input the map information, the navigation information, and each interactive text into the path planning model to obtain the planned path corresponding to each interactive text.
[0054] Understandably, when a large language model outputs multiple interactive texts, a path planning model can be used to perform path planning for each interactive text in turn; that is, map information, navigation information, and each interactive text are input into the path planning model to obtain the planned path output by the path planning model for each interactive text, thus obtaining multiple planned paths corresponding one-to-one with multiple interactive texts.
[0055] For example, if multiple interactive texts include the aforementioned interactive text 1, interactive text 2, interactive text 3, and interactive text 4, the path planning model can output the planned path 1 corresponding to interactive text 1, the planned path 2 corresponding to interactive text 2, the planned path 3 corresponding to interactive text 3, and the planned path 4 corresponding to interactive text 4 in sequence.
[0056] In some implementations, please refer to Figure 3 Step S230 may include the contents of steps S231 to S233:
[0057] Step S231: Input the interactive text, the vehicle's driving status parameters, the map information, and the navigation information into the path generation unit to obtain multiple candidate paths output by the path generation unit.
[0058] Specifically, please refer to Figure 4 The path planning model can include an input communication port, a prediction unit, a data preprocessing unit, a path generation unit, a path selection unit, a path tracking unit, and an output communication port. First, the vehicle's driving state parameters (i.e., the vehicle's state) and obstacle states can be selected from perception information, vehicle positioning information, and chassis information. These parameters are transmitted via the CAN (Controller Area Network) bus and then to the prediction unit through the input communication port. The vehicle state can include the vehicle's coordinates, position, speed, acceleration, and orientation, while the obstacle state can include the obstacle's coordinates, position, speed, acceleration, and orientation.
[0059] Furthermore, the prediction unit can predict the obstacle's trajectory within a preset time period after the current moment based on the obstacle's coordinate position, speed, acceleration, and orientation, thus obtaining the obstacle's trajectory and the predicted trajectory. The preset time period is a pre-set duration, such as 3 seconds or 5 seconds. Of course, this duration can be adjusted according to actual needs, and this embodiment does not impose any restrictions on it.
[0060] Next, the interactive text, the vehicle's status, and the predicted trajectories of obstacles are all input into the data preprocessing unit for preprocessing. This data preprocessing unit contains a transformer-based network model that uses an attention mechanism to filter out information with greater relevance and interactivity to the vehicle from the predicted trajectories of numerous obstacles, such as the predicted trajectories of obstacles closer to the vehicle. In other words, the data preprocessing unit can further encode the information input to this unit and finally output encoded information with greater relevance and interactivity to the vehicle.
[0061] Furthermore, the vehicle's driving status parameters, after being preprocessed by the data preprocessing unit, are transmitted to the path generation unit. The path generation unit can generate multiple feasible candidate paths based on the semantics of the interactive text, the vehicle's driving status parameters, map information, and navigation information.
[0062] For example, taking the interactive text "The vehicle first slows down to yield to pedestrian A, then accelerates to overtake vehicle B on the right" as an example, the following candidate paths can be generated: 1 "The vehicle slows down to yield to pedestrian A within the current lane, and then accelerates to overtake vehicle B from the adjacent lane to the left of vehicle B"; 2 "The vehicle slows down to yield to pedestrian A within the current lane, and then accelerates to overtake vehicle B from the adjacent lane to the right of vehicle B"; 3 "The vehicle slows down to yield to pedestrian A within the current lane, and then accelerates to overtake vehicle B from the second lane to the left of vehicle B".
[0063] Step S232: Input multiple candidate paths and the interactive text into the path selection unit to obtain the optimal path output by the path selection unit, which is then used as the planned path.
[0064] The path selection unit, also known as the Critic module, is an action value function responsible for evaluating the performance of the Actor module in a given state (s) and guiding the Actor module's actions in the next stage. In simpler terms, the path selection unit can select the candidate path with the highest action value function value from multiple candidate paths based on the vehicle's current driving state. This optimal path is then used as the planned path for the vehicle in the given interactive text.
[0065] Step S233: Input the planned path, the vehicle's driving status parameters, and the predicted trajectory information of the obstacles into the path tracking unit to obtain the planned path and the control sequence corresponding to the planned path output by the path tracking unit.
[0066] Furthermore, by inputting the vehicle's driving state parameters, the predicted trajectory information of obstacles, and the planned path output by the path selection unit to the path tracking unit, the planned path output by the path tracking unit and the corresponding control sequence can be obtained. The path tracking unit is the aforementioned Actor module, which outputs the action space of the intelligent agent (i.e., the vehicle). The formulas for expressing the planned path and control sequence can be found in the aforementioned embodiments and will not be repeated here.
[0067] Step S240: Using a preset scoring strategy, score each of the planned paths to obtain a path score for each of the planned paths.
[0068] Understandably, since the path planning model outputs multiple different planned paths, a preset scoring strategy can be used to score each planned path, and the optimal planned path can be selected for the vehicle to drive based on the path score of each planned path.
[0069] In some implementations, please refer to Figure 5 Step S240 may include the contents of steps S241 to S242:
[0070] Step S241: Calculate the index score of each planned path using multiple index calculation formulas to obtain multiple index scores for each planned path. The multiple index calculation formulas correspond one-to-one with multiple preset performance evaluation indicators. The multiple preset performance evaluation indicators include at least security indicators, comfort indicators, communication efficiency indicators, and compliance indicators.
[0071] Optionally, the calculation formula for the safety index can be expressed as follows:
[0072] I safety =Ave(PODAR)
[0073]
[0074] w(·)=w D (d)or w T (t)
[0075] w(0) = 1
[0076]
[0077] Among them, Isafety The safety score represents the safety index, where t is the prediction time, T is the prediction time domain [s], n is the vehicle number of the surrounding vehicles, and N is the total number of surrounding vehicles. To predict the risk value caused by the nth vehicle at time t, For the potential collision damage between the vehicle at time t and the vehicle in the nth cycle, ω T ω is the time reduction factor. D The spatial reduction factor is d, and the distance between the two workshop outlines is [m]. The circle method can be used for calculation.
[0078] Alternatively, the formula for calculating the comfort index can be expressed as follows:
[0079]
[0080] Among them, I comfort The comfort score, as represented by the comfort index, includes instantaneous lateral and longitudinal acceleration, in m / s². 2 ω is the instantaneous yaw rate in rad / s, and jerk is the longitudinal jerk in m / s². 3 T is the statistical time step interval, which is generally taken as 20. Of course, this value can be adjusted according to actual needs, and this embodiment does not limit it.
[0081] Alternatively, the formula for calculating the communication efficiency index can be expressed as follows:
[0082]
[0083] Among them, I traffic The traffic efficiency score represents the traffic efficiency index, where v is the vehicle's speed. This represents the average speed of vehicles surrounding the vehicle. In other words, the higher the traffic efficiency when the vehicle follows the planned route, the higher the traffic efficiency score of that planned route.
[0084] Optionally, the calculation formula for the compliance indicators can be expressed as follows:
[0085] I lawful =score-x
[0086] Among them, l lawfulThe legality score under the legality index is represented by: `score` represents the initial score of the vehicle assuming no traffic violations, which can be set to 0; `x` represents the points deducted for traffic violations. Understandably, the points deducted for different traffic violations can be different or the same, and this embodiment does not impose such restrictions. It should be noted that the value of `x` is a positive number. That is, the more traffic violations a vehicle commits while following the planned route, the higher its legality score. lawul The lower the value, the better.
[0087] Step S242: Based on the weight parameters corresponding to the scores of each indicator, perform a weighted summation of the scores of multiple indicators for each planned path to obtain the comprehensive score of each planned path, which is used as the path score for each planned path.
[0088] Specifically, the formula for calculating the overall score of the indicators can be expressed as follows:
[0089] I total =α1·I safety +α2·I comfort +α3·I traffic +α4·I lawful
[0090] Among them, I total The comprehensive score represents the planned path, with α1 representing the weight parameter corresponding to the safety indicator, α2 representing the weight parameter corresponding to the comfort indicator, α3 representing the weight parameter corresponding to the traffic efficiency indicator, and α4 representing the weight parameter corresponding to the compliance indicator. It should be noted that the parameter values of α1, α2, α3, and α4 are preset values. Different weight parameters can be set to the same or different values; this embodiment does not impose such restrictions. Optionally, considering that the most important factor during vehicle operation is driving safety, followed by compliance, and then comfort and traffic efficiency, α1 can be set to > α4. Of course, the weight parameters α3 corresponding to the traffic efficiency index and α2 corresponding to the comfort index can be adjusted according to the current intelligent driving mode of the vehicle. Optionally, if the current intelligent driving mode is the comfort intelligent driving mode, the aforementioned weight parameters can be set to α1>α4>α2>α3. Optionally, if the current intelligent driving mode is the high-efficiency intelligent driving mode, the aforementioned weight parameters can be set to α1>α4>α3>α2.
[0091] Step S250: Determine the planning path with the highest path score from multiple planning paths, and use it as the target planning path.
[0092] Understandably, a higher path score indicates higher safety, comfort, and traffic efficiency when the vehicle travels along that path, and a lower probability of violating traffic rules. Therefore, after obtaining the path score of each of the multiple planned paths, the planned path with the highest path score is determined as the target planned path; this target planned path can also be understood as the optimal planned path among the multiple planned paths.
[0093] Step S260: Control the vehicle to travel along the target planned path.
[0094] Similarly, the target control sequence corresponding to the target planning path includes at least the target acceleration parameter and the target steering wheel parameter. Based on this, the vehicle can be controlled to travel along the target planning path according to the target acceleration parameter and the target steering wheel parameter.
[0095] In this embodiment, a pre-trained large language model is used to understand the current driving scenario of the vehicle based on its driving-related parameters. Based on this understanding, multiple feasible driving strategies are deduced more quickly and comprehensively (i.e., multiple interactive texts are output). This deduction of all driving strategies is faster and more comprehensive than that of a human driver, significantly improving interpretability and interactive performance. Then, a pre-trained path planning model uses multiple feasible driving strategies, map information, and navigation information to identify multiple feasible driving paths. The path with the highest score is selected as the optimal driving path, and the vehicle is controlled to follow this optimal path, effectively improving safety, traffic efficiency, compliance, and comfort during driving. Furthermore, the interactive text containing driving strategies generated by the large language model considers the driving possibilities of a spatiotemporal segment compared to directly outputting trajectory / control sequences or control quantities. This allows for obtaining feasible solutions with a larger area and richer interaction methods. The corresponding trajectory / control sequence is then generated based on the interactive text containing the driving strategies. In other words, this application simulates the process of implicit semantic decision-making and interaction during human driving to flexibly deduce driving strategies and realize the deduction of driving strategies with spatiotemporal separation.
[0096] Please refer to Figure 6 , Figure 6 This is a flowchart illustrating a model training method according to another embodiment of this application. The following will be combined with... Figure 6 The model training method provided in the embodiments of this application will be described in detail. This model training method may include the following steps:
[0097] Step S301: Obtain a second sample set, which includes multiple second sample data. Each second sample data includes driving-related parameters of the vehicle at a historical time. Each second sample data carries a preset label, which includes preset interactive text. The preset interactive text is used to characterize the vehicle's preset driving strategy after the historical time.
[0098] The second sample set can be established based on the vehicle's driving dataset. Each second sample data includes driving-related parameters of the vehicle at a historical time. The specific parameter information contained in these driving-related parameters is the same type of information as that contained in the driving-related parameters in the aforementioned embodiments, which will not be repeated here. Furthermore, each second sample data carries a preset label, which contains preset interactive text used to characterize the vehicle's preset driving strategy after the historical time.
[0099] Step S302: Input each of the second sample data into the initial language model to obtain the second predicted interactive text output by the initial language model.
[0100] Step S303: Based on the degree of difference between the second predicted interactive text and the preset interactive text, iteratively update the model parameters of the initial language model until the second preset condition is met, and obtain the pre-trained initial language model.
[0101] Further, each second sample data is input into the initial language model to obtain the second predicted interactive text output by the initial language model for each second sample data; and a second loss value is determined based on the degree of difference between the second predicted interactive text corresponding to each second sample data and the preset interactive text carried by each second sample data; and the model parameters of the initial language model are iteratively updated based on the second loss value until the second preset condition is met, thus obtaining the pre-trained initial language model. Optionally, the above initial language model can be a Transformer model, and the network architecture of the initial language model is as follows: Figure 7 As shown, the Transformer model includes an encoder module and a decoder module. The encoder module includes... Figure 7 The multi-head attention layer, residual & normalization layer, feedforward network layer, and residual & normalization layer are included; the decoder module includes... Figure 7 The layers include multi-head attention layer, residual & normalization layer, multi-head attention layer, residual & normalization layer, feedforward network layer, and residual & normalization layer.
[0102] The second preset condition can be: the second loss value is less than a preset value, the second loss value no longer changes, or the number of training iterations reaches a preset number, etc. It is understood that after iteratively training the initial language model on the second sample set for multiple training cycles (each training cycle includes multiple iterations), continuously optimizing the parameters in the initial language model, the second loss value becomes smaller and smaller until it reaches a fixed value or is less than the preset value. At this point, it indicates that the initial language model has converged. Alternatively, it can be determined that the initial language model has converged after the number of training iterations reaches a preset number. In this case, the initial language model can be used as the pre-trained initial language model. The preset value and the preset number of iterations are pre-set and can be adjusted according to different application scenarios; this embodiment does not impose any restrictions on this.
[0103] In other words, by using a large amount of second sample data, the initial language model is trained to understand the vehicle's driving scenarios based on the vehicle's driving-related parameters, as well as to plan driving strategies for the understood driving scenarios. In this way, the pre-trained initial language model will initially have the ability to understand the vehicle's driving scenarios and plan driving strategies for the understood driving scenarios.
[0104] Step S304: Obtain a first sample set, which includes multiple first sample data, each of which includes driving-related parameters of the vehicle at a historical time and the historical driving path after the historical time.
[0105] Furthermore, after completing the pre-training of the initial language model, a first sample set can be obtained. Each first sample in the first sample set includes the vehicle's driving-related parameters at a historical moment and its historical driving path after that historical moment. It should be noted that the vehicle's driving behavior according to the historical driving path in each first sample is reasonable and safe.
[0106] Step S305: Input each of the first sample data into the pre-trained initial language model to obtain the first predicted interactive text output by the pre-trained initial language model.
[0107] Step S306: Input the first predictive interactive text into the initial planning model to obtain the first planning path output by the initial planning model.
[0108] Thus, by utilizing the pre-trained initial language model and based on each first sample data, a suitable driving strategy can be output for each first sample data, that is, a suitable first predicted interaction text can be output. The first predicted interaction text is the driving strategy of the vehicle after the historical time predicted by the initial language model. The first predicted interaction text is then input into the initial planning model to obtain the first planned path output by the initial planning model for the first predicted interaction text.
[0109] Step S307: Based on the degree of difference between the first planned path and the historical driving path, iteratively update the model parameters of the pre-trained initial language model and the initial planning model until the first preset condition is met, and obtain the updated initial language model as the large language model and the updated initial planning model as the path planning model.
[0110] Specifically, a first loss value is determined based on the degree of difference between the first planned path and the historical driving path; and the model parameters of both the pre-trained initial language model and the initial planning model are iteratively updated based on the first loss value until a first preset condition is met, resulting in an updated initial language model as the aforementioned large language model, and an updated initial planning model as the aforementioned path planning model. The model architecture of the initial language model can be found in [reference needed]. Figure 4 The model architecture shown will not be elaborated further here.
[0111] The first preset condition can be: the first loss value is less than a preset value, the first loss value no longer changes, or the number of training iterations reaches a preset number, etc. It is understood that after iteratively training the pre-trained initial language model and initial planning model on the first sample set for multiple training cycles (each training cycle includes multiple iterations), continuously optimizing the parameters in the pre-trained initial language model and initial planning model, the first loss value becomes smaller and smaller, eventually becoming a fixed value or less than the preset value. At this point, it indicates that the pre-trained initial language model and initial planning model have converged. Alternatively, it can be determined that the pre-trained initial language model and initial planning model have converged after the number of training iterations reaches a preset number. In this case, the converged pre-trained initial language model can be used as the large language model, and the converged initial planning model can be used as the path planning model. The preset value and preset number of iterations are pre-set and can be adjusted according to different application scenarios; this embodiment does not impose any restrictions on this.
[0112] In other words, by using labeled second sample data and pre-training the initial language model through supervised learning, and then using unlabeled first sample data to train both the pre-trained initial language model and the initial planning model simultaneously through self-supervised learning, the amount of labeling work can be greatly reduced, thus completing the training of both the large language model and the path planning model.
[0113] It should be noted that the model training process in this embodiment is performed on a computer device. After the large language model and path planning model are trained, they can be deployed in the vehicle. The vehicle can then generate the optimal planned path and drive along the optimal planned path during intelligent driving, as described in the aforementioned embodiment. Further details will not be elaborated here. Optionally, the computer device can be an electronic terminal with data processing capabilities, including but not limited to smartphones, laptops, and desktop computers. Alternatively, the computer device can be a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0114] In this embodiment, labeled second sample data is used to pre-train the initial language model through supervised learning. Then, unlabeled first sample data is used to train the pre-trained initial language model and the initial planning model simultaneously through self-supervised learning. This greatly reduces the workload of labeling the training sample data, thus enabling the training of both the large language model and the path planning model.
[0115] Please refer to Figure 8 The diagram illustrates a structural block diagram of a vehicle path planning device 400 according to an embodiment of this application. The device 400 may include: a parameter acquisition module 410, a strategy generation module 420, a path planning module 430, and a control module 440.
[0116] The parameter acquisition module 410 is used to acquire the driving-related parameters of the vehicle at the current moment. The driving-related parameters include at least the vehicle's positioning information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle.
[0117] The strategy generation module 420 is used to input the driving-related parameters into a pre-trained large language model to obtain the interactive text output by the large language model. The interactive text is used to represent the driving strategy of the vehicle after the current moment.
[0118] The path planning module 430 is used to input the map information, the navigation information and the interactive text into a pre-trained path planning model to obtain the planned path output by the path planning model.
[0119] The control module 440 is used to control the vehicle to travel along the planned path.
[0120] In some embodiments, the path planning model includes at least a path generation unit, a path selection unit, and a path tracking unit. The path planning module 430 can be specifically used to: input the interactive text, the vehicle's driving state parameters, the map information, and navigation information to the path generation unit to obtain multiple candidate paths output by the path generation unit; input the multiple candidate paths and the interactive text to the path selection unit to obtain the optimal path output by the path selection unit, which is used as the planned path; and input the planned path, the vehicle's driving state parameters, and the predicted trajectory information of obstacles to the path tracking unit to obtain the planned path output by the path tracking unit and the control sequence corresponding to the planned path.
[0121] In this mode, the control sequence includes at least the vehicle's acceleration parameters and steering wheel angle parameters. The control module 440 can specifically be used to: control the vehicle to travel along the planned path based on the acceleration parameters and the steering wheel angle parameters.
[0122] In some embodiments, the number of interactive texts is multiple, and the number of planned paths is multiple, with each planned path corresponding one-to-one with one of the interactive texts. The vehicle path planning device 400 may further include a scoring module and a target path determination module. The scoring module is used to score each planned path using a preset scoring strategy before controlling the vehicle to travel along the planned path, obtaining a path score for each planned path. The target path determination module is used to determine the planned path with the highest path score from the multiple planned paths as the target planned path. The control module 440 is used to control the vehicle to travel along the target planned path.
[0123] In this approach, the scoring module can be specifically used to: calculate the index score of each planned path using multiple index calculation formulas to obtain multiple index scores for each planned path, wherein the multiple index calculation formulas correspond one-to-one with multiple preset performance evaluation indicators, and the multiple preset performance evaluation indicators include at least security indicators, comfort indicators, communication efficiency indicators, and compliance indicators; and, based on the weight parameters corresponding to each index score, perform a weighted summation of the multiple index scores for each planned path to obtain a comprehensive index score for each planned path, which serves as the path score for each planned path.
[0124] In some embodiments, the vehicle path planning device 400 may further include a language model training module. The language model training module may be used to: acquire a first sample set, the first sample set including multiple first sample data, each first sample data including driving-related parameters of the vehicle at a historical time and a historical driving path after the historical time; input each first sample data into a pre-trained initial language model to obtain a first predicted interactive text output by the pre-trained initial language model; input the first predicted interactive text into an initial planning model to obtain a first planned path output by the initial planning model; and iteratively update the model parameters of the pre-trained initial language model and the initial planning model based on the degree of difference between the first planned path and the historical driving path, until a first preset condition is met, obtaining an updated initial language model as the large language model and an updated initial planning model as the path planning model.
[0125] In some embodiments, the vehicle path planning device 400 may further include a path planning model training module. The path planning model training module may be used to: acquire a second sample set, the second sample set including multiple second sample data, each second sample data including driving-related parameters of the vehicle at a historical time, each second sample data carrying a preset label, the preset label including preset interactive text, the preset interactive text being used to characterize the vehicle's preset driving strategy after the historical time; input each second sample data into the initial language model to obtain a second predicted interactive text output by the initial language model; and iteratively update the model parameters of the initial language model based on the degree of difference between the second predicted interactive text and the preset interactive text until a second preset condition is met, thereby obtaining the pre-trained initial language model.
[0126] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0127] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.
[0128] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0129] In summary, this application utilizes a pre-trained large language model to understand the current driving scenario of the vehicle based on its driving-related parameters. Based on the understood driving scenario, it can more quickly and comprehensively deduce multiple feasible driving strategies (i.e., output multiple interactive texts). Then, a pre-trained path planning model uses multiple feasible driving strategies, map information, and navigation information to identify multiple feasible driving paths. Finally, it selects the driving path with the highest path score as the optimal driving path and controls the vehicle to drive along the optimal driving path, thereby effectively improving the safety, traffic efficiency, compliance, and comfort of the vehicle during driving.
[0130] The following will combine Figure 9 This application describes one type of vehicle.
[0131] Reference Figure 9 , Figure 9 The diagram shows a structural block diagram of a vehicle 500 according to an embodiment of this application. The above-described method provided in this embodiment of the application can be executed by the vehicle 500.
[0132] The vehicle 500 in this application embodiment may include one or more of the following components: processor 501, memory 502, and one or more application programs, wherein the one or more application programs may be stored in memory 502 and configured to be executed by one or more processors 501, and the one or more programs are configured to perform the methods as described in the foregoing method embodiments.
[0133] Processor 501 may include one or more processing cores. Processor 501 connects to various parts within the vehicle 500 using various interfaces and lines, and performs various functions and processes data of the vehicle 500 by running or executing instructions, programs, code sets, or instruction sets stored in memory 502, and by calling data stored in memory 502. Optionally, processor 501 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). Processor 501 may integrate one or more of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the aforementioned modem can also be integrated into processor 501 and implemented using a separate communication chip.
[0134] The memory 502 may include random access memory (RAM) or read-only memory (ROM). The memory 502 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 502 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described below. The data storage area may also store data created by the vehicle 500 during use (such as the various correspondences described above).
[0135] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0136] In the several embodiments provided in this application, the coupling or direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication connection through some interface, device or module, and may be electrical, mechanical or other forms.
[0137] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0138] Please refer to Figure 10 This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable medium 600 stores program code that can be called by a processor to execute the methods described in the above method embodiments.
[0139] The computer-readable storage medium 600 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 600 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 600 has storage space for program code 610 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 610 may be compressed, for example, in a suitable form.
[0140] In some embodiments, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the steps in the above-described method embodiments.
[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A vehicle route planning method, characterized in that, The method includes: Obtain the vehicle's current driving-related parameters, which include at least the vehicle's location information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle; The driving-related parameters are input into a pre-trained large language model to obtain multiple interactive texts output by the large language model. The interactive texts are used to characterize the driving strategy of the vehicle after the current moment. The map information, the navigation information, and the multiple interactive texts are input into a pre-trained path planning model to obtain multiple planned paths output by the path planning model. Each of the multiple planned paths corresponds one-to-one with a single interactive text. Using a preset scoring strategy, each of the planned paths is scored to obtain a path score for each of the planned paths; The planning path with the highest score among the multiple planning paths is selected as the target planning path. Control the vehicle to travel along the target planned path.
2. The method according to claim 1, characterized in that, The path planning model includes at least a path generation unit, a path selection unit, and a path tracking unit. The map information, the navigation information, and each interactive text are input into a pre-trained path planning model to obtain the planned path corresponding to each interactive text, including: The interactive text, the vehicle's driving status parameters, the map information, and the navigation information are input into the path generation unit to obtain multiple candidate paths output by the path generation unit. The multiple candidate paths and the interactive text are input into the path selection unit to obtain the optimal path output by the path selection unit, which is then used as the planned path. The planned path, the vehicle's driving status parameters, and the predicted trajectory information of obstacles are input into the path tracking unit to obtain the planned path and the corresponding control sequence output by the path tracking unit.
3. The method according to claim 2, characterized in that, The control sequence includes at least the vehicle's acceleration parameters and steering wheel angle parameters; Controlling the vehicle to travel along the planned path includes: Based on the acceleration parameters and the steering wheel angle parameters, the vehicle is controlled to travel along the planned path.
4. The method according to claim 1, characterized in that, The step of scoring each planned path using a preset scoring strategy to obtain a path score for each planned path includes: The index score of each planned path is calculated using multiple index calculation formulas to obtain multiple index scores for each planned path. The multiple index calculation formulas correspond one-to-one with multiple preset performance evaluation indicators, which include at least security indicators, comfort indicators, communication efficiency indicators, and compliance indicators. Based on the weight parameters corresponding to the scores of each indicator, the scores of multiple indicators for each planned path are weighted and summed to obtain the comprehensive indicator score of each planned path, which is used as the path score of each planned path.
5. The method according to any one of claims 1-4, characterized in that, The training process of the large language model and the path planning model includes: Obtain a first sample set, which includes multiple first sample data, each of which includes driving-related parameters of the vehicle at a historical moment and the historical driving path after the historical moment. Each of the first sample data is input into the pre-trained initial language model to obtain the first predicted interactive text output by the pre-trained initial language model; The first predicted interactive text is input into the initial planning model to obtain the first planning path output by the initial planning model; Based on the degree of difference between the first planned path and the historical driving path, the model parameters of the pre-trained initial language model and the initial planning model are iteratively updated until the first preset condition is met, and the updated initial language model is obtained as the large language model, and the updated initial planning model is obtained as the path planning model.
6. The method according to claim 5, characterized in that, The pre-training process of the initial language model includes: Obtain a second sample set, which includes multiple second sample data. Each second sample data includes driving-related parameters of the vehicle at a historical time. Each second sample data carries a preset label, which includes preset interactive text. The preset interactive text is used to characterize the vehicle's preset driving strategy after the historical time. Each of the second sample data is input into the initial language model to obtain the second predicted interactive text output by the initial language model; Based on the degree of difference between the second predicted interactive text and the preset interactive text, the model parameters of the initial language model are iteratively updated until the second preset condition is met, thus obtaining the pre-trained initial language model.
7. A vehicle route planning device, characterized in that, The device includes: The parameter acquisition module is used to acquire the driving-related parameters of the vehicle at the current moment. The driving-related parameters include at least the vehicle's positioning information, obstacle information, map information, navigation information, and predicted trajectory information of traffic participants around the vehicle. The strategy generation module is used to input the driving-related parameters into a pre-trained large language model to obtain multiple interactive texts output by the large language model. The interactive texts are used to represent the driving strategy of the vehicle after the current moment. The path planning module is used to input the map information, the navigation information, and multiple interactive texts into a pre-trained path planning model to obtain multiple planned paths output by the path planning model, and the multiple planned paths correspond one-to-one with the multiple interactive texts; The scoring module is used to score each of the planned paths using a preset scoring strategy to obtain a path score for each of the planned paths. The target path determination module is used to determine the planning path with the highest path score from multiple planned paths, and use it as the target planning path. The control module is used to control the vehicle to travel along the target planned path.
8. A vehicle, characterized in that, The vehicles include: One or more processors; Memory; One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code that can be invoked by a processor to execute the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent agent decision-making method, control method, electronic equipment and storage medium
CN117151246A
Multi-dimensional intelligent driving path planning system
CN117824695A