Automatic driving behavior decision method and device, storage medium and electronic equipment
Patent Information
- Application Number
- CN202511996579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-26
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2045-12-26
AI Technical Summary
[0002]在相关技术钟,自动驾驶的行为决策方法大多依赖当前车辆与参考车辆之间的实时相对物理状态(如相对距离、相对速度等)进行分析,导致自动驾驶的行为规划过程中缺乏对参考车辆后续行驶趋势的预判,进而使得输出的行驶行为决策与实际交通场景的适配性不足,存在变道安全性不佳、决策合理性欠缺的问题
[0015]根据本申请实施例的又一方面,还提供了一种电子设备,包括存储器和处理器,上述存储器中存储有计算机程序,上述处理器被设置为通过上述计算机程序执行上述的自动驾驶的行为决策方法。
Smart Images

Figure CN121671641B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle detection technology, and more specifically, to a method and apparatus for making behavioral decisions for autonomous driving, a storage medium, and an electronic device. Background Technology
[0002] In related technologies, most autonomous driving behavior decision-making methods rely on analyzing the real-time relative physical state (such as relative distance and relative speed) between the current vehicle and a reference vehicle. This results in a lack of prediction of the reference vehicle's subsequent driving trend during the autonomous driving behavior planning process, leading to insufficient adaptability of the output driving behavior decisions to actual traffic scenarios. This results in problems such as poor lane-changing safety and insufficient decision-making rationality. In other words, the accuracy of the autonomous driving behavior decisions provided by related technologies is relatively low. Summary of the Invention
[0003] This application provides an autonomous driving behavior decision-making method and apparatus, storage medium and electronic device, to at least solve the technical problem of low accuracy in the behavior decision-making of autonomous vehicles in related technologies.
[0004] According to one aspect of the embodiments of this application, an autonomous driving behavior decision-making method is provided, comprising: determining a set of driving state parameters matching a current vehicle and a reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle; obtaining a driving style label of the reference vehicle, wherein the driving style label is determined based on historical driving data of the reference vehicle in a historical driving cycle; inputting the set of driving state parameters and the driving style label data into a behavior planning model to obtain the operation probability matching at least two driving behaviors respectively; and determining a target driving behavior from at least two driving behaviors based on the at least two operation probabilities.
[0005] According to another aspect of the embodiments of this application, an autonomous driving behavior decision-making device is also provided, comprising: a first determining unit, configured to determine a set of driving state parameters matching a current vehicle and a reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle; an acquiring unit, configured to acquire a driving style label of the reference vehicle, wherein the driving style label is determined based on historical driving data of the reference vehicle within a historical driving cycle; an input unit, configured to input the set of driving state parameters and the driving style label data into a behavior planning model to acquire the operation probability matching at least two driving behaviors; and a second determining unit, configured to determine a target driving behavior from at least two driving behaviors based on at least two operation probabilities.
[0006] As an optional solution, the above-mentioned input unit includes: a first determining module, used to determine a driving feature vector that matches the driving state parameter set and driving style label; an input module, used to input the driving feature vector into a behavior planning model, wherein the behavior planning model is a neural network model containing an input layer, at least one hidden layer and an output layer; and a first acquiring module, used to acquire the operation probabilities corresponding to at least two output nodes in the output layer of the behavior planning model.
[0007] As an optional solution, the second determining unit includes: a second acquisition module for acquiring the actual driving behavior of the current vehicle; a second determining module for determining the current predicted loss based on the actual driving behavior and at least two operation probabilities; and an adjustment module for adjusting at least one network parameter in the behavior planning model based on the current predicted loss.
[0008] As an optional approach, the input unit further includes: a third acquisition module, used to acquire an incremental sample from the incremental sample set as the current incremental sample when the number of incremental samples included in the incremental sample set is greater than or equal to the target number threshold, wherein the incremental sample is determined based on historical driving behavior that meets the loss condition, and the historical driving behavior does not match the reference driving behavior predicted by the behavior planning model at a historical moment; an update module, used to update the behavior planning model by gradient descent when the cross-entropy loss matched by the current incremental sample does not meet the convergence condition; and a repetition module, used to repeat the steps until the cross-entropy loss meets the convergence condition.
[0009] As an optional embodiment, the aforementioned second determining unit includes: a first triggering module, configured to trigger a first control signal matching a first driving behavior among at least two driving behaviors, wherein the first control signal is used to control the current vehicle to accelerate in the current lane; a second triggering module, configured to trigger a second control signal matching a second driving behavior among at least two driving behaviors, wherein the second control signal is used to control the current vehicle to decelerate in the current lane; a third triggering module, configured to trigger a third control signal matching a third driving behavior among at least two driving behaviors, wherein the third control signal is used to control the current vehicle to accelerate into a reference lane; and a fourth triggering module, configured to trigger a fourth control signal matching a fourth driving behavior among at least two driving behaviors, wherein the fourth control signal is used to control the current vehicle to decelerate before entering the reference lane.
[0010] As an optional solution, the acquisition unit includes: a fourth acquisition module for acquiring driving style tags sent by the reference vehicle; a fifth acquisition module for acquiring a set of driving description information of the reference vehicle in the first driving cycle, wherein the set of driving description information is used to indicate multiple driving decision behaviors of the reference vehicle in the first driving cycle; and inputting the set of driving description information into the driving style classification model to obtain driving style tags.
[0011] As an optional solution, the above-mentioned device further includes: a third determining unit, configured to determine a first position node and a second position node based on the trajectory position relationship between the first reference trajectory corresponding to the current vehicle in the spatiotemporal map and the second reference trajectory corresponding to the reference vehicle in the spatiotemporal map, wherein the spatiotemporal map is used to indicate the change of the object position of at least one vehicle object over time, and the at least one vehicle object includes the current vehicle and the reference vehicle; a fourth determining unit, configured to use the first position node as the trajectory starting point of the estimated driving trajectory of the current vehicle and the second position node as the trajectory ending point of the estimated driving trajectory; and a fifth determining unit, configured to determine at least one intermediate trajectory point in the estimated driving trajectory based on the driving style label and the target driving behavior.
[0012] As an optional solution, the fifth determining unit includes: a third determining module for determining a first collision risk weight based on a driving style label; a fourth determining module for determining a second collision risk weight based on the target driving behavior; and a fifth determining module for sequentially determining at least one intermediate trajectory point in the estimated driving trajectory based on the first collision risk weight and the second collision risk weight.
[0013] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described autonomous driving behavior decision method at runtime.
[0014] According to another aspect of the embodiments of this application, a computer program product is provided, the computer program product including a computer program / instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer program / instructions from the computer-readable storage medium, and executes the computer program / instructions, causing the computer device to perform the above-described autonomous driving behavior decision-making method.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-described autonomous driving behavior decision-making method through the computer program.
[0016] In this embodiment, the real-time relative physical state of the current vehicle and the reference vehicle is accurately obtained by determining the set of driving state parameters. Simultaneously, a driving style label determined based on the historical driving data of the reference vehicle is introduced to fully consider its driving characteristics. Both types of data are input into the behavior planning model, enabling the model to generate operational probabilities for each driving behavior by combining the real-time physical state and the long-term driving trend of the reference vehicle. Finally, the target driving behavior is determined based on the operational probabilities. This effectively compensates for the shortcomings of existing technologies that do not consider the driving style of the reference vehicle, improves the adaptability of driving behavior decisions to actual traffic scenarios, and thus enhances the safety and rationality of lane-changing decisions in autonomous driving. In this way, the adaptability of driving behavior decisions to actual traffic scenarios is improved, thereby increasing the accuracy of behavioral decisions during autonomous driving. Therefore, the above method solves the technical problem of low accuracy in behavioral decisions for autonomous driving in related technologies. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0018] Figure 1 This is a schematic diagram of the hardware environment for an optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0019] Figure 2 This is a flowchart of an optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of an optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of an optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0023] Figure 6 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0024] Figure 7 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0025] Figure 8 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0026] Figure 9 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0027] Figure 10 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0028] Figure 11 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0029] Figure 12 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0030] Figure 13 This is a schematic diagram of another optional autonomous driving behavior decision-making method according to an embodiment of this application;
[0031] Figure 14 This is a schematic diagram of an optional autonomous driving behavior decision-making device according to an embodiment of this application;
[0032] Figure 15 This is a schematic diagram of an optional electronic device according to an embodiment of this application. Detailed Implementation
[0033] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0034] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0035] According to one aspect of the embodiments of this application, an autonomous driving behavior decision-making method is provided. As an optional implementation, the above-described autonomous driving behavior decision-making method can be applied to, but is not limited to, [examples of other methods]. Figure 1 The illustrated hardware environment provides a behavior decision-making system for autonomous driving. Optionally, the aforementioned behavior decision-making method for autonomous driving can be applied to a vehicle terminal. Figure 1 A side view of a vehicle terminal 101 is shown, which is mounted on and capable of traversing a travel surface 113. The vehicle terminal 101 includes an onboard navigation system 103, a computer-readable storage device or medium (memory) 102 including a digital road map 104, a spatial monitoring system 117, a vehicle controller 109, a GPS (Global Positioning System) sensor 110, an HMI (Human / Machine Interface) device 111, and also includes an autonomous controller 112 and a telematics controller 114. The vehicle terminal 101 may include, but is not limited to, commercial vehicles, industrial vehicles, agricultural vehicles, passenger vehicles, all-terrain vehicles, personal mobile devices, robots, and similar mobile platforms to achieve the purposes of this application.
[0036] In one embodiment, the spatial monitoring system 117 includes: one or more spatial sensors and systems arranged to monitor a visible area 105 in front of the vehicle terminal 101; and a spatial monitoring controller 118. Spatial sensors for monitoring the visible area 105 include, for example, a lidar sensor 106, a radar sensor 107, a camera 108, and so on. The placement of the spatial sensors allows the spatial monitoring controller 118 to monitor traffic flow, including approaching vehicles, intersections, lane markings, and other objects surrounding the vehicle terminal 101. The spatial sensors of the spatial monitoring system 117 may include object location sensing devices. The lidar sensor 106 uses pulsed and reflected laser beams to measure the range or distance to an object. The radar sensor 107 uses radio waves to determine the range, angle, and / or speed of an object. The camera 108 includes an image sensor, a lens, and a camera controller.
[0037] Camera 108 is advantageously mounted and positioned on vehicle terminal 101 in a location that allows for capturing images of a visible area 105, wherein at least a portion of the visible area 105 includes the area in front of vehicle terminal 101 and a portion of the travel surface 113 of the trajectory of vehicle terminal 101. The visible area 105 may also include the surrounding environment. Other cameras (not shown) may also be employed, for example, including a second camera positioned on the rear or side portion of vehicle terminal 101 to monitor the rear of vehicle terminal 101 and one of the right or left sides of vehicle terminal 101.
[0038] The autonomous controller 112 is configured to implement autonomous driving or advanced driver assistance system (ADAS) vehicle functionality. Such functionality may include an onboard vehicle control system capable of providing a certain level of driving automation. Driving automation may include a series of dynamic driving and vehicle operations. Driving automation may include simultaneous automatic control of vehicle driving functions (including steering, acceleration, and braking), wherein the driver relinquishes control of the vehicle for a period of time during the journey. Driving automation may include simultaneous automatic control of vehicle driving functions (including steering, acceleration, and braking), wherein the driver relinquishes control of the vehicle terminal 101 for the entire journey. Driving automation includes hardware and controllers configured to monitor the spatial environment in various driving modes to perform various driving tasks during dynamic vehicle operations. By way of non-limiting example, autonomous vehicle functionality includes adaptive cruise control (ACC) operation, lane guidance and lane keeping operation, lane changing operation, steering assist operation, object avoidance operation, parking assist operation, vehicle braking operation, vehicle speed and acceleration operation, vehicle lateral movement operation, for example, as part of lane guidance, lane keeping, and lane changing operations, etc.
[0039] The aforementioned autonomous controller can be equipped with an operating system and an autonomous driving system. The operating system is responsible for managing the hardware resources (including sensors, system bus, network, etc.) of the vehicle terminal 101 and scheduling computing resources. The autonomous driving system can implement various algorithms required for autonomous driving, including localization, environmental perception, path planning, and control, and can make decisions in situations such as driving on curves, driving in straight lines, driving in complex road conditions, and changing lanes.
[0040] When using vehicle terminal 101 for autonomous driving behavior decision-making, steps S1-S4 can be executed in vehicle terminal 101 to determine the set of driving state parameters matching the current vehicle and the reference vehicle; obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle in the historical driving cycle; input the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probability matching at least two driving behaviors; and determine the target driving behavior from at least two driving behaviors based on the at least two operation probabilities.
[0041] Vehicle terminal 101 may include a telematics controller 114, which includes a wireless telematics communication system capable of performing off-vehicle communications (including communications with a communication network 115 having both wireless and wired communication capabilities). Optionally or additionally, the telematics controller 114 may directly perform off-vehicle communications by communicating with a non-airborne server 116 via the communication network 115.
[0042] In this embodiment, a set of driving state parameters matching the current vehicle and a reference vehicle is determined, wherein the set of driving state parameters indicates the relative physical state between the current vehicle and the reference vehicle; a driving style label of the reference vehicle is obtained, wherein the driving style label is determined based on historical driving data of the reference vehicle within a historical driving cycle; the set of driving state parameters and the driving style label data are input into a behavior planning model to obtain the operation probability matching at least two driving behaviors; and a target driving behavior is determined from at least two driving behaviors based on the at least two operation probabilities. This approach improves the adaptability of driving behavior decisions to actual traffic scenarios, thereby improving the accuracy of behavior decisions during autonomous driving. Furthermore, the above method solves the technical problem of low accuracy in behavior decisions for autonomous driving vehicles provided in related technologies.
[0043] In alternative implementations, such as Figure 2 As shown, the above-mentioned autonomous driving behavior decision-making method includes the following steps:
[0044] S202, determine a set of driving state parameters that match the current vehicle and the reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle;
[0045] Optionally, in this embodiment, the current vehicle may refer to, but is not limited to, the autonomous vehicle performing the lane-changing operation itself, which is the subject of the lane-changing decision. For example, a vehicle traveling on an urban road preparing to change lanes from a straight lane to a turning lane.
[0046] Optionally, in this embodiment, the reference vehicle may refer to, but is not limited to, surrounding vehicles related to the current vehicle's lane-changing behavior, including vehicles in the target lane and adjacent lanes, whose driving status will affect the current vehicle's lane-changing decision. For example, when the current vehicle is preparing to change lanes to the left lane, vehicle A traveling in front and vehicle B traveling behind in the left lane can both be considered reference vehicles; or when vehicle C is preparing to change lanes from the right lane to the lane the current vehicle is traveling in, vehicle C can also be considered a reference vehicle.
[0047] Optionally, in this embodiment, the driving state parameter set may include, but is not limited to, the relative physical state data between the current vehicle and the reference vehicle, providing basic information on environmental and vehicle interaction for lane-changing decisions. This may include, but is not limited to, relative speed difference, relative distance, estimated time-to-collision (TTC), and reference vehicle type code. For example, the relative distance between the current vehicle and the vehicle in front in the left lane is 50 meters, the relative speed difference is -10 km / h, and the estimated collision time is 8 seconds. These data may, but are not limited to, collectively constitute the driving state parameter set. S204, Obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within a historical driving cycle;
[0048] S204, Obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within the historical driving cycle;
[0049] Optionally, in this embodiment, the driving style label may be, but is not limited to, a driving style classification identifier determined based on the historical driving data (such as lane change frequency, braking frequency, acceleration frequency, average acceleration, etc.) of the reference vehicle within a historical driving cycle. It may include, but is not limited to, multiple types such as aggressive, normal, and mild, and may be used to reflect the driving behavior preferences of the reference vehicle, assisting the current vehicle in predicting its level of cooperation during lane changes. For example, a reference vehicle with frequent historical lane changes and high average acceleration is labeled as aggressive; a reference vehicle with stable historical following distance and smooth braking is labeled as mild.
[0050] Optionally, in this embodiment, the historical driving cycle may refer to, but is not limited to, past driving time periods used to analyze the driving style of the reference vehicle.
[0051] Optionally, in this embodiment, historical driving data may refer to, but is not limited to, the original driving records generated by the reference vehicle within the historical driving cycle. This data forms the basis for generating driving style labels and may include, but is not limited to, vehicle speed, acceleration, steering angle, lane centering error, following distance, etc. For example, data such as the interval between the reference vehicle's past 10 lane changes and the peak deceleration of each braking action.
[0052] S206, Input the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probability that matches at least two driving behaviors;
[0053] Optionally, in this embodiment, the behavior planning model can be, but is not limited to, a classification model built on a multilayer perceptron (MLP), thereby fusing driving state parameters and driving style labels to output the probability of different driving behaviors, providing a quantitative basis for lane-changing decisions. This model is trained with massive amounts of road-collected data and can adapt to different driving scenarios.
[0054] Optionally, in this embodiment, driving behavior may refer to, but is not limited to, the operating methods that the current vehicle can choose in a lane-changing scenario. It may be divided into aggressive behaviors such as directly changing lanes and forcing the other party to give way, and compromising behaviors such as slowing down and avoiding, and choosing to change lanes through a rear window.
[0055] S208, determine the target driving behavior from at least two driving behaviors based on at least two operational probabilities.
[0056] Optionally, in this embodiment, the operational probability can be, but is not limited to, the probability of occurrence of various driving behaviors output by the behavior planning model, used to quantify the adaptability of different behaviors. The higher the probability value, the more the behavior is in line with the current scenario and the driver's preferences. For example, if the model outputs an aggressive lane change probability of 0.7 and a compromise avoidance probability of 0.3, it indicates that the current scenario is more suitable for direct lane changing.
[0057] Optionally, in this embodiment, the target driving behavior may be, but is not limited to, the optimal operation selected from at least two driving behaviors, i.e., the behavior with the highest operation probability, and may be, but is not limited to, the lane-changing action ultimately performed by the current vehicle.
[0058] Optionally, in this embodiment, vehicle sensors such as radar and vision sensors are used to collect and calculate the relative physical state data between the current vehicle and the reference vehicle, and form a structured set of parameters to clarify the interaction relationship between the two.
[0059] Next, using an offline-trained classification model, the historical driving data of the reference vehicle is analyzed to derive its driving style classification label. This helps the current vehicle predict the behavioral tendencies of the reference vehicle.
[0060] Furthermore, the quantified relative state parameters and reference vehicle style labels are used as inputs to a pre-trained behavior planning model. The model uses multi-layer neural network reasoning to output probability values for various driving behaviors, thereby fusing scene data with style information, quantifying the adaptability of different behaviors, and providing an objective basis for decision-making.
[0061] Finally, the operation probabilities of all driving behaviors are compared, and the behavior with the highest probability is selected as the final lane change operation for the current vehicle.
[0062] It should be noted that by first determining the relative driving state parameters of the current vehicle and the reference vehicle, obtaining the historical driving style labels of the reference vehicle, and then inputting the two types of data into the behavior planning model to obtain the operation probability of different driving behaviors, the target driving behavior is finally selected based on the probability. This achieves a technical effect that balances lane change safety and human-likeness. It avoids lane change conflicts caused by not considering the driving style of the reference vehicle, and makes lane change behavior more in line with human driving habits through quantitative decision-making. At the same time, it improves the reliability and adaptability of autonomous driving lane change decision-making.
[0063] The embodiments provided in this application determine a set of driving state parameters matching the current vehicle and a reference vehicle, wherein the set of driving state parameters indicates the relative physical state between the current vehicle and the reference vehicle; obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within a historical driving cycle; input the set of driving state parameters and the driving style label data into a behavior planning model to obtain the operation probability matching at least two driving behaviors; and determine the target driving behavior from at least two driving behaviors based on the at least two operation probabilities. In this way, the adaptability of driving behavior decisions to actual traffic scenarios is improved, thereby increasing the accuracy of behavior decisions during autonomous driving. Therefore, the above method solves the technical problem of low accuracy in behavior decisions for autonomous driving vehicles provided in related technologies.
[0064] As an optional approach, the driving state parameter set and driving style label data are input into the behavior planning model to obtain the operation probabilities that match at least two driving behaviors, including:
[0065] S1-1, Determine the driving feature vector that matches the set of driving state parameters and driving style labels;
[0066] S1-2, input the driving feature vector into the behavior planning model, wherein the behavior planning model is a neural network model containing an input layer, at least one hidden layer and an output layer;
[0067] S1-3, obtain the operation probabilities corresponding to at least two output nodes in the output layer of the behavior planning model.
[0068] Optionally, in this embodiment, the input layer may be, but is not limited to, the first layer of the neural network model. It receives and transmits driving feature vectors to the hidden layer and serves as the interface between the model and external data. It needs to match the dimension of the driving feature vectors.
[0069] Optionally, in this embodiment, the hidden layer can be, but is not limited to, at least one network structure located between the input layer and the output layer. Its function is to perform nonlinear transformation and feature extraction on the input features through activation functions (such as ReLU), to mine deep scene-behavior correlation information in the data and improve the accuracy of model decision-making. For example, a hidden layer containing 7 nodes can perform operations on 5-dimensional input features to generate 7 intermediate data containing deep features, which are then passed to the output layer.
[0070] Optionally, in this embodiment, the output layer may be, but is not limited to, the last layer of the neural network model. Its function is to output the operation probability corresponding to each driving behavior. The number of nodes is the same as the number of driving behavior types, and the sum of the probability values may be, but is not limited to, 1, directly providing a basis for the selection of the target driving behavior.
[0071] Optionally, in this embodiment, the output node may be, but is not limited to, a single neuron in the output layer. Each node corresponds to a driving behavior and its function is to output the operation probability of that behavior, which is a direct reflection of the model's decision result. For example, the first node in the output layer corresponds to aggressive lane changing with an output probability of 0.65; the second node corresponds to compromise avoidance with an output probability of 0.35.
[0072] Optionally, in this embodiment, the set of driving state parameters is standardized and the driving style labels are encoded. The processed data is then concatenated to form a driving feature vector of a unified dimension, thereby transforming the original data from multiple sources into a structured vector that can be processed by the neural network, avoiding the impact of inconsistent data formats or differences in dimensions on the accuracy of model training and inference.
[0073] Next, the obtained driving feature vector is fed into the pre-trained behavior planning model, thereby utilizing the nonlinear fitting capability of the neural network to uncover the complex relationship between driving features and driving behavior, ensuring the accuracy of decision probabilities and scene adaptability.
[0074] Finally, the values of each node in the output layer of the behavior planning model are read. After normalization, these values correspond to the operation probability of a certain driving behavior. The higher the probability value, the more suitable the behavior is for the current scenario.
[0075] The embodiments provided in this application first process the driving state parameter set and driving style label into a standardized driving feature vector, and then input it into a behavior planning model containing an input layer, a hidden layer and an output layer. Finally, the operation probability corresponding to different driving behaviors is obtained in the output layer. This achieves the technical effect of converting multi-source data into a format that the model can process, accurately mining the relationship between scenarios and behaviors through neural networks, and outputting quantitative decision-making basis. This provides reliable algorithmic support for the subsequent selection of the optimal driving behavior and improves the accuracy and adaptability of autonomous driving lane change decisions.
[0076] As an optional approach, after determining the target driving behavior from at least two driving behaviors based on at least two operational probabilities, the method further includes:
[0077] S2-1, Obtain the current actual driving behavior of the vehicle;
[0078] S2-2, determine the current predicted loss based on the actual driving distance and at least two operational probabilities;
[0079] S2-3, Adjust at least one network parameter in the behavior planning model based on the current predicted loss.
[0080] Optionally, in this embodiment, actual driving behavior may refer to, but is not limited to, the actual driving operation performed by the vehicle in the lane-changing scenario. This serves as the basis for determining the accuracy of the behavior planning model's predictions and acts as a benchmark to measure the deviation between the model's predictions and the actual operation. For example, if the behavior planning model predicts an aggressive lane change probability of 0.65, but the driver takes over or the system ultimately performs a compromise avoidance maneuver, then the compromise avoidance maneuver is considered the actual driving behavior.
[0081] Optionally, in this embodiment, the prediction loss can be, but is not limited to, a quantification of the degree of inconsistency between the model prediction result calculated based on the cross-entropy loss function and the actual driving behavior. The larger the prediction loss, the greater the discrepancy between the model prediction and the actual behavior; the smaller the value, the higher the consistency, thereby providing an error signal for adjusting the model parameters.
[0082] Optionally, in this embodiment, the network parameters may be, but are not limited to, learnable parameters in the behavior planning model, including the weights and biases between neurons in each layer, which determine the mapping relationship from input features to output probabilities. Parameter adjustment is a means of iterative optimization of the model.
[0083] Optionally, in this embodiment, the specific driving operation ultimately performed by the vehicle in the lane-changing scenario is collected through the sensors or logs of the vehicle control system to determine whether it belongs to a specific type among preset behavior types such as aggressive lane changing or compromise avoidance.
[0084] Next, the actual driving behavior is encoded into a vector, and this vector, along with at least two operation probabilities output by the model, is substituted into the prediction loss function, such as cross-entropy loss, to calculate a quantified prediction loss value, which reflects the consistency between the model prediction and the actual behavior.
[0085] Finally, at least one network parameter in the behavioral planning model is adjusted based on the current predicted loss.
[0086] The embodiments provided in this application achieve the technical effect of allowing the model to be iteratively optimized through actual driving data, continuously reducing the discrepancy between prediction and actual behavior, and improving adaptability to specific scenarios and driver habits after the target driving behavior is determined.
[0087] As an optional approach, before inputting the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probabilities matching at least two driving behaviors, the following steps are also included:
[0088] S3-1, If the number of incremental samples included in the incremental sample set is greater than or equal to the target number threshold, an incremental sample is obtained from the incremental sample set as the current incremental sample. The incremental sample is determined based on the historical driving behavior that meets the loss condition. The historical driving behavior does not match the reference driving behavior predicted by the behavior planning model at the historical moment.
[0089] S3-2, if the cross-entropy loss of the current incremental sample matching does not meet the convergence condition, the behavior planning model is updated by gradient descent.
[0090] S3-3, Repeat the above steps until the cross-entropy loss meets the convergence condition.
[0091] Optionally, in this embodiment, the incremental sample set may refer to, but is not limited to, a dataset that stores incremental samples that meet the loss conditions. The incremental samples consist of historical scene data where the model predictions do not match the actual driving behavior, providing effective training data for the incremental learning of the behavior planning model and avoiding meaningless data from occupying computing resources.
[0092] Optionally, in this embodiment, the incremental sample may be, but is not limited to, a single scene data selected from historical driving data that meets the criteria of mismatch between model prediction and actual behavior and cross-entropy loss exceeding a threshold. Each sample contains at least a driving feature vector and an actual driving behavior label, providing a targeted correction basis for updating model parameters, thereby solving the prediction bias of the model in a specific scenario.
[0093] Optionally, in this embodiment, the target number threshold can be, but is not limited to, the minimum number of samples required to trigger a model update in the incremental sample set. Its purpose is to avoid model parameter oscillations caused by an insufficient sample size, ensuring that parameter updates are based on a sufficient amount of valid scenario data and improving optimization stability. For example, setting the target number threshold to 64 means that sample extraction for model updates only begins when the number of samples in the incremental sample set is ≥64.
[0094] Optionally, in this embodiment, the loss condition may be, but is not limited to, the criteria for determining whether historical driving behavior can constitute incremental samples. It may be, but is not limited to, the reference driving behavior predicted by the model not matching the actual driving behavior and the cross-entropy loss being greater than a preset threshold, thereby filtering out samples that are valuable for model optimization and excluding meaningless data that are accurately predicted or have small deviations.
[0095] Optionally, in this embodiment, historical driving behavior can be, but is not limited to, the actual driving operations performed by the current vehicle or reference vehicle in past driving cycles. Its function is to serve as a true benchmark for judging the accuracy of the model's predictions and to screen incremental samples.
[0096] Optionally, in this embodiment, the reference driving behavior may be, but is not limited to, the driving behavior predicted by the behavior planning model for a certain scenario at a historical time, and is compared with the historical driving behavior to determine whether there is a deviation in the model prediction, and then incremental samples are selected.
[0097] Optionally, in this embodiment, the convergence condition may be, but is not limited to, the criteria for determining whether the model parameter update has stopped, and may be, but is not limited to, the average cross-entropy loss of multiple consecutive iterations being limited to a preset threshold, or the loss decrease being less than a minimum threshold, thereby avoiding overtraining of the model and ensuring a balance between prediction accuracy and generalization ability.
[0098] Optionally, in this embodiment, gradient descent updates can be, but are not limited to, calculating model network parameters such as the gradients of weights and biases based on cross-entropy loss, and adjusting parameters along the negative gradient direction as an optimization method. By iteratively reducing the model's prediction loss, the prediction accuracy for the scenarios corresponding to incremental samples is improved.
[0099] Optionally, in this embodiment, it is first determined whether the number of samples in the incremental sample set reaches the target number threshold. If it does, a sample is extracted as the current incremental sample. The sample must meet the loss condition that the historical driving behavior does not match the reference driving behavior predicted by the model in the past and the cross-entropy loss exceeds the threshold. This ensures that the sample is valuable for model optimization, thereby avoiding model optimization instability caused by insufficient sample size. At the same time, targeted samples are selected to provide a valid basis for subsequent parameter updates and ensure that model optimization has a clear direction.
[0100] Next, the driving feature vector of the current incremental sample is input into the behavior planning model to obtain the predicted operation probability, and the cross-entropy loss is calculated with the actual behavior label in the sample. If the loss does not meet the convergence condition, the gradient descent algorithm is used to calculate the parameter gradient, and the network parameters such as the weights and biases of the model are adjusted along the negative gradient direction.
[0101] New incremental samples are drawn from the incremental sample set again, and the sample acquisition and parameter update steps are repeated. The cross-entropy loss is continuously calculated until the loss of multiple consecutive iterations meets the convergence condition, and the model update is stopped.
[0102] The embodiments provided in this application determine whether the number of samples in the incremental sample set reaches the target threshold before inputting data into the behavior planning model. If the threshold is reached, incremental samples that meet the loss conditions are extracted. When the sample cross-entropy loss does not meet the convergence condition, the model is updated by gradient descent, and this process is repeated until the loss converges. This achieves the technical effect of enabling the behavior planning model to be iteratively optimized based on effective incremental samples, avoiding performance problems caused by insufficient samples or overtraining, and continuously improving the prediction accuracy for complex scenarios. It realizes efficient incremental learning of the model and enhances the adaptability and reliability of autonomous driving lane change decisions.
[0103] As an optional approach, after determining the target driving behavior from at least two driving behaviors based on at least two operational probabilities, it also includes one of the following:
[0104] S4-1, trigger a first control signal that matches the first driving behavior in at least two driving behaviors, wherein the first control signal is used to control the current vehicle to accelerate in the current lane;
[0105] S4-2, trigger a second control signal that matches the second driving behavior in at least two driving behaviors, wherein the second control signal is used to control the current vehicle to decelerate in the current lane;
[0106] S4-3, trigger a third control signal that matches the third driving behavior in at least two driving behaviors, wherein the third control signal is used to control the current vehicle to accelerate into the reference lane;
[0107] S4-4, trigger a fourth control signal that matches the fourth driving behavior in at least two driving behaviors, wherein the fourth control signal is used to control the current vehicle to decelerate before entering the reference lane.
[0108] Optionally, in this embodiment, the first driving behavior may refer to, but is not limited to, the specific operation of accelerating in the current lane, which can guide the vehicle to perform acceleration operations in scenarios where there is no need to change lanes but it is necessary to improve the driving efficiency of the current lane.
[0109] Optionally, in this embodiment, the first control signal may be, but is not limited to, an instruction signal that matches the first driving behavior and is used to control the vehicle hardware to perform acceleration operations, which is generated by the autonomous driving control system and sent to the power system.
[0110] Optionally, in this embodiment, the second driving behavior may refer to, but is not limited to, the specific operation of decelerating in the current lane. When there is a risk in the current lane, such as an obstacle ahead or a reference vehicle cutting into the current lane, the vehicle is guided to decelerate to avoid the risk and avoid a collision.
[0111] Optionally, in this embodiment, the second control signal may be, but is not limited to, a deceleration command signal that matches the second driving behavior, and sent to the vehicle braking system to convert the deceleration intention into hardware operation and control the vehicle to reduce its speed in the current lane.
[0112] Optionally, in this embodiment, the third driving behavior may refer to, but is not limited to, the lane-changing operation of accelerating into the reference lane. The lane-changing driving behavior output by the model is such that there is a safe acceleration lane-changing window in the reference lane. For example, when the distance to the vehicle behind in the reference lane is far and the speed is slow, the vehicle is guided to quickly complete the lane change by accelerating, thereby improving the lane-changing efficiency.
[0113] Optionally, in this embodiment, the third control signal may be, but is not limited to, a composite acceleration and steering command signal that matches the third driving behavior, and is sent to both the power system and the steering EPS system. Its function is to coordinate the control of vehicle acceleration and steering to achieve fast and smooth lane changes.
[0114] Optionally, in this embodiment, the fourth driving behavior may refer to, but is not limited to, a lane-changing operation of decelerating and entering the reference lane. When the lane-changing window in the reference lane is narrow, or when the distance to the vehicle in front in the reference lane is close, the risk of lane changing can be reduced by first decelerating to adjust the distance before safely entering the reference lane.
[0115] Optionally, in this embodiment, the fourth control signal may be, but is not limited to, a deceleration and steering composite command signal that matches the fourth driving behavior. It is first sent to the braking system to achieve deceleration, and then sent to the steering EPS system to control steering, thereby controlling the vehicle to complete the lane change in stages and ensuring that the lane change process is safe and smooth.
[0116] Optionally, in this embodiment, the reference lane may refer to, but is not limited to, the target lane that the current vehicle plans to enter. It is usually an adjacent lane, such as the left lane or the right lane. It is the target of the third and fourth driving behaviors and its function is to clarify the target direction of the lane change and provide a basis for the steering angle and lane change path of the control signal.
[0117] Optionally, in this embodiment, after the behavior planning model determines that the target driving behavior is acceleration in the current lane, the autonomous driving control system generates and sends a first control signal to the vehicle power system. This signal includes specific acceleration parameters, such as target vehicle speed and torque output value, to control the vehicle to increase its speed without changing its driving direction in the current lane.
[0118] Optionally, in this embodiment, when the target driving behavior is deceleration in the current lane, the control system generates a second control signal and sends it to the braking system. The signal includes parameters such as the magnitude of the braking force and the vehicle speed after the target deceleration, controlling the vehicle to maintain its driving direction in the current lane and reduce its driving speed.
[0119] Optionally, in this embodiment, if the target driving behavior is a third driving behavior, such as accelerating into the reference lane, the control system generates a third control signal that simultaneously includes acceleration and steering commands, and coordinates the vehicle to accelerate and steer while entering the reference lane.
[0120] Optionally, in this embodiment, when the target driving behavior is to decelerate and enter the reference lane, the control system generates a fourth control signal in stages. First, a deceleration command is sent to the braking system. After the vehicle speed drops to the target value, a steering command is sent to the steering EPS system to control the vehicle to decelerate and adjust the distance before turning and entering the reference lane.
[0121] The embodiments provided in this application, by determining the target driving behavior and triggering the corresponding control signal according to the behavior type, transforms the driving behavior intention of the decision layer into an operation instruction that can be executed by the vehicle hardware, thereby ensuring that the vehicle accurately executes the target driving behavior, adapts to the efficiency and safety requirements in different scenarios, and improves the reliability and practicality of autonomous driving lane changing and driving within the lane.
[0122] As an optional option, obtain the driving style label of the reference vehicle, including one of the following:
[0123] S5-1, retrieve the driving style label sent by the reference vehicle;
[0124] S5-2, Obtain the driving description information set of the reference vehicle in the first driving cycle, wherein the driving description information set is used to indicate multiple driving decision behaviors of the reference vehicle in the first driving cycle; Input the driving description information set into the driving style classification model to obtain the driving style label.
[0125] Optionally, in this embodiment, the first driving cycle may refer to, but is not limited to, a specific time range used to analyze the driving style of the reference vehicle, such as the most recent hour or the past 50 kilometers of driving mileage, to limit the collection range of driving description information, ensure that the data is timely and representative, and avoid inaccurate style labels due to outdated data or an excessively large range.
[0126] Optionally, in this embodiment, the driving description information set may refer to, but is not limited to, a multi-dimensional data set describing the driving decision-making behavior of the reference vehicle within the first driving cycle, and may include, but is not limited to, relevant data such as lane change frequency, braking frequency, acceleration frequency, average acceleration, and following distance.
[0127] Optionally, in this embodiment, the driving decision-making behavior may be, but is not limited to, specific operational decisions made by the vehicle based on the environment during driving, such as frequency, frequency of emergency braking, and time for maintaining a safe following distance.
[0128] Optionally, in this embodiment, the driving style classification model may be, but is not limited to, a machine learning model deployed by the current vehicle for analyzing the driving description information of the reference vehicle and outputting style labels, which transforms the set of driving description information into structured driving style labels, thereby enabling the current vehicle to autonomously determine the style of the reference vehicle.
[0129] Optionally, in this embodiment, the current vehicle receives a generated driving style tag actively transmitted by the reference vehicle through communication protocols such as vehicle networking and Bluetooth. This tag is determined by the reference vehicle's own system based on its historical driving data, and the current vehicle can use it directly without any additional processing.
[0130] Optionally, in this embodiment, the driving description information of the reference vehicle during the first driving cycle is collected through the sensors of the current vehicle, such as radar, vision camera, or road network data platform, to form a data set; the data set is input into the driving style classification model pre-trained by the current vehicle, and the model outputs the driving style label of the reference vehicle through feature extraction and classification reasoning.
[0131] The embodiments provided in this application offer different options for obtaining reference vehicle driving style labels, thereby achieving the technical effect of adapting to different scenarios and flexibly obtaining reference vehicle style information. The efficiency of label acquisition can be guaranteed by direct reception, or the availability and accuracy of the labels can be guaranteed by autonomous analysis.
[0132] As an optional approach, after inputting the driving description information set into the driving style classification model to obtain the driving style label, the method further includes:
[0133] S6-1, Based on the trajectory position relationship between the first reference trajectory corresponding to the current vehicle in the spatiotemporal map and the second reference trajectory corresponding to the reference vehicle in the spatiotemporal map, determine the first position node and the second position node, wherein the spatiotemporal map is used to indicate the change of the object position of at least one vehicle object over time, and the at least one vehicle object includes the current vehicle and the reference vehicle.
[0134] S6-2, the first position node is used as the starting point of the estimated driving trajectory of the current vehicle, and the second position node is used as the ending point of the estimated driving trajectory.
[0135] S6-3, Based on the driving style label and the target driving behavior, determine at least one intermediate trajectory point in the estimated driving trajectory.
[0136] Optionally, in this embodiment, the driving description information set may refer to, but is not limited to, a multi-dimensional data set used to indicate multiple driving decision behaviors of the reference vehicle within the first driving cycle. This includes specific decision-related data such as lane change frequency, braking / acceleration frequency, and following distance, thereby providing input features for the driving style classification model and serving as the basis for obtaining the driving style label of the reference vehicle. For example, the driving description information set of the reference vehicle within the first driving cycle might be "lane change frequency 2 times / 10 km, average number of emergency braking times 1 time / hour, and average following distance 2 seconds," which fully reflects its driving decision characteristics.
[0137] Optionally, in this embodiment, the driving style classification model may refer to, but is not limited to, a model used to analyze the driving description information of a reference vehicle and output style labels, thereby transforming the unstructured set of driving description information into structured driving style labels, and achieving accurate determination of the driving style of the reference vehicle. For example, after inputting the driving description information set of the reference vehicle, such as "low lane change frequency and stable following distance," into the model, the model outputs a "mild" driving style label.
[0138] Optionally, in this embodiment, the training process of the driving style classification model is centered on labeled driving data. First, a large set of driving description information (including lane change frequency, braking / acceleration frequency, following distance, and other multi-dimensional data) of different vehicle objects in various driving scenarios is collected, and each data is labeled with a corresponding driving style label (such as "mild", "normal", "aggressive") to construct a training dataset covering different style characteristics. Then, the dataset is preprocessed, including removing outlier data, standardizing the data format to eliminate differences in units, and dividing the training set and validation set. Next, the preprocessed driving description information set is used as input features, and the corresponding style labels are used as output labels. The data is then input into a preset model architecture (such as a neural network or decision tree model) for iterative training. During the training process, the model parameters are adjusted to minimize the error between the predicted label and the true label. At the same time, the model performance is monitored in real time using the validation set. If overfitting occurs, regularization, data augmentation, and other methods are used to optimize the model until the classification accuracy, recall, and other indicators on the validation set reach the preset standards. The training is then completed, the model parameters are solidified, and a driving style classification model that can be directly used for driving style determination is formed.
[0139] Optionally, in this embodiment, the driving style classification model adopts a three-stage architecture of "input layer - feature extraction layer - classification output layer". The input layer receives a set of preprocessed driving description information and transforms multi-dimensional data such as lane change frequency, braking / acceleration frequency, and following distance into vector form that the model can recognize. The feature extraction layer is the core part of the model. It can automatically mine deep correlation features in the input data (such as the combination features of rapid acceleration and lane change frequency, the stability features of following distance, etc.) through structures such as fully connected layers, convolutional layers, or decision tree nodes, to realize the mapping from raw data to style features. The classification output layer receives the high-dimensional feature vector output by the feature extraction layer and transforms it into the probability distribution of corresponding driving style labels through activation functions such as Softmax. Finally, it outputs the style label with the highest probability (such as "aggressive"). Some models will also add dropout layers, normalization layers, etc. between the layers to improve the generalization ability and classification stability of the model, and ensure that the corresponding driving style label can be accurately output for driving description information in different scenarios.
[0140] Optionally, in this embodiment, the current vehicle may refer to, but is not limited to, the vehicle object performing the estimated driving trajectory planning. It is the subject of trajectory planning, thereby generating an estimated driving trajectory that conforms to the driving scenario based on reference vehicle information and its own needs. For example, an autonomous vehicle A traveling on a highway needs to plan a trajectory to overtake a reference vehicle ahead; vehicle A is the current vehicle.
[0141] Optionally, in this embodiment, the spatiotemporal diagram may refer to, but is not limited to, a visualization tool used to indicate the change of the position of at least one vehicle object over time. The horizontal axis represents the time dimension, and the vertical axis represents the spatial position dimension, thereby intuitively presenting the trajectory evolution pattern of the vehicle object and providing a visual basis for trajectory position relationship analysis. For example, the spatiotemporal diagram can clearly show the position changes of the current vehicle A and the reference vehicle B within 0-10 seconds, helping to determine their relative motion state.
[0142] Optionally, in this embodiment, the first reference trajectory may refer to, but is not limited to, the historical or preset trajectory corresponding to the current vehicle in the spatiotemporal map. This trajectory records the spatial position of the current vehicle at different points in time, thus serving as a benchmark for trajectory analysis and providing a basis for determining the starting point of the estimated driving trajectory. For example, the trajectory of the current vehicle A in the spatiotemporal map from 0 to 3 seconds is "position 5 meters at 0 seconds, position 15 meters at 1 second, position 25 meters at 2 seconds, and position 35 meters at 3 seconds," and this trajectory is the first reference trajectory.
[0143] Optionally, in this embodiment, the reference vehicle may refer to, but is not limited to, other vehicle objects that provide driving style reference for the current vehicle. Their driving status and style labels provide a reference for the current vehicle's trajectory planning, helping the current vehicle adapt to the surrounding traffic environment and generate a safer and more reasonable predicted trajectory. For example, if vehicle B is traveling in front of current vehicle A and its driving style label is "aggressive," then vehicle B is the reference vehicle.
[0144] The second reference trajectory can be, but is not limited to, the historical or real-time trajectory of the reference vehicle in the spatiotemporal map. It records the spatial position of the reference vehicle at different points in time, and is thus compared and analyzed with the first reference trajectory to provide a basis for determining the key nodes of the predicted driving trajectory. For example, the trajectory of reference vehicle B in the spatiotemporal map from 0 to 3 seconds is "position 10 meters at 0 seconds, position 20 meters at 1 second, position 30 meters at 2 seconds, and position 40 meters at 3 seconds". This trajectory is the second reference trajectory.
[0145] Optionally, in this embodiment, the trajectory positional relationship may refer to, but is not limited to, the relative positional characteristics of the first reference trajectory and the second reference trajectory in the spatiotemporal diagram, including distance, positional overlap trend, consistency of movement direction, etc., thereby serving as the basis for determining the first position node and the second position node. For example, if the distance between the first reference trajectory and the second reference trajectory is 5 meters at 2 seconds, and both move in the same direction, this characteristic constitutes the trajectory positional relationship between the two.
[0146] Optionally, in this embodiment, the first position node may, but is not limited to, a spatiotemporal node (containing time and spatial location information) determined based on the positional relationship between the first and second reference trajectories, serving as the starting point of the current vehicle's estimated driving trajectory. This clarifies the starting position of the estimated driving trajectory and ensures the continuity of trajectory planning. For example, by analyzing the trajectory positional relationship, the first position node may be determined as "position 35 meters at 3 seconds," meaning the current vehicle's estimated trajectory begins from this time and spatial location.
[0147] Optionally, in this embodiment, the second position node may, but is not limited to, a spatiotemporal node (containing time and spatial location information) determined based on the positional relationship between the first and second reference trajectories, serving as the endpoint of the current vehicle's estimated driving trajectory. This clarifies the termination position of the estimated driving trajectory and defines the target range for trajectory planning. For example, by analyzing the trajectory positional relationship, the second position node may be determined as "position 85 meters at 8 seconds," meaning the current vehicle's estimated trajectory needs to extend to this time point and spatial location.
[0148] Optionally, in this embodiment, the estimated driving trajectory may refer to, but is not limited to, the future driving path planned for the current vehicle from the starting point to the ending point of the trajectory, which includes a continuous sequence of multiple spatiotemporal nodes, thereby providing clear driving guidance for the current vehicle and ensuring that the driving process is safe, efficient, and meets expectations. For example, the estimated driving trajectory of the current vehicle A is "3 seconds 35 meters, 4 seconds 45 meters, 5 seconds 55 meters, 6 seconds 65 meters, 7 seconds 75 meters, 8 seconds 85 meters", which clearly defines the target position at different time points.
[0149] Optionally, in this embodiment, the trajectory starting point may, but is not limited to, the initial spatiotemporal node of the estimated driving trajectory, i.e., the first position node, thereby determining the starting point of the estimated trajectory and providing a reference starting point for subsequent trajectory planning. For example, the trajectory starting point of the estimated driving trajectory is "position 35 meters at 3 seconds", and the current vehicle starts executing the estimated trajectory from this node.
[0150] Optionally, in this embodiment, the trajectory endpoint may, but is not limited to, the termination spatiotemporal node of the estimated driving trajectory, i.e., the second position node, thereby determining the final target position of the estimated trajectory and clarifying the endpoint requirement of the trajectory planning. For example, if the endpoint of the estimated driving trajectory is "position 85 meters at 8 seconds", the current vehicle needs to travel along the estimated trajectory to this node.
[0151] Optionally, in this embodiment, the driving style label may refer to, but is not limited to, a structured identifier (such as "mild," "normal," or "aggressive") output by a driving style classification model to characterize the driving style of a reference vehicle. This provides a style adaptation basis for the current vehicle's trajectory midpoint planning, ensuring that the estimated trajectory is consistent with the reference vehicle's style. For example, if the reference vehicle's driving style label is "mild," the current vehicle needs to avoid aggressive position changes when planning midpoint trajectory points.
[0152] Optionally, in this embodiment, the target driving behavior may refer to, but is not limited to, the driving behavior preset and expected to be achieved by the current vehicle (such as "smooth following," "safe lane change," and "quick overtaking"), thereby providing behavioral guidance for the planning of intermediate points of the estimated driving trajectory and ensuring that the estimated trajectory meets the driving needs of the current vehicle. For example, the target driving behavior of the current vehicle may be "smooth following," and a safe distance from the reference vehicle must be maintained when planning intermediate trajectory points.
[0153] Optionally, in this embodiment, intermediate trajectory points may refer to, but are not limited to, key spatiotemporal nodes located between the trajectory start point and the trajectory end point that constitute the estimated driving trajectory, thereby connecting the trajectory start point and the trajectory end point and ensuring that the estimated driving trajectory is continuous, smooth, and meets the requirements of the driving style label and the target driving behavior. For example, the intermediate trajectory points of the estimated driving trajectory are "4 seconds 45 meters, 5 seconds 55 meters, 6 seconds 65 meters, 7 seconds 75 meters", and these nodes together constitute a continuous driving path.
[0154] Optionally, in this embodiment, the first reference trajectory (including spatial positions at different time points) recorded by the current vehicle in the spatiotemporal map and the second reference trajectory recorded by the reference vehicle in the same spatiotemporal map are first obtained. Then, the relative positional relationship (such as distance, movement trend, and overlap) of the two trajectories in the spatiotemporal dimension is analyzed. Based on this positional relationship, key nodes that meet the trajectory planning requirements are selected and used as the first position node and the second position node, respectively. The spatiotemporal map will synchronously present the position changes of the current vehicle and the reference vehicle over time to ensure that the analysis of the two trajectories is based on a unified spatiotemporal dimension.
[0155] Next, the time and space coordinates of the first position node are determined and set as the starting node (trajectory start point) of the current vehicle's estimated driving trajectory. Then, the time and space coordinates of the second position node are extracted and set as the ending node (trajectory end point) of the estimated driving trajectory, thus completing the definition of the beginning and end nodes of the estimated driving trajectory.
[0156] Finally, the driving style label of the reference vehicle (such as "mild") and the target driving behavior of the current vehicle (such as "smooth following") are obtained. Then, based on these two constraints, combined with the spatiotemporal information of the trajectory start and end points, multiple spatiotemporal nodes that meet the "style adaptation + behavior guidance" criteria are calculated. These nodes are the intermediate trajectory points of the predicted driving trajectory, and the intermediate trajectory points must ensure the continuity and smoothness from the trajectory start to the end point.
[0157] The embodiments provided in this application first determine the starting point (first position node) and ending point (second position node) of the estimated driving trajectory based on the positional relationship between the current vehicle and the reference vehicle in the same spatiotemporal map. Then, by combining the driving style label of the reference vehicle and the target driving behavior of the current vehicle, the midpoint of the trajectory is determined. This achieves the technical effect of generating a continuous, reasonable, and adaptable estimated driving trajectory for the current vehicle that is adapted to the surrounding traffic environment and its own driving needs. This ensures that the estimated trajectory not only conforms to the driving style characteristics of the reference vehicle, but also accurately achieves the target driving behavior of the current vehicle, thereby improving the pertinence and safety of trajectory planning.
[0158] As an optional approach, determining at least one intermediate trajectory point in the estimated driving trajectory based on the driving style label and the target driving behavior includes:
[0159] S7-1, Determine the first collision risk weight based on the driving style label;
[0160] S7-2, Determine the second collision risk weight based on the target driving behavior;
[0161] S7-3, based on the first collision risk weight and the second collision risk weight, at least one intermediate trajectory point in the estimated driving trajectory is determined sequentially.
[0162] Optionally, in this embodiment, the first collision risk weight may, but is not limited to, a weight coefficient (ranging from 0 to 1, with a larger value indicating a higher risk correlation) calculated based on the driving style label of the reference vehicle to quantify the collision risk caused by the reference vehicle's own driving style. This transforms the style characteristics of the reference vehicle into a quantifiable risk indicator, providing a risk constraint basis for the planning of intermediate trajectory points. For example, when the reference vehicle label is "mild," the first collision risk weight is set to 0.3, indicating that its driving style has a low impact on collision risk; when the label is "aggressive," the weight is set to 0.8, indicating that its style is prone to causing risk and needs to be avoided.
[0163] Optionally, in this embodiment, the second collision risk weight may, but is not limited to, a weight coefficient (ranging from 0 to 1, with a larger value indicating a higher risk) calculated based on the current vehicle's target driving behavior to quantify the collision risk associated with the target behavior itself. This transforms the current vehicle's driving requirements into a quantifiable risk indicator, ensuring that intermediate trajectory point planning satisfies the behavioral objective while keeping the risk within an acceptable range. For example, when the target behavior is "smooth following," the second collision risk weight is set to 0.2; when the target behavior is "rapid overtaking," the weight is set to 0.7, reflecting the difference in risk inherent in the behavior itself.
[0164] Optionally, in this embodiment, the estimated driving trajectory may refer to, but is not limited to, the future driving path planned for the current vehicle from the starting point to the ending point of the trajectory, which includes a continuous sequence of multiple spatiotemporal nodes, thereby providing a planning carrier for intermediate trajectory points. The intermediate trajectory points need to be distributed along the spatiotemporal range of the trajectory to ensure the continuity and integrity of the trajectory. For example, if the range of the estimated driving trajectory is "3 seconds 35 meters, 8 seconds 85 meters", the intermediate trajectory points need to be selected within this time and space interval, and the trajectory should remain smooth.
[0165] Optionally, in this embodiment, intermediate trajectory points may refer to, but are not limited to, key spatiotemporal nodes located between the trajectory start point and the trajectory end point, constituting the estimated driving trajectory, thereby connecting the trajectory start point and the trajectory end point. These points must be determined under the joint constraints of the first collision risk weight and the second collision risk weight to ensure that the selection of each node takes into account both risk control and driving requirements. For example, the intermediate trajectory points of the estimated driving trajectory may be "4 seconds 45 meters, 5 seconds 55 meters, 6 seconds 65 meters, 7 seconds 75 meters," and the spatiotemporal coordinates of each node are verified by risk weights to avoid collision risks with reference vehicles.
[0166] Optionally, in this embodiment, the driving style label of the reference vehicle (such as "mild" or "aggressive") is first obtained, and then based on the preset "style label-weight mapping rule" (which is derived from a large amount of driving data statistics, and different labels correspond to fixed or dynamically calculated weight coefficients), the first collision risk weight matching the label is calculated and determined. The weight value directly reflects the degree of influence of the reference vehicle style on the collision risk.
[0167] Next, the target driving behavior of the current vehicle is defined (such as "smooth following" or "rapid overtaking"). Then, based on the preset "target behavior-weight mapping rule" (which combines the risk level classification of behavior, and different behaviors correspond to weight coefficients that adapt to their safety requirements), the second collision risk weight matching the target behavior is calculated and determined. The weight value directly reflects the risk level of the target behavior itself.
[0168] Finally, based on the determined first collision risk weight and second collision risk weight, the comprehensive risk weight of the two is calculated (such as weighted summation, product, or other preset algorithms). Then, using the comprehensive risk weight as a constraint, and combined with the spatiotemporal information of the starting and ending points of the estimated driving trajectory, spatiotemporal nodes that meet the "risk controllable" requirement are sequentially screened or calculated. Each node must meet the "safe distance threshold corresponding to the comprehensive risk weight". These nodes are the intermediate trajectory points of the estimated driving trajectory.
[0169] The embodiments provided in this application first determine the first collision risk weight based on the driving style label of the reference vehicle, then determine the second collision risk weight based on the target driving behavior of the current vehicle, and finally determine the intermediate trajectory points of the estimated driving trajectory in sequence based on the comprehensive constraints of the two weights. This achieves "controllable risk quantification" in the planning of intermediate trajectory points, ensuring that each intermediate trajectory point is both adapted to the driving style characteristics of the reference vehicle and meets the target driving needs of the current vehicle, while minimizing the collision risk and improving the safety and rationality of the estimated driving trajectory.
[0170] As an optional approach, obtain the driving style tags of the reference vehicle, including:
[0171] S8-1, Obtain the driving deviation operation of the reference vehicle, wherein the driving deviation operation is used to indicate the driving operation that does not match the driving style label;
[0172] S8-2, If the reference vehicle's deviation operation in the second driving cycle meets the deviation conditions, adjust the reference vehicle's driving style label.
[0173] Optionally, in this embodiment, the deviation driving operation may refer to, but is not limited to, a driving operation performed by the reference vehicle during driving that does not match the currently labeled driving style tag, as a signal to determine whether the reference vehicle style has changed, so as to avoid the tag becoming invalid due to the change of the reference vehicle style.
[0174] Optionally, in this embodiment, the second driving cycle may refer to, but is not limited to, a specific time or mileage range used to statistically analyze the deviation of the reference vehicle from the driving operation, thereby limiting the statistical dimension of the deviation operation, avoiding unnecessary label adjustments triggered by a single accidental deviation, ensuring that the label adjustment is based on sufficient behavioral data, and improving the rationality of the adjustment.
[0175] Optionally, in this embodiment, the deviation condition may be, but is not limited to, the criteria for determining whether the reference vehicle driving style label needs to be adjusted. It may be, but is not limited to, the number of deviation driving operations within the second driving cycle being greater than or equal to a preset number threshold, or the proportion of deviation operations to total driving operations being greater than or equal to a preset proportion threshold, thereby filtering out scenarios where the reference vehicle style changes and avoiding accidental label adjustments due to accidental deviations.
[0176] Optionally, in this embodiment, adjusting the driving style label of the reference vehicle may, but is not limited to, updating the original driving style label to a new label that matches the current actual driving style based on the deviation driving operation characteristics of the reference vehicle in the second driving cycle, so as to ensure that the style label of the reference vehicle is always consistent with the actual behavior and avoid outdated labels causing the current vehicle decision deviation.
[0177] Optionally, in this embodiment, the driving operation of the reference vehicle is monitored in real time by the sensors of the current vehicle or the driving data sent by the reference vehicle, and these operations are compared with the current driving style label of the reference vehicle. Operations that do not match the label are filtered out, which are deviations from the driving operation.
[0178] Next, the number or proportion of deviation driving operations of the reference vehicle in the second driving cycle is counted. If the statistical results meet the preset deviation conditions, the original driving style label of the reference vehicle is updated to a new label based on the characteristics of these deviation operations.
[0179] It should be noted that by first acquiring the deviation driving operation of the reference vehicle that does not match the current driving style label, and then adjusting the driving style label of the reference vehicle when the deviation operation meets the deviation conditions in the second driving cycle, the technical effect of dynamically correcting the style label of the reference vehicle is achieved, ensuring that the label is consistent with the actual driving behavior. This avoids the current vehicle decision deviation caused by the change of the reference vehicle style but the label not being updated, and provides real-time and accurate style basis for subsequent anthropomorphic lane change decisions, thereby improving the safety and adaptability of autonomous driving decisions.
[0180] As an optional solution, in order to better understand the process of the above-mentioned autonomous driving behavior decision-making method, the following describes the execution flow of the above-mentioned autonomous driving behavior decision-making method in conjunction with optional embodiments, but it is not intended to limit the technical solution of the embodiments of this application.
[0181] With the rapid development of automotive electronics, advanced driver assistance systems (ADAS) have become a crucial research topic globally. Decision-making and planning are key modules in autonomous driving algorithms. This involves acquiring information about surrounding targets and obstacles from perception, localization, and prediction modules, as well as vehicle data, and processing this data. Combining the current vehicle state with environmental factors, decisions are made regarding the vehicle's current behavior. Decisions involving multi-vehicle interactions, particularly during lane changes, are particularly challenging. For example:
[0182] Scenario 1: The vehicle is currently maintaining its lane when a vehicle from another lane suddenly cuts into the vehicle's lane. The vehicle decides whether to yield or proceed directly.
[0183] Scenario 2: The vehicle needs to change lanes, but there is a vehicle in the target lane. The vehicle decides whether to let the other vehicle pass first or change lanes directly. This decision is made through speed planning.
[0184] During vehicle operation, speed planning needs to be able to change the relevant speed sequence along the path in a timely manner. From the current position of the vehicle to the target position, multiple speed curves can be planned, and the optimal curve with the lowest cost can be selected by defining a cost function.
[0185] However, the relevant technologies have the following problems: lane change decisions do not take into account the degree of cooperation of the participants. Although the lane change decision value is obtained through data-driven methods, the profile of the participants is not taken into account. For example, if the car behind is an "experienced driver", it is unlikely to give up space to change lanes, and may even accelerate and compress the lane change window to force the car behind it to give up changing lanes or choose the window behind it. This can lead to lane change failures, or even traffic accidents; rigid window selection decision-making style: the model function weights are fixed parameters, which cannot adapt to different driver style preferences (such as aggressive or conservative), resulting in system behavior that does not match driver expectations and reducing trust; huge computational consumption: in the spatio-temporal graph (ST), dense sampling and full-time dynamic planning of all obstacle trajectories are required, and the computational complexity increases exponentially with the scene complexity, placing an excessive burden on the onboard computing platform; difficulty in covering complex scenes: the window selection strategy does not fully consider the real-time changes of windows in dynamic traffic flow (such as other vehicles cutting in and out), resulting in insufficient decision reliability in complex congestion scenarios; the speed DP cost function design does not consider the profiles and types of participants: when planning based on ST graph DP+QP, the weights of the cost function are the same set of parameters, without considering the profiles and types of traffic participants.
[0186] It should be noted that, for example, if the vehicle behind is an "experienced driver," then the cost definition for selecting the Node above this vehicle (overtaking) during the DP process on the ST diagram should be slightly higher than in the normal scenario; otherwise, it will cause the same problem of squeezing lane-changing space as in Problem 1. For example, using the same set of cost weights for different types of participants will result in a slightly stronger sense of pressure on the VRU, which does not conform to the driving habit of human drivers who prefer to maintain a sufficient safe distance from the VRU. For example, using the same set of cost weights for different types of participants will result in the same planning logic for trucks as for other types, which does not conform to the driving habit of human drivers who are unwilling to change lanes behind trucks.
[0187] In this embodiment, in order to solve the problems existing in the related technologies, multiple problem-solving stages are included. In the first stage: Agent profiling, the core of this stage is to solve the problem of "who is the other party". This method automatically identifies and classifies the driver's driving style into three types: "aggressive", "normal" or "mild" by analyzing the driver's historical behavior data, and provides probabilistic output for online real-time inference.
[0188] The core of the first phase is a two-stage machine learning framework, which includes offline model training and online style recognition.
[0189] Offline model training includes: data preprocessing and feature engineering, feature standardization, unsupervised clustering, and supervised classification model training.
[0190] Data preprocessing and feature engineering: The system collects raw data from the vehicle's CAN bus and sensing system (such as vehicle speed, acceleration, steering angle, etc.), and divides it into continuous historical segments using a sliding window slicing method. From each segment, features that significantly reflect driving style are extracted: lane change frequency, braking / acceleration frequency, mean absolute acceleration / deceleration, mean / standard deviation of acceleration, standard deviation of lane centering error, following distance, and longitudinal speed relative to the vehicle ahead via TTC.
[0191] Feature standardization: The extracted multidimensional feature vectors are standardized to eliminate the influence of dimensions and prepare for subsequent cluster analysis.
[0192] Unsupervised clustering: The K-Means clustering algorithm is used to perform unsupervised learning on the standardized feature data, automatically clustering driving behavior patterns into 3 categories. Experts assign semantic labels of "aggressive," "normal," or "mild" to each category based on the feature performance of the cluster centers.
[0193] Supervised classification model training: Using the clustering results as labeled data, a supervised classifier (such as a Support Vector Machine, SVM) is trained. This step aims to learn the complex mapping relationship from feature vectors to driving styles, resulting in a high-precision classification model that can be deployed online.
[0194] The second stage is personalized behavioral decision-making. The core of this stage is to solve the question of "what to do", that is, whether the vehicle should choose aggressive or compromising behavior in the interaction scenario.
[0195] like Figure 3 As shown, Figure 3 (a) in the diagram represents a scenario where the vehicle wants to change lanes to the right to exit the ramp and there are adjacent vehicles in the target lane. The vehicle is vehicle 302, and the adjacent vehicle is vehicle 304. The vehicle has two lane-change window options: one is... Figure 3 (b) shows the selection of the target lane in front of the vehicle window, i.e., aggressive behavior, represented by vehicle 302 changing lanes in front of vehicle 304; the second is as follows Figure 3 As shown in (c), selecting the window behind the vehicle in the target lane is a compromise behavior, which is represented by vehicle 302 changing lanes behind vehicle 304.
[0196] like Figure 4 As shown, Figure 4 (a) in the diagram represents a scenario where an adjacent vehicle is changing lanes to the right to exit a ramp and the target lane contains the vehicle itself. The adjacent vehicle is vehicle 402, and the vehicle itself is vehicle 404. The vehicle has two driving window options: one is... Figure 4 As shown in (b), the vehicle (vehicle 404) slows down to avoid the adjacent vehicle (vehicle 402) changing lanes in front of the vehicle (vehicle 404), and the second is as follows. Figure 4The vehicle (vehicle 404) shown in type (c) accelerates, causing the adjacent vehicle (vehicle 402) to change lanes from behind the vehicle (vehicle 404).
[0197] The decision on what behavior to take is based on parameters from the current scene, such as the vehicle's speed and acceleration, the target vehicle's speed and acceleration, the size of the window, and the agent profile derived in the first stage. These parameters are then input into the model to determine the behavior.
[0198] Decision-making mechanism: The behavioral decision-making model is a classification model based on a multilayer perceptron (MLP). Its input layer nodes correspond to a set of standardized scene feature vectors F_behavior, including: the relative speed difference Δv between the vehicle and the target vehicle, the relative distance d, the estimated collision time TTC, and the target type encoding (one-hot encoding, such as [1,0] representing a vehicle, [0,1] representing a VRU), etc. The output layer consists of two nodes, representing the probability of 'aggressive' behavior P_aggressive and the probability of 'conservative' behavior P_conservative, respectively. The final behavior is determined by the higher probability.
[0199] The factory settings of the behavioral decision model: The initial parameters of the model are trained based on massive road data or shadow pattern data, representing a universal safe driving style.
[0200] Online Adaptation: After the vehicle is delivered to the user, the system continuously records the driver's actual choices in specific scenarios using a shadow mode (e.g., whether the driver chooses to decelerate or maintain speed during a cut-in). Once enough effective data has been accumulated, the system uses an incremental learning algorithm (e.g., online logistic regression update based on stochastic gradients) to adjust the decision boundary of the classification model, gradually approximating the driver's actual behavioral preferences, thereby achieving highly human-like decision-making.
[0201] The third stage is efficient motion planning. The core of this stage is to solve the "how to do it" problem under the constraints of the decisions made in the first stage, and to generate smooth and safe trajectories with extremely low computing power.
[0202] Control quantity generation: Based on the behavior determined in the second stage (such as "aggressive lane change"), window information and environmental characteristics, as well as the agent profile determined in the first stage (such as "aggressive"), an ST diagram is built. The target and the predicted future trajectory are projected onto the ST diagram. DP+QP is used to obtain reference values for control quantities such as longitudinal / lateral acceleration and jerk of the vehicle.
[0203] Low-computing-power trajectory optimization: In longitudinal velocity planning, this embodiment abandons the method of densely sampling all obstacle trajectories throughout the ST diagram and proposes a key spatiotemporal event point sampling method. Specifically, the system only focuses on key event points where surrounding dynamic obstacles interact with the planned path of the vehicle, mainly including "entry" and "exit" points. These points define the start and end times of obstacle occupation of the conflict area.
[0204] When constructing the DP (Dynamic Programming) search graph, nodes are only set near the spatiotemporal coordinates corresponding to these key event points. In this way, the number of nodes that DP needs to search is significantly reduced from O(NT) (N is the number of obstacles, T is the time step) in the traditional method to O(M) (M is the number of key events, M << NT), which significantly reduces the computational complexity and enables the algorithm to run in real time on a low-cost computing platform.
[0205] It should be noted that in this embodiment, the participant's agent profile is considered: based on the analysis of the traffic participant's historical driving behavior, driving style is evaluated and categorized into three labels: "aggressive," "normal," and "mild." Subsequent lane-changing decisions, derived through data-driven methods, take the participant's profile into account, using the profile labels as one of the inputs to the window selection classification model. This aims to minimize the selection of windows before aggressive drivers, avoiding situations where drivers fail to proactively create space for lane changes, or even accelerate and compress lane-changing windows, forcing their own vehicles to abandon the lane change or choose windows after them. This could lead to lane-changing failures or even traffic accidents.
[0206] Window decision-making style is anthropomorphized: including factory settings and online adaptation.
[0207] Among them, factory settings: The initial parameters of this model are trained based on massive road data or shadow mode data, representing a universal safe driving style.
[0208] Online Adaptation: After the vehicle is delivered to the user, the system continuously records the driver's actual choices in specific scenarios using a shadow mode (e.g., whether the driver chooses to decelerate or maintain speed during a cut-in). Once enough valid data has been accumulated, the system employs an incremental learning algorithm (e.g., online logistic regression update based on stochastic gradients) to adjust the decision boundary of the classification model, gradually approximating the driver's actual behavioral preferences in terms of the model's output probability distribution, thus achieving highly human-like decision-making.
[0209] Low computational cost: In longitudinal speed planning, this embodiment abandons the traditional method of densely sampling all obstacle trajectories throughout the ST diagram. It innovatively proposes a key spatiotemporal event point sampling method. Specifically, the system only focuses on key event points where surrounding dynamic obstacles interact with the planned path of the vehicle, mainly including "entry" and "exit" points. These points define the start and end times of obstacle occupation of the conflict area, significantly reducing computational requirements.
[0210] Complex scenarios can be covered: The window selection strategy takes into account the real-time changes of the window in the dynamic traffic flow (such as other vehicles entering and exiting), and the decision is reliable in complex congestion scenarios.
[0211] Transforming "identity" and "intention" into cost weights: The cost function considers the participant's profile and type. When planning using DP+QP based on the ST graph, the profile and type of traffic participants are taken into account. For example, if the car behind is an experienced driver, the cost of selecting the Node above this car (overtaking) during the DP process on the ST graph should be defined slightly higher than in a typical scenario. Furthermore, the avoidance cost weight for VRUs should be greater than for other target types, allowing the system to decelerate earlier and more smoothly, aligning with human drivers' ethics and habits. The code that yields to trucks should have a greater weight to avoid changing lanes behind trucks, also consistent with human drivers' habits.
[0212] Further examples, such as Figure 5 As shown, this is a low-computing-power anthropomorphic autonomous driving decision-making and planning method and system based on driver profiles. Its components and their functions are as follows: Figure 5 As shown: including radar sensor 502, body sensor 504, vision sensor 506, environment model 508, target selection module 510, trajectory prediction module 512, behavior decision module 514, control quantity planning module 516, steering assist system 518, steering EPS system 520, braking system 522 and power system 524.
[0213] The radar sensor, vehicle body sensor, and vision sensor are used to detect surrounding target information, such as lane lines, obstacles, and road signs. The map EHR module outputs navigation map information, including map reference lines, positioning information, and information on road convergence and divergence.
[0214] Environment Model: The environment model module is a crucial component providing road environment information across various functions. It receives visual and high-precision map information, performs data parsing and processing, and outputs information such as road separation points, reference lines, and lane lines required by the functions.
[0215] Target selection module: S9-1, through coordinate system transformation, calculates the information of the vehicle and each input target in the SL coordinate system, mainly including position, velocity, acceleration, etc.
[0216] S9-2 filters out four targets in the vehicle's lane: compare the lateral bounding boxes of the vehicle's lane with those of surrounding targets to see if they overlap. If they overlap, select the preceding and following vehicles based on their longitudinal distance. When determining the size of the bounding box, factors such as lane changing via the vehicle's lever, active lane changing, and driver steering wheel movements are all taken into account.
[0217] S9-3, after the target selection in the vehicle's own lane is completed, the target selection in adjacent lanes and the next adjacent lane begins: by comparing the lateral distances of each target to the center lines of each lane, the lane in which the target is located is determined, thus determining the lateral relative position of the target to the vehicle; then, the targets are sorted by longitudinal distance to determine their relative positions to the vehicle. After the above operations, the target selection module outputs information on 16 targets around the vehicle.
[0218] Further examples, such as Figure 6 As shown, vehicle 602 is a private vehicle, vehicles 604 and 606 are vehicles in the lane to the right of vehicle 602, vehicle 608 is the vehicle in front of vehicle 602, and vehicle 610 is the vehicle to the left of vehicle 602.
[0219] Trajectory Prediction: Before predicting the target's trajectory, the target's behavior is predicted. Behavior prediction involves processing information about each target vehicle relative to its own lane centerline and the target's current state. The information about a target vehicle relative to its own lane centerline mainly includes the distances from the vehicle's four rear left, front left, rear right, and front right edge points to its lane centerline, its lane ID, and its secondary lane ID. The vehicle's direction of travel is taken as the positive longitudinal (x-axis) direction, and the left side perpendicular to the longitudinal direction is taken as the positive lateral (y-axis) direction. If the lateral distance of the target vehicle's left front edge point is greater than the x-coordinate y of the left lane line, and its right rear edge point is less than the x-coordinate y of the left lane line, then its secondary lane is the left lane (left turn) of its current lane. If the lateral distance of the target vehicle's left front edge point is greater than the x-coordinate y of the right lane line, and its right rear edge point is less than the x-coordinate y of the right lane line, then its secondary lane is the right lane (left turn) of its current lane. If the lateral distance of the target vehicle's right front edge point is less than the x-coordinate y of the right lane line, and its left rear edge point is greater than the x-coordinate y of the right lane line, then its secondary lane is the right lane (right turn) of its current lane. If the lateral distance q of the target vehicle's right front edge point is less than the x-coordinate y of the left lane line, and its left rear edge point is greater than the x-coordinate y of the left lane line, then its secondary lane is the left lane (right turn) of its current lane.
[0220] Further examples, such as Figure 7 The diagram shows a preprocessing visualization of target vehicle behavior prediction. With the vehicle's driving direction as the positive longitudinal (x-axis) and the left side perpendicular to the longitudinal direction as the positive lateral (y-axis), it illustrates the positional relationship between target vehicle 702 and the center line and lane lines of its lane via four edge points: left rear, left front, right rear, and right front. It also shows four scenarios for determining the secondary lane ID: (e.g., ...) Figure 7 As shown in (a), the lateral distance of the left front edge point of vehicle 702 is greater than the y value of the left lane line and the right rear edge point is less than the value, and the corresponding secondary lane is the left lane (left turn) of the lane it is in. Figure 7 As shown in (b), the lateral distance of the left front edge point of vehicle 702 is greater than the y-value of the right lane line and the right rear edge point is less than the value, and the corresponding secondary lane is the right lane (left turn) of the lane it is in. Figure 7 As shown in (c), the lateral distance of the right front edge point of vehicle 702 is less than the y-value of the right lane line, and the left rear edge point is greater than the value. The corresponding secondary lane is the right lane (right turn) of the lane it is in. Figure 7 As shown in (d), the lateral distance of the right front edge point of vehicle 702 is less than the y value of the left lane line, and the left rear edge point is greater than the value. The corresponding secondary lane is the left lane (right turn) of the lane it is in. These scenarios intuitively present the determination logic of the secondary lane ID through the positional relationship between the edge point and the lane line, providing basic information for subsequent target behavior prediction.
[0221] There are three main target states: LaneKeep, TurnLeft, and TurnRight. If the target vehicle is a car, and its left front wheel crosses the left lane line by 0.26 meters (a calibrable variable) and its lateral speed exceeds the threshold of 0.3 m / s (a calibrable variable), then it is determined that it intends to turn left. If the target vehicle is a truck, and its left front wheel crosses the lane line by 0.4 meters (a calibrable variable) and its lateral speed exceeds the threshold of 0.3 m / s (a calibrable variable), then it is determined that it intends to turn left. The principle for determining whether a target vehicle intends to turn right is the same as the logic for determining whether it intends to turn left.
[0222] After determining the predicted behavior of the target vehicle, the trajectory is then predicted, ultimately resulting in a trajectory containing num = 40 points (variables can be labeled). When predicting the trajectory, both its lateral and longitudinal positions need to be predicted separately.
[0223] Longitudinal: Given the target's velocity and acceleration information, with a time interval of 0.2s (the variable can be calibrated), according to formula (1), we can obtain:
[0224] (1);
[0225] Where vx is its longitudinal velocity and ax is its longitudinal acceleration, the longitudinal position xt of the target vehicle within 8s (the variable can be calibrated) can be obtained; according to formula (2), we can get:
[0226] (2);
[0227] Its longitudinal speed vx can be obtained. The maximum longitudinal speed limit needs to consider two aspects: ensuring that the predicted speed does not exceed the lane speed limit and maintaining a safe distance from the vehicle in front. If the calculated speed is greater than the lane speed limit, then the lane speed limit is taken as the speed, as shown in formula (3):
[0228] (3);
[0229] in This indicates the speed limit of the road, and this method is used to ensure that the predicted speed does not exceed the lane speed limit;
[0230] To avoid a collision with the vehicle in front, the longitudinal acceleration of the target vehicle is predicted by considering the speed and distance of the vehicle in front. This longitudinal acceleration is determined by looking up a table based on empirical values. Two nine-dimensional arrays, TableDeltaR and TableDeltaV, are used to store the distance difference (m) and speed difference (m / s) between the target vehicle and the vehicle in front, respectively.
[0231] TableDeltaR=[1 5 10 15 20 25 30 35 40];
[0232] TableDeltaV=[1 4 8 12 14 16 18 20 21] (Variables can be labeled).
[0233] The acceleration values can be obtained from Table 1. The values in the table are determined based on practical experience (the variables can be calibrated).
[0234] Table 1
[0235]
[0236] Assuming the target vehicle will be centered within the target lane after changing lanes, its final lateral position is... The x-coordinate of the road centerline of the target lane, and the lateral velocity. The value is 0. The initial lateral position of the target vehicle is known. lateral velocity Its lateral position relative to the final state lateral velocity We need to find its lateral position Y and lateral velocity V_y within 8 seconds.
[0237] The four control points are , , , . , , , Therefore, its lateral position can be obtained from formula (4):
[0238] (4);
[0239] Here, Y and t are both arrays of dimension 40, and the value of the j-th element of t is... .
[0240] The longitudinal velocity of the target vehicle can be obtained from formulas (5) and (6). :
[0241] (5);
[0242] (6).
[0243] This embodiment also includes a driver profile classification module (unsupervised clustering + supervised classification): this module consists of two parts: offline training and online classification.
[0244] Optional, such as Figure 8 The diagram shows the process flow of data processing and model training. The process proceeds sequentially through S802 data acquisition, S804 data analysis, S806 data cleaning, S808 data slicing, S810 feature extraction, S812 standardization, S814 feature selection, S816 dimensionality reduction, S818 model training, and S822 labeling, finally entering S824 model training. This fully presents the entire process from raw data acquisition to feature processing and model training.
[0245] In this embodiment, the collected data is analyzed, and the data of the region of interest is selected for sliding window slicing to extract the feature quantities that can reflect the driver profile, such as: lane change frequency, braking / acceleration frequency, mean absolute acceleration / deceleration, average value / standard deviation of acceleration, standard deviation of lane centering error, following distance, and TTC longitudinal speed with the vehicle in front.
[0246] After feature extraction, the features are standardized, and then unsupervised learning algorithms such as K-Means or Gaussian Mixture Models (GMM) are used for clustering. The clustering results are then labeled (aggressive, general, mild). GMM is a soft clustering algorithm that provides the probability that a sample belongs to each cluster.
[0247] Then, model training is performed to obtain the weight values of each feature. The process, or components, are as follows: Figure 9 As shown, the algorithm consists of three core algorithm modules: Feature extraction algorithm 902 obtains sample data through road sampling data, data analysis, data slicing, and feature extraction; Clustering analysis algorithm 904 completes the corresponding processing based on model training and label addition; Classification algorithm 906 implements the function through model training, model selection, and weight data. Each of the three modules contains corresponding processing steps, which together support the data processing and model-related task flow.
[0248] Clustering is performed on the normalized feature values. After clustering, classification algorithms such as Support Vector Machine (SVM), KNN, and classification tree can be selected for classification. The classification model is trained and iterated to obtain the weight values of the classification network, as shown in formulas (7), (8), and (9).
[0249] (7);
[0250] (8);
[0251] (9);
[0252] In the formula, exp is an exponential function with the natural constant e as its base. , , These represent the probability values for the driver's aggressive, average, and mild driving profiles, respectively. , , The values represent the linear classification values of the driver's aggressive, normal, and mild driving profiles, respectively. These values are the results of the previous clustering step. Then, the regression equation coefficients are obtained based on formulas (10), (11), and (12).
[0253] (10);
[0254] (11);
[0255] (12);
[0256] in, , , The regression weights are assigned to each feature of the driver's aggressive, average, and mild driving profiles. , , These represent the intercepts of the linear classification equations for the driver's aggressive, average, and mild driving profiles, respectively.
[0257] Online classification: Online classification first obtains the model's input values, then performs feature preprocessing, and finally uses a classification algorithm to classify the data, outputting probability values for three driver profile labels. The specific equations are shown in reverse in the above formula, and will not be elaborated further. Feature extraction steps are as follows: Figure 10 As shown, the process includes S1002, data preprocessing; S1004, scene analysis; S1006, feature extraction; and S1008, feature normalization.
[0258] Personalized Behavior Decision Stage: The core of this stage is to solve the question of "what to do", that is, whether the vehicle should choose aggressive or compromising behavior in the interaction scenario.
[0259] The desired behavior is determined by parameters from the current scene, such as the vehicle's speed and acceleration, the target vehicle's speed and acceleration, and the size of the window. These parameters are then input into the model to derive the behavior.
[0260] Decision-making mechanism: The behavior decision-making model is a classification model based on a multilayer perceptron (MLP). Its input layer nodes correspond to a set of standardized scene feature vectors F_behavior, including: the relative speed difference Δv between the vehicle and the target vehicle, the relative distance d, the estimated collision time TTC, and the target type encoding (one-hot encoding, [1,0] represents a vehicle, [0,1] represents a VRU), etc. The output layer consists of two nodes, representing the probability of 'aggressive' behavior P_aggressive and the probability of 'conservative' behavior P_conservative, respectively. The final behavior is determined by the higher probability. The main features, i.e., feature symbols, are described in Table 2:
[0261] Table 2
[0262]
[0263] Feature standardization: These features have different physical meanings and units (e.g., distance is in meters, speed is in meters per second), and must be standardized, otherwise it will affect the stability and convergence speed of model training. This embodiment uses Z-score standardization, as shown in formula (13):
[0264] (13);
[0265] in and These are the mean and standard deviation of each feature on the training dataset. These values are calculated using massive amounts of data at the time of manufacture. and Standardize it.
[0266] Factory Setup: S10-1, Data Collection: Collect massive amounts of shadow pattern data of human drivers, including scene features F_behavior and the driver's true behavior label y_true in that scene ([1, 0] represents attack, [0, 1] represents compromise). S10-2, Training: Train the MLP network using the backpropagation algorithm and cross-entropy loss function. The training goal is to make the model's predicted probability distribution [P_agg, P_con] as close as possible to the distribution of the true label y_true. S10-3, Deployment: After training converges, the obtained network weight matrix (Weights) and bias vector (Biases) are solidified into the vehicle-side system program. This is the "factory setup".
[0267] Online Adaptive: After the vehicle is delivered to the user, the system continuously records the driver's actual choices in specific scenarios using a shadow mode (e.g., whether the driver chooses to decelerate or maintain speed during a cut-in). Once enough effective data has been accumulated, the system uses an incremental learning algorithm to adjust the decision boundary of the classification model, gradually bringing the model's output probability distribution closer to the driver's actual behavioral preferences, thereby achieving highly human-like decision-making.
[0268] The system will not update the model for every decision, as this would introduce too much noise and involve frequent computation. It requires a reasonable triggering mechanism, which is as follows in this embodiment:
[0269] S11-1, Recording and Comparison: After the system executes the decision, it will continuously record the driver's actual behavior y_true (i.e. what the driver actually did).
[0270] S11-2, Calculate uncertainty / discrepancy: The system re-inputs the current scene feature F_behavior into the factory model to obtain the predicted probability [P_agg_factory, P_con_factory] of the factory model. Key indicator: CrossEntropy H. It measures the difference between the predicted probability distribution of the factory model and the driver's actual "distribution". Since the actual behavior is a definite value (such as [1, 0]), the crossentropy is calculated as shown in formula (14):
[0271] (14);
[0272] For example, if the actual behavior is an attack (y_true = [1, 0]), and the factory model predicts P_factory = [0.3, 0.7], then H = -log(0.3) ≈ 1.2. If the prediction is [0.8, 0.2], then H = -log(0.8) ≈ 0.22. The larger the H value, the more uncertain the factory model's prediction is, and the greater the discrepancy between the prediction and the driver's behavior.
[0273] S11-3, Triggering condition: Set a cross-entropy threshold H_threshold (0.8). When H > H_threshold, it indicates a significant discrepancy between the factory model's decision-making and the driver's habits in the current scenario. This data (F_behavior, y_true) is stored in an experience replay buffer. Incremental learning is enabled: When the amount of data in the buffer reaches a batch size B (e.g., B = 64), an incremental learning process is triggered.
[0274] Optional, such as Figure 11 As shown: S1102, feature extraction and standardization are performed in the real-time scene; S1104, MLP forward propagation inference; S1106, output the probability of compromise behavior and the probability of attack behavior; S1108, determine whether the probability of compromise behavior is greater than the probability of attack behavior. If the result is "yes", execute S1112; if it is "no", execute S1110; S1110, execute the attack behavior; S1112, execute the compromise behavior; S1114, record the driver's actual behavior; S1116, determine whether incremental learning is triggered. If the result is "yes", execute S1118; if it is "no", return to S1102; S1118, incremental learning process.
[0275] The specific method of incremental learning: The goal of incremental learning is to fine-tune the large factory model using a small amount of new data from new car owners, rather than retraining it. The algorithm used in this embodiment is based on online learning using Mini-batch Stochastic Gradient Descent (SGD).
[0276] Loss Function: Cross-entropy loss is used. For a batch of B data points {(x_i, y_i)}, the loss function is as shown in formula (15):
[0277] (15);
[0278] in It is the probability vector output by the current MLP model (with the latest parameters θ).
[0279] Parameter update formula: The goal is to minimize the loss function and update the network parameters through gradient descent, as shown in formula (16):
[0280] (16);
[0281] in: : A set of parameters containing all weights and biases of the MLP. : Learning rate. This is a hyperparameter, 0.0001, to ensure that learning is stable and gradual. Loss function with respect to parameters The gradient is calculated through backpropagation.
[0282] Incremental learning process: S12-1, retrieve a batch of data from the experience buffer. S12-2, use the current model to predict this batch of data and calculate the loss. S12-3, perform one backpropagation and calculate the gradient. S12-4, update the parameters of the current model using the above formula. S12-5, clear the buffer and wait for the next trigger.
[0283] The efficient motion planning includes five modules: ST construction, target projection, node selection, DP curve generation, and cost calculation. Specifically, ST construction determines the parameter information of the constructed position-time graph, including the maximum value range of the horizontal and vertical axes. This patent selects a maximum value of 8s (calibration value) for the horizontal axis time and 100m (calibration value) for the vertical axis.
[0284] Target Projection: This module projects the predicted trajectory of surrounding targets and the target's trajectory over the next 8 seconds (calibrated value, TBD) onto the ST map. Each projected target is bound to its attributes: type and profile. The key to projection lies in two points: the targets that can be projected and how long the projection can last. A target only begins to be projected if a surrounding target encroaches on the driver's lane or if the driver changes lanes into the target vehicle's lane. Similarly, a target does not begin to be projected if a surrounding target cuts out of the driver's lane or the driver changes lanes to another lane. In summary, if a target interacts with the driver, the target's trajectory will be projected onto the ST map.
[0285] Node selection: This module selects two Nodes for the trajectory of the target projection. The purpose of these nodes is to be used for dynamic planning traversal in subsequent steps. The Node selection method is to select the first point of the trajectory and the first point after the trajectory ends.
[0286] Cost calculation: In related technologies, the ST-based dynamic programming algorithm has the following problems:
[0287] Question 1: In the ST diagram, overtaking means the curve is above the obstacle curve. If the car behind is an aggressive, experienced driver, it is likely that it will not slow down but will choose to approach. In this case, if the car behind attempts to overtake (choosing the upper node), the risk is greater.
[0288] Question 2: For example, using the same set of cost weights for different types of participants results in a slightly stronger sense of pressure on the VRU, which does not conform to the driving habits of human drivers who tend to keep a safe distance from the VRU.
[0289] Question 3: For example, using the same set of cost weights for different types of participants results in the same planning logic for trucks as for other types, which does not conform to the driving habits of human drivers who are unwilling to change lanes behind trucks.
[0290] To address the first issue mentioned above, and specifically for "aggressive" following vehicles, the cost of overtaking decisions is increased by adjusting the weight of the collision cost term. When the vehicle's decision is "aggressive" and the target vehicle's profile is "aggressive," its weight is significantly increased. This causes the cost value of any node point less than the safety threshold T_Distance to increase sharply for aggressive following vehicles. To find the lowest cost, the DP algorithm will tend to choose nodes farther away from the vehicle, meaning it is more likely to avoid passing over it, thus achieving a human-like approach of "being more cautious about overtaking behavior."
[0291] Regarding questions two and three above: For different types of participants (VRUs, trucks), the safe distance threshold T_Distance and offset cost are dynamically adjusted. For VRUs, the threshold becomes very large, causing the DP process to incur significant costs at greater distances, forcing the planner to slow down earlier and maintain a larger margin. For trucks, on the one hand, the safe distance increases, and on the other hand, a high fixed cost is imposed on the path scheme "following behind the truck," causing the DP algorithm to actively avoid these nodes, choosing to change lanes to the front of them or not to change lanes at all, perfectly replicating the preferences of human drivers.
[0292] After traversing all nodes, multiple curves are generated. The performance of the curves is evaluated by calculating the Cost. Factors to consider include: recommended speed cost, acceleration cost, collision cost, distance to the target offset, and travel efficiency (time taken to reach the same distance / distance traveled in the same time). For the i-th node currently being calculated and related to obstacle j, the Cost calculation process is as follows:
[0293] The method for calculating the portrait factor is shown in formula (17):
[0294] (17);
[0295] in, For image lookup table functions, A portrait of the j-th target.
[0296] The method for calculating the type safety factor is shown in formula (18):
[0297] (18);
[0298] in, This is a type lookup table function. Let j be the type of the j-th target.
[0299] The dynamic safety factor is calculated as shown in formula (19):
[0300] (19);
[0301] in, To fix the safety distance threshold, It is a speed-related coefficient that simulates the common sense that the safe distance increases with increasing speed.
[0302] Finally, the Cost function of the current i-th Node is obtained as shown in formula (20):
[0303] (20);
[0304] Where exp is an exponential function with the natural constant e as its base. This indicates the recommended speed for that lane. / This represents the current Node velocity and acceleration. / This represents the velocity and acceleration of the previous Node, and T_Distance represents the threshold distance at which a safe distance from the target is considered. This indicates the actual distance between the Node and the target at each point. If the distance between the Node and the obstacle is less than a threshold value... When the distance is large, the cost of distance can be ignored. This indicates whether the current node is in front of the truck.
[0305] The total cost along this path is as shown in formula (21):
[0306] (twenty one);
[0307] Where N is the number of Node nodes.
[0308] Optimal path selection: After calculating the cost of all curves, this module selects the curve with the lowest cost as the optimal speed curve;
[0309] Decision-making output: Based on the position of the curve relative to the target, determine whether the vehicle should yield or proceed straight. For lane-changing scenarios, if the optimal curve is below the target, yield; if it is above the target, change lanes directly. For scenarios where the target is cutting in, if the optimal curve is below the target, decelerate and yield; if it is above the target, proceed directly.
[0310] Steering EPS module: Angle control ensures that the vehicle's driving trajectory follows the planned path.
[0311] Braking and power systems: control the vehicle's longitudinal travel parameters.
[0312] It should be noted that a low-computing-power anthropomorphic autonomous driving decision-making and planning method and system based on driver profiles, such as... Figure 12 and Figure 13 As shown:
[0313] as follows Figure 12 As shown, Figure 12 The diagram below illustrates dynamic programming after sampling. The horizontal axis represents time with a sampling period of 1 second, and the vertical axis represents the longitudinal distance from the current position of the vehicle. Curve 1202 is the predicted distance and time curve of the vehicle's adjacent vehicles, curve 1204 is the predicted distance and time curve of another adjacent vehicle, and curves 1206 and 1208 are the distance and time relationship curves of the vehicle after dynamic programming. The points on curves 1206 and 1208 are the sampling points selected during dynamic programming.
[0314] like Figure 13 The diagram illustrates the selection of the target's appearance time and the next moment after its end (next moment = end time + sampling time) as sampling points. The horizontal axis represents time, with a sampling period of 1 second. The vertical axis represents the longitudinal distance from the vehicle's current position. Curve 1302 shows the predicted distance and time for adjacent vehicles around the vehicle, curve 1304 shows the predicted distance and time for another adjacent vehicle around the vehicle, and curves 1306 and 1308 show the distance-time relationship curves after dynamic programming. The points on curves 1306 and 1308 are the sampling points selected during dynamic programming. Using dynamic programming, the system dynamically traverses backward from the current position. The DP process is shown in curves 1306 and 1308, selecting the curve with the lowest cost that satisfies the constraints as the optimal speed curve. Therefore, the method of using the target's appearance time and the next moment after its end as Node nodes significantly reduces the computational power of dynamic programming and improves operational efficiency.
[0315] It should be noted that the moment of target appearance and the next moment after target exit are defined as follows: the moment of target appearance can be understood as the specific moment when the target vehicle enters the vehicle's perception range (such as the initial time point when the target is detected by sensors), and when its driving status and location information are first collected and recorded by the system. This is the starting time anchor point for the system to begin predicting the behavior and analyzing the trajectory of the target. The next moment after target exit can be understood as the moment immediately following the last moment when the target vehicle leaves the vehicle's perception range (such as driving out of the detection area) or when its driving status and location information are no longer collected and recorded by the system. This is the time node when the system terminates continuous monitoring and analysis of the target, stops updating its behavior prediction results and trajectory planning, and is also the cutoff reference point for subsequent data aggregation and processing related to the target (such as driving feature statistics and interaction event review).
[0316] The target occurrence time and end time can also be the time following the start time of the interaction and the time following the end time of the interaction. The target occurrence time is the moment when the vehicle and the adjacent vehicle just begin to interact. To further illustrate, suppose vehicle A is in the right lane of the vehicle. When vehicle A begins to change lanes into the vehicle's lane, the moment when vehicle A turns to the left and any part of vehicle A begins to overlap with the vehicle's lane can be determined as the target occurrence time. Similarly, when vehicle A and the vehicle's lanes no longer overlap, if vehicle A changes lanes to the left again, and when vehicle A changes lanes to the left lane of the vehicle, and after completing the lane change, there is no longer any position of vehicle A overlapping with the vehicle's lane, this moment can be considered the target end time.
[0317] From the perspective of the spatiotemporal channel, the moment a target appears is the first time that the spatiotemporal information (location and status data) of the target vehicle is accessed by the system's spatiotemporal channel. At this moment, the target's "time coordinates" and "spatial coordinates" are effectively mapped for the first time and recorded by the channel. The spatiotemporal channel begins to continuously receive and transmit the target's dynamic spatiotemporal data, building a continuous spatiotemporal data link for subsequent behavior prediction and trajectory analysis. The moment after the target ends is the time when the target vehicle's spatiotemporal information officially exits the system's spatiotemporal channel. The last set of spatiotemporal data of the target has been transmitted and recorded. After this moment, the spatiotemporal channel no longer receives new data from the target, and the original target spatiotemporal link terminates. At the same time, this moment serves as the cutoff anchor point for the target data in the spatiotemporal channel, providing a clear boundary marker for the segmented processing of subsequent spatiotemporal data (such as tracing interactive events within a certain spatiotemporal interval).
[0318] Table 3 shows an example of model training and validation for a low-computing-power anthropomorphic autonomous driving decision-making and planning method and system based on driver profiles, with the variables for model training as follows:
[0319] Table 3
[0320]
[0321] In this embodiment, the training dataset is generated by a MATLAB script according to certain rules, and the selection of window A or B is done manually based on distance and speed. F1 represents the first vehicle in front of the current vehicle, and R1, R2, and R3 represent vehicles to the left or right of the current vehicle.
[0322] Dataset constraints: The lateral velocity of a vehicle is (-1, 1) m / s; the lateral acceleration of a vehicle is (-0.5, 0.5) m / s²; the distance between a vehicle and the lane line is (-2, 2) m; the longitudinal velocity of a vehicle is (11.1, 16.7) m / s; the longitudinal acceleration of a vehicle is (-3, 3) m / s²; the difference in longitudinal velocity between vehicles in the same lane is (-2, 2) m / s; the difference in longitudinal acceleration between vehicles in the same lane is (-1, 1) m / s; the interval between vehicles is related to the vehicle's velocity and acceleration.
[0323] Label settings: The label defaults to 1 (B window). If F1=0 (does not exist), and if the length of LF is greater than or equal to 9, and the speed of R1 is greater than or equal to the speed of R1 - 1;
[0324] If F1=1 (exists), if LA>=LB and LF>=offset+6, if LA>=9 and LF>=offset+6, if the speed of R1>=R1 and LF>=offset+6, then the label is set to 0 (Window A).
[0325] Save the data generated by Matlab, use a Python script to read the .mat file, distinguish the feature vectors and labels, and divide the training set and test set into .h5 types with key values X_train, Y_train, X_test, and Y_test, respectively.
[0326] Save the data generated by Matlab, use a Python script to read the .mat file, distinguish the feature vectors and labels, and divide the training set and test set into .h5 types with key values X_train, Y_train, X_test, and Y_test, respectively.
[0327] The decision model consists of an input layer, seven fully connected layers, seven Dropout layers, and an output layer.
[0328] The input to the decision model is (None, 73, 1), and the output is (None, 2, 1).
[0329] Model Training: Model parameter initialization: kernel_initializer = 'he_normal'; Activation function: activation='relu'; Regularization: kernel_regularizer = regularizers.l2(0.02); Learning rate: lr= 0.0001; Loss function: loss='categorical_crossentropy'; Optimizer: optimizer=RMSprop; Evaluation metrics: metrics=['accuracy']; Number of iterations: epochs=100; Batch size per training session: batch_size=128.
[0330] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0331] According to another aspect of the embodiments of this application, an autonomous driving behavior decision-making device for implementing the above-described autonomous driving behavior decision-making method is also provided. For example... Figure 14 As shown, the device includes:
[0332] The first determining unit 1402 is used to determine a set of driving state parameters that match the current vehicle and the reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle.
[0333] The acquisition unit 1404 is used to acquire the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within the historical driving cycle;
[0334] Input unit 1406 is used to input the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probability that matches at least two driving behaviors.
[0335] The second determining unit 1408 is used to determine the target driving behavior from at least two driving behaviors based on at least two operational probabilities.
[0336] As an optional solution, the input unit 1406 includes: a first determining module, used to determine a driving feature vector that matches the driving state parameter set and the driving style label; an input module, used to input the driving feature vector into a behavior planning model, wherein the behavior planning model is a neural network model containing an input layer, at least one hidden layer and an output layer; and a first acquiring module, used to acquire the operation probabilities corresponding to at least two output nodes in the output layer of the behavior planning model.
[0337] As an optional solution, the second determining unit 1408 includes: a second acquisition module for acquiring the actual driving behavior of the current vehicle; a second determining module for determining the current predicted loss based on the actual driving behavior and at least two operation probabilities; and an adjustment module for adjusting at least one network parameter in the behavior planning model based on the current predicted loss.
[0338] As an optional approach, input unit 1406 further includes: a third acquisition module, used to acquire an incremental sample as the current incremental sample when the number of incremental samples included in the incremental sample set is greater than or equal to the target number threshold, wherein the incremental sample is determined based on historical driving behavior that meets the loss condition, and the historical driving behavior does not match the reference driving behavior predicted by the behavior planning model at a historical moment; an update module, used to update the behavior planning model by gradient descent when the cross-entropy loss matched by the current incremental sample does not meet the convergence condition; and a repetition module, used to repeat the steps until the cross-entropy loss meets the convergence condition.
[0339] As an optional solution, the second determining unit 1408 includes: a first triggering module for triggering a first control signal matching a first driving behavior among at least two driving behaviors, wherein the first control signal is used to control the current vehicle to accelerate in the current lane; a second triggering module for triggering a second control signal matching a second driving behavior among at least two driving behaviors, wherein the second control signal is used to control the current vehicle to decelerate in the current lane; a third triggering module for triggering a third control signal matching a third driving behavior among at least two driving behaviors, wherein the third control signal is used to control the current vehicle to accelerate into a reference lane; and a fourth triggering module for triggering a fourth control signal matching a fourth driving behavior among at least two driving behaviors, wherein the fourth control signal is used to control the current vehicle to decelerate before entering the reference lane.
[0340] As an optional solution, the acquisition unit 1404 includes: a fourth acquisition module for acquiring the driving style label sent by the reference vehicle; a fifth acquisition module for acquiring the driving description information set of the reference vehicle in the first driving cycle, wherein the driving description information set is used to indicate multiple driving decision behaviors of the reference vehicle in the first driving cycle; and inputting the driving description information set into the driving style classification model to obtain the driving style label.
[0341] As an optional solution, the device further includes: a third determining unit, configured to determine a first position node and a second position node based on the trajectory position relationship between the first reference trajectory corresponding to the current vehicle in the spatiotemporal map and the second reference trajectory corresponding to the reference vehicle in the spatiotemporal map, wherein the spatiotemporal map is used to indicate the change of the object position of at least one vehicle object over time, and the at least one vehicle object includes the current vehicle and the reference vehicle; a fourth determining unit, configured to use the first position node as the trajectory starting point of the estimated driving trajectory of the current vehicle and the second position node as the trajectory ending point of the estimated driving trajectory; and a fifth determining unit, configured to determine at least one intermediate trajectory point in the estimated driving trajectory based on the driving style label and the target driving behavior.
[0342] As an optional solution, the fifth determining unit includes: a third determining module for determining a first collision risk weight based on a driving style label; a fourth determining module for determining a second collision risk weight based on the target driving behavior; and a fifth determining module for sequentially determining at least one intermediate trajectory point in the estimated driving trajectory based on the first collision risk weight and the second collision risk weight.
[0343] Optionally, in this embodiment, the implementation of each of the above-mentioned unit modules can be referred to the above-mentioned method embodiments, which will not be repeated here.
[0344] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described autonomous driving behavior decision-making method is also provided. This electronic device may be... Figure 15 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 15 As shown, the electronic device includes a memory 1502 and a processor 1504. The memory 1502 stores a computer program, and the processor 1504 is configured to execute the steps of any of the above method embodiments via the computer program.
[0345] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0346] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0347] S1, determine a set of driving state parameters that match the current vehicle and the reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle;
[0348] S2, obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within the historical driving cycle;
[0349] S3, input the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probability that matches at least two driving behaviors;
[0350] S4, determine the target driving behavior from at least two driving behaviors based on at least two operational probabilities.
[0351] Alternatively, as those skilled in the art will understand, Figure 15 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 15 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 15 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 15 The different configurations shown.
[0352] The memory 1502 can be used to store software programs and modules, such as the program instructions / modules corresponding to the autonomous driving behavior decision-making method and device in this embodiment. The processor 1504 executes various functional applications and data processing by running the software programs and modules stored in the memory 1502, thereby realizing the aforementioned autonomous driving behavior decision-making method. The memory 1502 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1502 may further include memory remotely located relative to the processor 1504, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1502 may be used, but is not limited to, to store information such as target driving behavior. As an example, such as Figure 15As shown, the memory 1502 may include, but is not limited to, the first determining unit 1402, the acquiring unit 1404, the input unit 1406, and the second determining unit 1408 of the autonomous driving behavior decision-making device. Furthermore, it may include, but is not limited to, other module units of the autonomous driving behavior decision-making device, which will not be elaborated upon in this example.
[0353] Optionally, the aforementioned transmission device 1506 is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1506 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one example, the transmission device 1506 is a radio frequency (RF) module used to communicate with the Internet wirelessly.
[0354] In addition, the aforementioned electronic device also includes: a display 1508 for displaying target driving behavior; and a connection bus 1510 for connecting the various module components in the aforementioned electronic device.
[0355] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0356] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0357] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0358] According to one aspect of this application, a computer-readable storage medium is provided, from which a processor of a computer device reads computer instructions, and the processor executes the computer instructions, causing the computer device to perform the aforementioned autonomous driving behavior decision-making method.
[0359] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0360] S1, determine a set of driving state parameters that match the current vehicle and the reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle;
[0361] S2, obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within the historical driving cycle;
[0362] S3, input the set of driving state parameters and driving style label data into the behavior planning model to obtain the operation probability that matches at least two driving behaviors;
[0363] S4, determine the target driving behavior from at least two driving behaviors based on at least two operational probabilities.
[0364] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0365] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0366] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0367] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0368] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0369] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0370] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A behavior decision-making method for autonomous driving, characterized in that, include: Determine a set of driving state parameters that match the current vehicle and the reference vehicle, wherein the set of driving state parameters is used to indicate the relative physical state between the current vehicle and the reference vehicle; Obtain the driving style label of the reference vehicle, wherein the driving style label is determined based on the historical driving data of the reference vehicle within a historical driving cycle; input the driving state parameter set and the driving style label data into the behavior planning model to obtain the operation probability that matches at least two driving behaviors respectively; The target driving behavior is determined from at least two driving behaviors based on at least two of the said operational probabilities; Based on the trajectory position relationship between the current vehicle and the second reference trajectory of the reference vehicle in the spatiotemporal map, a first position node and a second position node are determined. The spatiotemporal map is used to indicate the change of the object position of at least one vehicle object over time. The at least one vehicle object includes the current vehicle and the reference vehicle. The first location node is used as the starting point of the estimated driving trajectory of the current vehicle, and the second location node is used as the ending point of the estimated driving trajectory. The first collision risk weight is determined based on the driving style label; The second collision risk weight is determined based on the target driving behavior; Based on the first collision risk weight and the second collision risk weight, at least one intermediate trajectory point in the estimated driving trajectory is determined sequentially.
2. The method according to claim 1, characterized in that, The step of inputting the driving state parameter set and the driving style label data into the behavior planning model to obtain the operation probability matching at least two driving behaviors includes: Determine the driving feature vector that matches the set of driving state parameters and the driving style label; The driving feature vector is input into the behavior planning model, wherein the behavior planning model is a neural network model containing an input layer, at least one hidden layer and an output layer; Obtain the operation probabilities corresponding to at least two output nodes in the output layer of the behavior planning model.
3. The method according to claim 2, characterized in that, After determining the target driving behavior from at least two driving behaviors based on at least two of the said operational probabilities, the method further includes: Obtain the actual driving behavior of the current vehicle; The current predicted loss is determined based on the actual driving distance and at least two of the operational probabilities. At least one network parameter in the behavior planning model is adjusted based on the current predicted loss.
4. The method according to claim 2, characterized in that, Before inputting the driving state parameter set and the driving style label data into the behavior planning model to obtain the operation probability matching at least two driving behaviors, the method further includes: If the number of incremental samples included in the incremental sample set is greater than or equal to the target number threshold, an incremental sample is obtained from the incremental sample set as the current incremental sample. The incremental sample is determined based on the historical driving behavior that meets the loss condition. The historical driving behavior does not match the reference driving behavior predicted by the behavior planning model at the historical moment. If the cross-entropy loss of the current incremental sample matching does not meet the convergence condition, the behavior planning model is updated by gradient descent. Repeat the above steps until the cross-entropy loss satisfies the convergence condition.
5. The method according to claim 1, characterized in that, After determining the target driving behavior from at least two driving behaviors based on at least two of the operational probabilities, the method further includes one of the following: Trigger a first control signal that matches a first driving behavior in at least two of the driving behaviors, wherein the first control signal is used to control the current vehicle to accelerate in the current lane; A second control signal is triggered that matches the second driving behavior in at least two of the driving behaviors, wherein the second control signal is used to control the current vehicle to decelerate in the current lane; Trigger a third control signal that matches a third driving behavior among at least two of the said driving behaviors, wherein the third control signal is used to control the current vehicle to accelerate into the reference lane; A fourth control signal is triggered that matches the fourth driving behavior among at least two of the said driving behaviors, wherein the fourth control signal is used to control the current vehicle to decelerate before entering the reference lane.
6. The method according to claim 1, characterized in that, The acquisition of the driving style label of the reference vehicle includes one of the following: Obtain the driving style tag sent by the reference vehicle; Obtain a set of driving description information of the reference vehicle during a first driving cycle, wherein the set of driving description information is used to indicate multiple driving decision behaviors of the reference vehicle during the first driving cycle; input the set of driving description information into a driving style classification model to obtain the driving style label.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method according to any one of claims 1 to 6.
8. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 6 through the computer program.
Citation Information
Patent Citations
Decision-making method for forced lane change of unmanned vehicle, computer equipment and medium
CN116767218A
Intelligent driving decision-making method and device, storage medium and vehicle
CN117068204A