Driving decision determination method, device and electronic equipment
By combining special models and large language models to evaluate the probability density of driving behavior, the problem of insufficient accuracy of autonomous driving decision planning is solved, and more efficient and safe driving decisions are achieved.
Patent Information
- Application Number
- CN202510542429.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing autonomous driving decision planning strategies have insufficient human-like and generalization capabilities, resulting in low decision-making accuracy, especially in a mixed-car environment, which is difficult to meet safe and efficient driving needs.
Combining the preset driving decision-making special model and the preset driving decision-making large language model, by obtaining the current driving scenario data, using the specific optimization capabilities of the dedicated model and the language understanding and knowledge reasoning capabilities of the large language model, comprehensively assess the probability density of driving behavior and generate target driving decisions.
It improves the accuracy and reliability of autonomous driving decisions, can better deal with various complex and specific driving scenarios, and enhances the safety and efficiency of the vehicle.
Smart Images

Figure CN120057041B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent driving technology, and in particular to a driving decision-making method, device, and electronic equipment. Background Art
[0002] From a technical perspective, the safe and efficient operation of autonomous vehicles on the road requires the coordinated work of three modules: perception, decision-making and planning, and control execution. Decision-making and planning, considered the intelligent brain of an autonomous vehicle, coordinates the overall situation, planning short-term avoidance maneuvers and long-term routes based on road conditions, traffic regulations, and vehicle status. It is crucial for the safe and efficient operation of autonomous vehicles.
[0003] Lane changing is often considered one of the most unpredictable driving tasks and a major contributor to traffic congestion and vehicle collisions. Therefore, research on lane-changing decision-making and planning strategies is highly desirable, especially as we transition into an era of mixed autonomous and human-driven driving.
[0004] In mixed traffic environments, a key challenge for autonomous vehicles to maintain social compatibility is the lack of human-likeness and generalizability of their decision-making and planning strategies. Whether these strategies are human-like and generalizable is crucial to a vehicle's ability to accurately complete various driving tasks, significantly impacting the operational safety and efficiency of transportation systems. However, current autonomous driving decision-making and planning research is typically rule-based or end-to-end modeling, with strong specialized reasoning capabilities and insufficient common-sense reasoning.
[0005] Therefore, how to improve the accuracy of autonomous driving decision-making and planning has become an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a driving decision determination method, device and electronic device to solve the problem of how to improve the accuracy of autonomous driving decision planning.
[0007] In a first aspect, the present invention provides a driving decision determination method, the method comprising:
[0008] Obtaining current driving scene data corresponding to the vehicle, the current driving scene data including at least one of video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data;
[0009] Inputting current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model;
[0010] Based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model, the target driving decision corresponding to the vehicle is determined.
[0011] The driving decision determination method provided in an embodiment of the present application obtains current driving scene data corresponding to a vehicle. The current driving scene data includes at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data. By collecting the current driving scene data, the vehicle can fully perceive its driving environment, providing a rich and accurate information basis for subsequent decision-making. The current driving scene data is input into a preset dedicated driving decision model and a preset large language model for driving decisions. This allows the respective advantages of these two models to be fully utilized. Dedicated models are typically optimized for specific driving scenarios and tasks, excelling at handling common, patterned driving situations and providing quick and accurate decision recommendations. The large language model, on the other hand, has powerful language understanding and knowledge reasoning capabilities, capable of handling complex semantic information and rare driving scenarios, and better handling situations requiring comprehensive judgment and knowledge application. The combined use of these two models enables the vehicle to make more flexible decisions in various driving scenarios. Based on the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model for driving decisions, the target driving decision corresponding to the vehicle is determined. This fusion decision-making approach combines the results of both models, avoiding the potential limitations of a single model. By analyzing and fusing the two probability densities, the likelihood of various driving behaviors can be more accurately assessed, leading to targeted driving decisions that better align with actual driving scenarios and improve decision accuracy and reliability.
[0012] In an optional embodiment, the training process of the preset driving decision-making large language model includes:
[0013] Get the original large language model;
[0014] Decompose the weight matrix of the original large language model into at least two low-rank matrices;
[0015] Freeze the original parameters corresponding to the original large language model;
[0016] Obtain a preset driving decision expert database; the preset driving decision expert database includes expert experience, industry rules, and historical cases;
[0017] Based on the preset driving decision expert database, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix;
[0018] Fuse the original parameters with the updated parameters to obtain the target parameters;
[0019] Based on the target parameters, a preset driving decision language model is obtained.
[0020] The driving decision determination method provided in the embodiments of the present application obtains an original large language model, thereby leveraging the existing strong foundation of the original large language model. This avoids the large computational resources and time required to train the model from scratch. The weight matrix of the original large language model is decomposed into at least two low-rank matrices, thereby reducing the complexity of the original large language model and facilitating model optimization and understanding. The original parameters corresponding to the original large language model are then frozen, preserving the general language knowledge and capabilities learned by the original large language model. This pre-trained knowledge is very helpful for understanding and processing driving-related natural language information. Furthermore, freezing the original parameters helps stabilize the model training process. A pre-set driving decision expert database is obtained to incorporate professional knowledge and experience to meet actual driving needs. Based on the pre-set driving decision expert database, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix, thereby customizing the model according to the specific needs of driving decisions. The information in the expert database can guide the pre-set driving decision large language model to learn specific patterns and laws related to driving, allowing the pre-set driving decision large language model to gradually adapt to various factors in the driving scenario, such as traffic flow and changing road conditions. In this way, the pre-set driving decision large language model can better capture the relationship between driving decisions and various factors, thereby generating decisions that are more consistent with actual driving situations. The original parameters are fused with the updated parameters to obtain the target parameters, leveraging the strengths of both. The original parameters retain the general language understanding capabilities and extensive knowledge base of the large language model, while the updated parameters imbue the model with specific capabilities and expertise for driving decision-making tasks. By fusing these two parameters, the model can leverage its general language processing capabilities to understand and process a variety of driving-related natural language information while also making accurate driving decisions based on the knowledge in the expert database, achieving an organic combination of general and specialized capabilities. Based on the target parameters, the pre-set driving decision large language model is obtained. This makes the pre-set driving decision large language model a model specifically optimized for driving decision-making tasks. Combining the strong foundation of the original large language model with the specialized driving decision-making capabilities gained through parameter updates from the expert database, it can efficiently process a variety of driving-related information and generate accurate and reasonable driving decisions. This model can better meet the real-time, accuracy, and reliability requirements of autonomous or assisted driving systems, providing strong support for safe vehicle operation.
[0021] In an optional embodiment, based on a preset driving decision expert database, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix, including:
[0022] Identify expert experience, industry rules, and historical cases in the pre-set driving decision expert database to build a knowledge graph;
[0023] Input the knowledge graph into the original large language model;
[0024] The original large language model recognizes the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments;
[0025] Reward evaluation is performed on each simulated driving decision based on a preset reward function to obtain reward feedback data; the preset reward function includes safety reward, efficiency reward, comfort reward, and social compatibility reward;
[0026] Based on the reward feedback data, the direction of updating the parameters of each low-rank matrix is determined;
[0027] The parameters of each low-rank matrix are updated based on the update direction to obtain the updated parameters corresponding to each low-rank matrix.
[0028] The driving decision determination method provided in the embodiments of the present application identifies expert experience, industry rules, and historical cases in a preset driving decision expert database to construct a knowledge graph. This structured knowledge representation facilitates the understanding and utilization of subsequent models, improving the accessibility and operability of knowledge. The knowledge graph is input into the original large language model, thereby enriching the original large language model's knowledge base and enabling it to better understand driving-related concepts, rules, and experience. For example, the original large language model can use the knowledge graph to understand the meaning of different traffic signs and driving strategies under various road conditions, thereby enabling it to reason and make decisions based on more comprehensive knowledge when handling driving-related tasks. The original large language model identifies the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments. This allows the original large language model to rehearse and analyze a variety of possible driving situations in a virtual environment, covering everything from simple everyday driving scenarios to complex special situations such as inclement weather and traffic accident scenes. Through this simulation, the model can learn and grasp the optimal driving decisions in different scenarios in advance, improving its ability to cope with actual driving scenarios. Rewards are evaluated for each simulated driving decision based on a preset reward function to obtain reward feedback data. A preset reward function evaluates simulated driving decisions across multiple dimensions, including safety, efficiency, comfort, and social compatibility, providing a comprehensive assessment of decision quality. This multi-dimensional evaluation approach avoids focusing solely on a single metric (such as safety) while ignoring other important aspects (such as comfort and social compatibility). Reward feedback data determines the direction in which the parameters of each low-rank matrix are updated, making parameter updates targeted and precise. Based on reward feedback, the model can prioritize adjustments to parameters that have a significant impact on achieving high rewards, avoiding blind parameter updates that can degrade model performance. The parameters of each low-rank matrix are updated based on the update direction, resulting in updated parameters for each low-rank matrix. These updated parameters enable the original large speech model to better handle driving decision-making tasks, improving decision accuracy and reliability. For example, the updated model may be able to more accurately judge the driving intentions of other vehicles in complex road conditions, leading to safer and more effective driving decisions.
[0029] In an optional embodiment, the preset reward function formula is: ;
[0030] Where N is the number of simulations, Reward for safety, Reward for efficiency, Reward for comfort, rewards for social compatibility;
[0031] ; Where C is the number of collisions, α is the safety penalty coefficient, and α>0;
[0032] ; Among them, T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency bonus coefficient, and β>0;
[0033] ; The smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration rate of change, which is recorded as ; n is the number of time intervals during driving, Δa i is the acceleration change in the i-th time interval, and the reasonable following distance is maintained by the actual following distance d actual and ideal following distance d ideal The deviation is measured as , m is the number of times of following the car; γ is the comfort reward coefficient, S a and S d are the normalized constants of the sum of squares of acceleration change rate and following distance deviation, respectively, and γ>0; when the acceleration changes smoothly and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative;
[0034] ; Among them, the degree of traffic flow can be measured by the average speed change rate of surrounding vehicles, which is recorded as , when simulated driving decisions promote smooth traffic flow, , let the number of lane changes that are reasonable and completed successfully be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ >0, when the simulated driving decision helps to improve the average speed of surrounding vehicles and the lane change is completed reasonably, the social compatibility reward is positive.
[0035] The driving decision determination method provided in this embodiment uses a preset reward function to evaluate simulated driving decisions across multiple dimensions, including safety, efficiency, comfort, and social compatibility. This allows for a comprehensive assessment of the quality of these decisions. This multi-dimensional evaluation approach avoids focusing on a single metric (e.g., safety) while ignoring other important aspects (e.g., comfort and social compatibility).
[0036] In an optional embodiment, determining a target driving decision corresponding to the vehicle based on a first driving behavior probability density output by a preset driving decision-specific model and a second driving behavior probability density output by a preset driving decision large language model includes:
[0037] Input the current driving scene data into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle;
[0038] Inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into a preset weight determination model, outputting weight information corresponding to a preset driving decision-specific model and a preset driving decision-making large language model, respectively;
[0039] Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are fused to determine the target driving decision corresponding to the vehicle.
[0040] The driving decision determination method provided in an embodiment of the present application inputs current driving scene data into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle. This allows the complexity of the driving scene to be accurately assessed based on the input current driving scene data. A first driving behavior probability density, a second driving behavior probability density, and the current driving scene complexity are input into a preset weight determination model, which outputs weight information corresponding to a preset driving decision-specific model and a preset driving decision-based large language model, respectively. This flexible adjustment of model weights based on scene complexity optimizes the entire driving decision-making process and fully leverages the advantages of both models. This avoids the rigid use of a single model or a fixed weight distribution scheme for all scenarios, thereby improving decision accuracy and adaptability. Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are fused to determine the target driving decision corresponding to the vehicle. By fusing the driving behavior probability densities output by the two models using the weight information, the specific domain expertise of the preset driving decision-specific model and the general understanding and reasoning capabilities of the preset driving decision-based large language model are fully utilized. Furthermore, this fusion process reduces the errors and uncertainties that may arise from a single model, improving the reliability and stability of driving decisions.
[0041] In an optional embodiment, the current driving scene data is input into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle, including:
[0042] Perform feature extraction on the video data in front of the vehicle in the current driving scene data to extract visual features of multiple objects; the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles;
[0043] Identify the vehicle's operating information, surrounding vehicle motion information, and surrounding road data in the current driving scene data, and construct an initial graph structure using vehicles, pedestrians, road facilities, and obstacles as nodes and the spatial relationships and interactions between them as edges.
[0044] Identify the front video data of the vehicle in the current driving scene data to determine the current driving scene type;
[0045] Based on the current driving scenario type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure.
[0046] Fuse visual features, target graph structure, environmental data, and time data to generate target features;
[0047] Based on the target features, the complexity of the current driving scene corresponding to the vehicle is determined.
[0048] The driving decision determination method provided in the embodiments of the present application performs feature extraction on the video data in front of the vehicle in the current driving scene data, extracting visual features of various objects. By extracting visual features of various objects, such as vehicles, pedestrians, road facilities, and obstacles, the system can accurately perceive various objects in the environment in front of the vehicle. The method identifies the vehicle's operating information, surrounding vehicle motion information, and surrounding road data in the current driving scene data. Using vehicles, pedestrians, road facilities, and obstacles as nodes and the spatial relationships and interactions between them as edges, it constructs an initial graph structure that clearly presents the relationships between various elements in the current driving scene. This graph structure facilitates the analysis of complex driving situations. The system can use the graph structure to quickly analyze the interactions between multiple objects. The method identifies the video data in front of the vehicle in the current driving scene data and determines the current driving scene type. Different driving scenarios have different characteristics and requirements. Once the scene type is determined, targeted handling of various situations can be achieved, thereby improving the accuracy and effectiveness of decision making. Based on the current driving scene type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure. The information contained in the graph structure is further enriched. The target graph structure, with its semantic information, provides a more optimized basis for system decision-making. Visual features, the target graph structure, environmental data, and temporal data are fused to generate target features. These fused target features possess enhanced representational capabilities and can more accurately describe the complexity and characteristics of the current driving scene. They can capture information that cannot be captured by a single data source, providing a more reliable basis for subsequent decision-making. Based on the target features, the complexity of the current driving scene corresponding to the vehicle is determined, thereby accurately assessing the complexity of the current driving scene.
[0049] In an optional embodiment, the visual features, target graph structure, environmental data, and time data are fused to generate target features, including:
[0050] Perform feature encoding on environmental data and time data to generate environmental data features and time data features;
[0051] Input visual features, target graph structure, environmental data features, and temporal data features into a preset cross-modal attention interaction network;
[0052] A preset cross-modal attention interaction network calculates the attention scores between visual features, target graph structure, environmental data features, and temporal data features;
[0053] Based on the attention scores among visual features, target graph structure, environmental data features and temporal data features, the visual features, target graph structure, environmental data features and temporal data features are fused to generate target features.
[0054] The driving decision determination method provided in the embodiment of the present application performs feature encoding on environmental data and time data to generate environmental data features and time data features, thereby converting these data of different types and formats into a unified feature representation that can be processed by the model. Visual features, target graph structure, environmental data features, and time data features are input into a preset cross-modal attention interaction network to achieve the integration of multi-source information. The preset cross-modal attention interaction network calculates the attention scores between the visual features, target graph structure, environmental data features, and time data features, enabling automatic focus on the most important information for understanding and making decisions in the current driving scene. Based on the attention scores between the visual features, target graph structure, environmental data features, and time data features, the visual features, target graph structure, environmental data features, and time data features are fused to generate target features. The generated target features can more comprehensively and accurately describe the driving scene. The fused target features combine the information advantages of different modalities, including the appearance and relationship information of objects, and taking into account the influence of environmental and time factors. They provide high-quality feature representations for subsequent determination of driving scene complexity and making driving decisions, helping to improve the accuracy and reliability of decisions.
[0055] In an optional embodiment, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene are input into a preset weight determination model, and weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model is output, including:
[0056] Calculating a difference measure between the first driving behavior probability density and the second driving behavior probability density;
[0057] Splicing the difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene to generate a splicing feature;
[0058] Obtaining the current driving decision output accuracy corresponding to the current driving scenario type, the preset driving decision-specific model, and the preset driving decision large language model;
[0059] Generate dynamic mapping parameters based on the current driving scenario type and the current driving decision output accuracy corresponding to the preset driving decision-specific model and the preset driving decision-making large language model respectively;
[0060] Mapping the splicing features based on dynamic mapping parameters to generate low-dimensional features;
[0061] Input the low-dimensional features into the preset hidden layer and perform feature processing on the low-dimensional features;
[0062] The output layer in the preset weight determination model outputs weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model respectively.
[0063] The driving decision determination method provided in an embodiment of the present application calculates a difference measure between a first driving behavior probability density and a second driving behavior probability density. By calculating the difference measure, the degree of difference between the first driving behavior probability density output by a preset driving decision-specific model and the second driving behavior probability density output by a preset driving decision-making large language model can be clearly understood. This helps to identify differences between the two models in their understanding of the current driving scenario and decision-making tendencies, providing a basis for subsequent comprehensive consideration of the outputs of the two models. The difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scenario are concatenated to generate a concatenated feature. This multi-dimensional information can more comprehensively describe the relevant factors of the current driving decision, providing a more comprehensive data foundation for subsequent model analysis. The accuracy of the current driving decision output corresponding to the current driving scenario type, the preset driving decision-specific model, and the preset driving decision-making large language model is obtained. Based on the current driving scenario type and the accuracy of the current driving decision output corresponding to the preset driving decision-specific model and the preset driving decision-making large language model, dynamic mapping parameters are generated. This enables dynamic adaptation to different driving scenarios. The performance of the models varies in different scenarios. By dynamically generating mapping parameters, the processing of the concatenated features can be optimized for each specific scenario. The concatenated features are mapped based on dynamic mapping parameters to generate low-dimensional features, effectively reducing the dimensional complexity of the data. The low-dimensional features are then fed into a preset hidden layer and processed to reveal deeper information and patterns within them. The output layer in the model, which uses preset weights, outputs the weights corresponding to the preset dedicated driving decision model and the preset large language model for driving decisions. This achieves a reasonable weight allocation between the two models for the current driving scenario, improving the quality of driving decisions.
[0064] In a second aspect, the present invention provides a driving decision determination device, the device comprising:
[0065] an acquisition module for acquiring current driving scene data corresponding to the vehicle, the current driving scene data including at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data;
[0066] An input module, configured to input current driving scene data into a preset driving decision-making dedicated model and a preset driving decision-making large language model;
[0067] The determination module is used to determine the target driving decision corresponding to the vehicle based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model.
[0068] The driving decision determination device provided in an embodiment of the present application obtains current driving scene data corresponding to the vehicle. The current driving scene data includes at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data. By collecting the current driving scene data, a rich and accurate information foundation is provided for subsequent decision-making. The current driving scene data is input into a preset dedicated driving decision model and a preset large language model for driving decisions. This allows the respective advantages of these two models to be fully utilized. Dedicated models are typically optimized for specific driving scenarios and tasks, excelling at handling common, patterned driving situations and providing quick and accurate decision recommendations. The large language model, on the other hand, possesses powerful language understanding and knowledge reasoning capabilities, capable of handling complex semantic information and rare driving scenarios, and better handling situations requiring comprehensive judgment and knowledge application. The combined use of these two models enables a more flexible decision-making approach for the vehicle in various driving scenarios. Based on the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model for driving decisions, the target driving decision corresponding to the vehicle is determined. This fusion decision-making approach combines the results of both models, avoiding the limitations that may exist with a single model. By analyzing and fusing the two probability densities, the likelihood of various driving behaviors can be evaluated more accurately, thereby arriving at target driving decisions that are more in line with actual driving scenarios and improving the accuracy and reliability of decisions.
[0069] In a third aspect, the present invention provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the driving decision determination method of the first aspect or any corresponding embodiment thereof by executing the computer instructions. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0071] Figure 1 is a flow chart of a driving decision determination method according to an embodiment of the present invention;
[0072] Figure 2 is a flow chart of another driving decision determination method according to an embodiment of the present invention;
[0073] Figure 3 is a timing framework diagram of a driving decision determination method according to an embodiment of the present invention;
[0074] Figure 4 is a structural block diagram of a driving decision determination device according to an embodiment of the present invention;
[0075] Figure 5 FIG. 4 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0076] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0077] It should be noted that the driving decision determination method provided in the embodiments of this application may be executed by a device that determines the driving decision. This device may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. The electronic device may be a control unit in the vehicle. The following method embodiments are all described using the electronic device as an example.
[0078] According to an embodiment of the present invention, an embodiment of a driving decision determination method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0079] In this embodiment, a driving decision determination method is provided, which can be used in the above-mentioned electronic device. Figure 1 is a flow chart of a driving decision determination method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0080] Step S101: Acquire current driving scene data corresponding to the vehicle.
[0081] The current driving scene data includes at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data.
[0082] Specifically, electronic devices can use a camera mounted on the front of the vehicle to collect video data from the vehicle's front. These devices then use the vehicle's own sensors, such as speed sensors, accelerometers, steering wheel angle sensors, and gyroscopes, to obtain corresponding information about the vehicle's movement. Electronic devices primarily rely on on-board radar (such as millimeter-wave radar and lidar) and connected vehicle technology to obtain information about surrounding vehicle movements. For example, millimeter-wave radar can measure information such as the distance, relative speed, and angle between surrounding vehicles and the vehicle itself, with high accuracy and real-time performance. Lidar can generate three-dimensional point cloud data of the surrounding environment, more accurately depicting the position and shape of surrounding vehicles. Connected vehicle technology, on the other hand, can obtain driving status information about surrounding vehicles, such as speed, direction, and braking status, through inter-vehicle communication (such as vehicle-to-vehicle communication).
[0083] Electronic devices can obtain surrounding road data through onboard mapping systems, road sensors (such as geomagnetic sensors and road profile sensors), and cameras. For example, the onboard mapping system provides basic road information, such as road grade, number of lanes, speed limit signs, and intersection information. Road sensors can monitor road conditions in real time, such as surface smoothness and the presence of obstacles. Cameras can use image recognition technology to identify road markings, lane markings, and other information.
[0084] Electronic devices can also collect information from environmental sensors installed on the vehicle, such as temperature, humidity, light, and rainfall sensors. Furthermore, they can exchange data with external environmental monitoring systems (such as weather stations) to obtain more comprehensive environmental information. Electronic devices primarily obtain time information from the vehicle's clock system and can also use the navigation system to obtain specific date and time.
[0085] Step S102 : inputting the current driving scene data into a preset driving decision-specific model and a preset driving decision-making language model.
[0086] Specifically, pre-built driving decision-making models are typically designed specifically for driving decision-making tasks. Their architecture may be based on traditional machine learning algorithms (such as decision trees and support vector machines) or deep learning models (such as convolutional neural networks and recurrent neural networks). Taking deep learning models as an example, pre-built driving decision-making models often take into account the characteristics of driving scene data. For video data from the ego vehicle's front, a convolutional neural network (CNN) is used to extract visual features from the image, such as those of objects like vehicles, pedestrians, and road signs. Through multiple layers of convolution and pooling operations, CNNs automatically learn feature representations at different levels of the image, from simple edge and texture information to complex object shape and structure. Numerical data, such as ego vehicle operation information and surrounding vehicle motion information, is processed through fully connected layers and integrated with visual features to fully understand the driving scene. After calculation and processing by the model, the pre-built driving decision-making model outputs a first driving behavior probability density. This probability density represents the probability distribution of various possible driving behaviors (such as acceleration, deceleration, lane change left, lane change right, and maintaining the current state) in the current driving scenario. For example, the output probability density may indicate that in the current scenario, the probability of acceleration is 0.1, the probability of deceleration is 0.3, the probability of changing lanes to the left is 0.1, the probability of changing lanes to the right is 0.1, the probability of maintaining the current state is 0.4, etc. The probability density output by the preset driving decision-making model is based on its learning and understanding of the input data and reflects the model's judgment and prediction of the current driving scenario.
[0087] The pre-configured large language model for driving decisions is an improvement and optimization based on large language models (such as the GPT series and BERT). In driving decision-making applications, the pre-configured large language model converts driving scenario data into language-based input. Numerical data requires conversion to appropriate text descriptions, such as describing a vehicle speed of "60 km / h" as "The current vehicle speed is 60 kilometers per hour." For image data, an image description generation algorithm may be used to convert key information from the image into textual statements. During this data conversion process, the accuracy and completeness of the information must be ensured so that the model can correctly understand the driving scenario. Furthermore, operations such as word segmentation and word embedding may be required on the input text to convert the text into a vector form that the model can process.
[0088] The pre-configured large language model for driving decisions outputs a second driving behavior probability density based on its understanding and reasoning of the input text. Similar to the pre-configured dedicated driving decision model, this probability density represents the probability distribution of various driving behaviors in the current driving scenario. With its powerful language understanding and reasoning capabilities, the large language model can process complex driving scenario descriptions and semantic information, outputting a more logical and reasonable driving behavior probability density. For example, when encountering a complex traffic sign or traffic rule description, the large language model can accurately determine the appropriate driving behavior based on its understanding of the language and output the corresponding probability density.
[0089] Step S103 : determining a target driving decision corresponding to the vehicle based on a first driving behavior probability density output by a preset driving decision-specific model and a second driving behavior probability density output by a preset driving decision-specific language model.
[0090] Specifically, the electronic device may determine weight information corresponding to a preset dedicated driving decision model and a preset large language model for driving decisions. It then performs a weighted calculation on the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model. Specifically, the weighted calculation may be a weighted sum of the probabilities of each driving behavior. For example, for acceleration, the fused probability = the probability of acceleration in the first driving behavior probability density × the dedicated model weight + the probability of acceleration in the second driving behavior probability density × the large language model weight.
[0091] After merging the new driving behavior probability density, the driving behavior with the highest probability is typically selected as the target driving decision. For example, if the probability of maintaining the current state in the fused probability density is 0.6, the highest probability among all driving behaviors, then the target driving decision is to maintain the current state. This method is simple and direct, enabling quick decision-making and is suitable for most common driving scenarios.
[0092] The driving decision determination method provided in the embodiments of the present application obtains current driving scene data corresponding to the vehicle. The current driving scene data includes at least one of the following: video data in front of the vehicle, vehicle operation information, surrounding vehicle motion information, surrounding road data, environmental data, and time data. By collecting the current driving scene data, the vehicle can fully perceive its driving environment, including the road conditions ahead, its own driving status, surrounding vehicle dynamics, road conditions, environmental factors, and time factors, providing a rich and accurate information foundation for subsequent decision-making. The current driving scene data is input into a preset dedicated driving decision model and a preset large language model for driving decision-making. This allows the respective advantages of these two models to be fully utilized. Dedicated models are typically optimized for specific driving scenarios and tasks, excelling at handling common, patterned driving situations and providing quick and accurate decision recommendations. The large language model, on the other hand, has powerful language understanding and knowledge reasoning capabilities, capable of handling complex semantic information and rare driving scenarios, and better handling situations requiring comprehensive judgment and knowledge application. The combined use of these two models enables the vehicle to make more flexible decisions in various driving scenarios. Whether in simple everyday driving scenarios or complex, special situations, the two models work together to generate appropriate driving decisions, improving the vehicle's ability to cope with a variety of complex environments. The target driving decision for the vehicle is determined based on the first driving behavior probability density output by the pre-set dedicated driving decision model and the second driving behavior probability density output by the pre-set large language model. This fusion decision-making approach integrates the results of both models, avoiding the limitations of a single model. By analyzing and fusing the two probability densities, the likelihood of various driving behaviors can be more accurately assessed, resulting in a target driving decision that better reflects the actual driving scenario, improving decision accuracy and reliability.
[0093] In this embodiment, a driving decision determination method is provided, which can be used in the above-mentioned electronic device. Figure 2 is a flow chart of a driving decision determination method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0094] Step S201: Acquire current driving scene data corresponding to the vehicle.
[0095] The current driving scene data includes at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data.
[0096] For details about this step, please refer to the above description of step S101 and will not be repeated here.
[0097] Step S202 : inputting the current driving scene data into a preset driving decision-specific model and a preset driving decision-making language model.
[0098] The training process of the preset driving decision-making language model may include the following steps:
[0099] Step a1: Obtain the original large language model.
[0100] Specifically, the electronic device may receive the original large language model input by the user, or may receive the original large language model sent by other devices. The embodiment of the present application does not specifically limit the manner in which the electronic device obtains the original large language model.
[0101] Step a2: decompose the weight matrix of the original large language model into at least two low-rank matrices.
[0102] Specifically, the electronic device decomposes the weight matrix of the original large language model into two low-rank matrices using a low-rank decomposition method.
[0103] Among them, the low-rank decomposition method can be any of the low-rank decomposition methods such as the singular value decomposition method and the non-negative matrix decomposition method.
[0104] Step a3: Freeze the original parameters corresponding to the original large language model.
[0105] Specifically, the electronic device can freeze the original parameters corresponding to the original large language model by setting the requires_grad attribute (in PyTorch) or the trainable attribute (in TensorFlow) of the parameter.
[0106] Step a4: Obtain a preset driving decision expert database.
[0107] Among them, the preset driving decision expert database includes expert experience, industry rules, and historical cases.
[0108] Specifically, the electronic device may receive a preset driving decision expert library input by a user, or may receive a preset driving decision expert library sent by other devices. The electronic device may also search for the preset driving decision expert library in a storage space.
[0109] Step a5: Based on the preset driving decision expert database, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix.
[0110] Specifically, the above step a5 may include the following steps:
[0111] Step a51: Identify the expert experience, industry rules, and historical cases in the preset driving decision expert database to construct a knowledge graph.
[0112] Specifically, the electronic device conducts an in-depth analysis of the textual knowledge in the preset driving decision expert database. Using natural language processing technology, it transforms textual content such as expert experience, industry rules, and historical cases into a structured knowledge representation, and identifies key entities from the knowledge representation. These entities include "driver," "vehicle," "road," "traffic signs," and "weather." For example, in the sentence "On a rainy day, the driver drove at 80 kilometers per hour on the highway," "rainy day" is the weather entity, "driver" is the person entity, "highway" is the road entity, and "80 kilometers per hour" is the speed entity.
[0113] The electronic device then determines the relationships between entities, such as "the driver operates the vehicle," "the vehicle is traveling on the road," and "traffic signs indicate the driver's behavior." Taking "the light ahead is red, and the driver steps on the brakes" as an example, the "indication" relationship between "red light (traffic sign)" and "driver" and the "execution" relationship between "driver" and "brake (operational behavior)" can be extracted. Next, the electronic device assigns corresponding attributes to the entities, such as the color, model, and speed of the vehicle, the age and driving experience of the driver, and the grade and width of the road. For example, for "a black car is traveling at 50 kilometers per hour on a city road," "black" can be extracted as the color attribute of the vehicle, "car" as the model attribute of the vehicle, "city road" as the type attribute of the road, and "50 kilometers per hour" as the speed attribute of the vehicle.
[0114] The electronic device then uses the extracted entities as nodes in the knowledge graph, with each node representing a specific concept or object. For example, "driver," "vehicle," "road," and "traffic sign" all exist as independent nodes in the knowledge graph. Edges are created based on the relationships between entities, with the direction and type of the edge representing the direction and nature of the relationship. For example, a "direction" edge is created from the "traffic sign" node to the "driver" node, indicating the direction relationship of the traffic sign to the driver; and an "operation" edge is created from the "driver" node to the "vehicle" node, indicating the operation relationship of the driver to the vehicle.
[0115] Finally, the electronic device adds the extracted attributes to the corresponding nodes to enrich the node information and generate a knowledge graph. For example, it adds attributes such as color, model, and speed to the "vehicle" node, and attributes such as age and driving experience to the "driver" node. In this way, each node in the knowledge graph has detailed description information, which can more accurately represent the knowledge related to driving decisions.
[0116] Step a52: input the knowledge graph into the original large language model.
[0117] Specifically, the electronic device converts the knowledge graph into a graph vector representation, and then inputs the graph vector representation into the original large language model.
[0118] In step a53, the original large language model recognizes the knowledge graph and outputs the corresponding simulated driving decisions under various simulated driving environments.
[0119] Specifically, the original large language model, such as a model based on the Transformer architecture, processes input data through a multi-layer self-attention mechanism and a feedforward neural network. When processing input related to the knowledge graph, the self-attention mechanism can capture the relationship between different parts of the input data, thereby understanding the semantic information in the knowledge graph. For example, the original large language model can use the self-attention mechanism to focus on the relationship between "traffic signs" and "driver behavior" in the knowledge graph, as well as their importance in the current simulated driving environment. The multi-layer structure of the original large language model enables it to extract features of the input data layer by layer, from simple word semantic understanding to complex knowledge reasoning.
[0120] The original large language model performs reasoning and logical judgment based on its understanding of the knowledge graph and other input information. It can infer reasonable driving decisions based on the expert experience, industry rules, historical cases, and other knowledge stored in the knowledge graph, combined with the characteristics of the current simulated driving environment. For example, if the knowledge graph records expert experience that speed should be reduced and a safe distance should be maintained when driving on rainy urban roads, the model will reason based on this knowledge when it identifies the current simulated driving environment as "driving on urban roads in the rain" and determine the appropriate driving behavior in this situation, such as recommending that the driver reduce speed to an appropriate range and maintain a certain safe distance from the vehicle in front.
[0121] Next, the original large language model generates simulated driving decisions based on the inference results. This may involve extracting relevant decision suggestions from the knowledge graph and adjusting and optimizing them based on the specific conditions of the current simulated driving environment. For example, under different vehicle speeds and road conditions, the original large language model will generate specific operating instructions based on the rules and experience in the knowledge graph, such as "Current speed 60 km / h, intersection ahead, please slow down to 30 km / h and watch for pedestrians." The original large language model may also consider the combined influence of multiple factors, such as weather, traffic flow, and vehicle status, to generate comprehensive driving decisions.
[0122] Step a54: Perform reward evaluation on each simulated driving decision based on a preset reward function to obtain reward feedback data.
[0123] The preset reward functions include safety reward, efficiency reward, comfort reward and social compatibility reward.
[0124] In an optional embodiment of the present application, the preset reward function formula is: ;
[0125] Where N is the number of simulations, Reward for safety, Reward for efficiency, Reward for comfort, rewards for social compatibility;
[0126] ; Where C is the number of collisions, α is the safety penalty coefficient, and α>0; for example, each time a collision occurs, α points are deducted. If no collision occurs, C=0, then R safety =0.
[0127] ; Among them, T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency reward coefficient, and β>0; for example, when the actual completion time is equal to the expected time, the efficiency reward is 0; when the actual completion time is less than the expected time, the reward is positive, otherwise it is negative.
[0128] ; The smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration rate of change, which is recorded as ; n is the number of time intervals during driving, Δa i is the acceleration change in the i-th time interval, and the reasonable following distance is maintained by the actual following distance d actual and ideal following distance d ideal The deviation is measured as , m is the number of times of following the car; γ is the comfort reward coefficient, S a and S d are the normalized constants of the sum of squares of acceleration change rate and following distance deviation, respectively, and γ>0; when the acceleration changes smoothly and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative;
[0129] ; Among them, the degree of traffic flow can be measured by the average speed change rate of surrounding vehicles, which is recorded as , when simulated driving decisions promote smooth traffic flow, , let the number of lane changes that are reasonable and completed successfully be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ>0, when the simulated driving decision helps to improve the average speed of surrounding vehicles and the lane change is completed reasonably, the social compatibility reward is positive.
[0130] Specifically, the electronic device performs a reward evaluation on each simulated driving decision based on a preset reward function to obtain reward feedback data.
[0131] In step a55, based on the reward feedback data, the direction of updating the parameters of each low-rank matrix is determined.
[0132] Specifically, to determine the relationship between reward feedback data and low-rank matrix parameters, the electronic device can establish an association model. Machine learning methods can be used to train a model using low-rank matrix parameters as input features and reward feedback data as output labels to learn the mapping relationship between them. For example, a neural network model can be used to input the parameter values of the low-rank matrix into the network, and after multiple layers of calculation and transformation, the predicted reward score is output. The model parameters are adjusted using a large amount of training data, allowing the model to predict the relationship between reward scores and low-rank matrix parameters as accurately as possible.
[0133] After establishing the association model, the feature importance of the low-rank matrix parameters is analyzed. Feature selection methods, such as gradient-based methods and information gain-based methods, can be used to determine which low-rank matrix parameters have a significant impact on the reward feedback data. For example, by calculating the gradient of each parameter, a larger gradient indicates a more significant impact on the reward score. This identifies key parameters with significant impact on the reward feedback data, providing a basis for determining the direction of parameter updates. The electronic device can then use a gradient-based optimization method to determine the direction of parameter updates based on the gradients calculated from the association model between the reward feedback data and the low-rank matrix parameters. Parameters that increase the reward score are increased (following the gradient's ascending direction); parameters that decrease the reward score are decreased (following the gradient's descending direction). For example, if the gradient of a low-rank matrix parameter is positive, indicating that increasing the parameter's value is likely to increase the reward score, the parameter's value is appropriately increased during the parameter update. By continuously updating the parameters along the gradient's direction, the low-rank matrix is gradually optimized to improve the reward score of the simulated driving decision.
[0134] Step a56: Update the parameters of each low-rank matrix based on the update direction to obtain the updated parameters corresponding to each low-rank matrix.
[0135] Specifically, the electronic device determines the direction and magnitude of the parameter update by calculating the gradient of the reward feedback data with respect to the parameters of the low-rank matrix. It then updates the parameters of the low-rank matrix, obtaining the updated parameters corresponding to each low-rank matrix. The learning rate η is a key factor in gradient-based update methods. If the learning rate is too large, the parameter update step size will be too large, potentially causing the parameters to skip the optimal value or even prevent the model from convergence. If the learning rate is too small, the parameter update will be very slow, and the training time will be very long. Therefore, the learning rate needs to be dynamically adjusted according to the training process. A common method is learning rate decay, which gradually reduces the learning rate as training progresses. For example, the initial learning rate is set to 0.1. After a certain number of training steps (e.g., 1000 steps), the learning rate is multiplied by a decay factor (e.g., 0.9), resulting in a learning rate of 0.1 × 0.9. This allows for rapid parameter adjustment in the early stages of training, and reduces the step size as the optimal value is approached, improving the accuracy of parameter updates.
[0136] Step a6: Fusing the original parameters with the updated parameters to obtain the target parameters.
[0137] Specifically, the electronic device may obtain weight information corresponding to the original parameters and each updated parameter, and then fuse the original parameters with each updated parameter based on the weight information corresponding to the original parameters and each updated parameter to obtain the target parameter.
[0138] Step a7: Based on the target parameters, a preset driving decision language model is obtained.
[0139] Specifically, the electronic device generates a preset driving decision-making language model based on the target parameters.
[0140] Step S203 : determining a target driving decision corresponding to the vehicle based on a first driving behavior probability density output by a preset driving decision-specific model and a second driving behavior probability density output by a preset driving decision-specific language model.
[0141] Specifically, the above step S203 may include the following steps:
[0142] Step S2031: input the current driving scene data into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle.
[0143] Specifically, the above step S2031 may include the following steps:
[0144] Step b1: extract features from the video data in front of the vehicle in the current driving scene data to extract visual features of multiple objects.
[0145] The multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles.
[0146] Specifically, electronic devices can use deep learning models to detect and locate at least one of the following objects: vehicles, pedestrians, road facilities, and obstacles in video data from the vehicle's forward direction. These deep learning models can include Faster R-CNN, the You Only Look Once (YOLO) series, and Single Shot MultiBox Detector (SSD). The electronic device can then use a CNN to extract visual features from the detected objects. A CNN consists of multiple convolutional layers, pooling layers, and fully connected layers. Convolutional layers slide kernels across an image to extract local features, such as edges and textures. Pooling layers downsample the output of the convolutional layers, reducing the data size while retaining important feature information. Fully connected layers integrate the extracted features and output a feature vector for the object. For example, based on classic CNN models such as ResNet and VGG, feature extraction can be performed on image regions of detected objects such as vehicles and pedestrians, thereby obtaining visual features for various objects. After pre-training on large-scale image datasets, these models can learn rich visual feature representations that effectively extract features such as shape, color, and texture.
[0147] In step b2, the vehicle operation information, surrounding vehicle movement information, and surrounding road data in the current driving scene data are identified, and the initial graph structure is constructed with vehicles, pedestrians, road facilities, and obstacles as nodes and the spatial position relationships and interactions between vehicles, pedestrians, road facilities, and obstacles as edges.
[0148] Specifically, the electronic device identifies the vehicle's operating information, surrounding vehicle motion information, and surrounding road data within the current driving scene data. Both the vehicle and surrounding vehicles are considered nodes in a graph structure. Each vehicle node contains relevant information about the vehicle, such as its type (car, truck, bus, etc.), speed, location, and direction of travel. Pedestrians, as important road users, are also identified as nodes in the graph structure. Pedestrian nodes contain information such as their location, speed, walking direction, and posture. Road infrastructure includes traffic lights, traffic signs, streetlights, and guardrails. These facilities play a crucial role in guiding and ensuring safety in the transportation system. Using them as nodes clearly demonstrates the relationship between road infrastructure and vehicles and pedestrians. Obstacles on the road, such as construction site fences and fallen objects, are also identified as nodes. Obstacle nodes contain information such as their location, size, and shape.
[0149] Then, based on the spatial positional information between vehicles, pedestrians, road facilities, and obstacles, edges are constructed to represent spatial relationships. For example, if a vehicle has a certain distance and angle relationship with the traffic light ahead, an edge can be established between the vehicle node and the traffic light node. Edge attributes can include information such as the distance and angle between the two. Similarly, spatial relationships between vehicles and pedestrians, or between vehicles and obstacles, can be represented by edges. These edges intuitively display the relative positional relationships between nodes in the graph structure, providing a basis for subsequent path planning and obstacle avoidance decisions. In addition to spatial relationships, interactions between vehicles, pedestrians, road facilities, and obstacles are considered, and edges representing these interactions are constructed. For example, when a traffic light turns green, vehicles start or accelerate according to the light's instructions. This control relationship between the light and the vehicle can be represented by establishing an edge between the light node and the vehicle node, representing the interaction. Similarly, when pedestrians cross the road, vehicles must slow down or stop to yield. This interaction between vehicles and pedestrians can also be represented by edges. These interaction edges can reflect the dynamic relationships between nodes in the graph structure, which helps to analyze and predict behaviors and events in traffic scenarios.
[0150] Through the above steps, vehicles, pedestrians, road facilities and obstacles are identified as nodes, and edges are constructed based on the spatial position relationship and interaction between them, eventually forming the initial graph structure.
[0151] Step b3: Identify the video data in front of the vehicle in the current driving scene data to determine the current driving scene type.
[0152] Specifically, the current driving scene data is identified from the front-view video data of the vehicle to determine the current driving scene type. The driving scene type can be high-altitude driving scene, urban road driving scene, school area driving scene, rural road driving scene, etc., and different current driving scene types correspond to different attention levels.
[0153] Step b4: Based on the current driving scene type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure.
[0154] Specifically, the electronic device determines the semantic information to add based on the current driving scenario type and the characteristics of the edge. For example, for the edge between a vehicle and a pedestrian, semantic information can be added, such as "pedestrians may suddenly cross the road, so slow down and avoid them." For the edge between a vehicle and a traffic light, semantic information can be added, such as "stop and wait when the light is red, and proceed only when the light is green." For the edge between a vehicle and a vehicle in the adjacent lane, semantic information can be added, such as "maintain a safe distance and appropriate speed when overtaking." This semantic information should have clear guiding significance and help understand and process various relationships in the driving scenario.
[0155] Then, the determined semantic information is appropriately added to the edges of the initial graph structure. A semantic information field can be added to the edge attributes to store the corresponding semantic description. Alternatively, the semantic information can be graphically annotated next to or near the edge to provide a more intuitive presentation. For example, a text label can be used to indicate the semantic information next to the edge, or a dedicated "Semantic Description" item can be added to the edge attribute list to record the detailed semantic content.
[0156] Step b5: fuse the visual features, target graph structure, environmental data and time data to generate target features.
[0157] Specifically, the above step b5 may include the following steps:
[0158] Step b51: feature-encode the environmental data and time data to generate environmental data features and time data features.
[0159] Specifically, for environmental data, one-hot encoding or label encoding is often used for categorized environmental data, such as weather conditions. Numerical environmental data, such as temperature, humidity, and light intensity, may require normalization or standardization. To more comprehensively represent the characteristics of environmental data, feature combination and extraction can also be performed. For example, temperature and humidity can be combined into a new feature called "humidity index," or the main components of the data can be extracted through methods such as principal component analysis (PCA), reducing the data dimension while retaining the majority of the information.
[0160] Time data generally includes information such as specific time points (such as year, month, day, hour, minute, and second) or time intervals. First, the time data needs to be converted into a suitable format for encoding. For example, date and time data can be converted into a timestamp (the number of seconds or milliseconds from a fixed starting time point to the current time) to facilitate subsequent numerical calculations. Time data has distinct periodicity, such as 24 hours a day, 7 days a week, and 12 months a year. These periodic features can be extracted, for example, by dividing a day into different time periods (morning, noon, afternoon, and evening) and representing them with one-hot encoding; or by numerically encoding the days of the week, with Monday as 1, Tuesday as 2, and so on. By extracting periodic features, the model can better capture time-related patterns.
[0161] The processed environmental and temporal data features are combined to form a final feature vector. Different types of features can be arranged in a certain order, for example, first the environmental data features and then the temporal data features, to form a complete feature vector that serves as input for subsequent models.
[0162] Step b52: input the visual features, target graph structure, environmental data features and time data features into a preset cross-modal attention interaction network.
[0163] Specifically, the electronic device inputs visual features, target graph structure, environmental data features and time data features into a preset cross-modal attention interaction network.
[0164] In step b53, a cross-modal attention interaction network is preset to calculate the attention scores among the visual features, target graph structure, environmental data features, and temporal data features.
[0165] Specifically, to enable features from different modalities to interact in the same space, a pre-defined cross-modal attention interaction network performs a linear transformation on the features of each modality. By multiplying the features by the corresponding weight matrix, the features of each modality are mapped into a new feature space. Then, attention scores are calculated between the visual features, the target graph structure, the environmental data features, and the temporal data features.
[0166] For example, to calculate the attention score between visual features and other modal features, dot product or other similarity measurement methods are usually used. Suppose we want to calculate the visual features V ′ and target graph structure G ′The attention score between A VG , can be done through the dot product operation A VG =( V ′) T G′ to get a preliminary similarity value. In order to make the attention score more interpretable and numerically stable, it may be scaled. Similarly, the attention score A between the visual feature and the environmental data feature and the time data feature can be calculated VE and A VT , and attention scores between other modality features, such as A GE 、A GT 、A ET wait.
[0167] Step b54, based on the attention scores among the visual features, the target graph structure, the environmental data features and the time data features, the visual features, the target graph structure, the environmental data features and the time data features are fused to generate the target features.
[0168] Specifically, the visual features, target graph structure, environmental data features, and time data features are multiplied by the corresponding attention scores and then added together to obtain the target features.
[0169] Step b6: Determine the complexity of the current driving scene corresponding to the vehicle based on the target features.
[0170] Specifically, the preset scene complexity determination model outputs the current driving scene complexity corresponding to the vehicle based on the target features.
[0171] In step S2032, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene are input into a preset weight determination model, and weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model is output.
[0172] Specifically, the above step S2032 may include the following steps:
[0173] Step c1: Calculate the difference measure between the first driving behavior probability density and the second driving behavior probability density.
[0174] Specifically, the electronic device may calculate a difference measure between the first driving behavior probability density and the second driving behavior probability density based on a preset measurement method.
[0175] Among them, the preset measurement method can be a KL divergence measurement method, a JS divergence measurement method, a Wasserstein distance (bulldozer distance) measurement method, a Euclidean distance, etc. The embodiment of the present application does not specifically limit the preset measurement method.
[0176] In step c2, the difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene are concatenated to generate a concatenated feature.
[0177] Specifically, the electronic device may splice the difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene to generate a splicing feature.
[0178] Step c3: Obtain the current driving decision output accuracy corresponding to the current driving scene type, the preset driving decision-specific model, and the preset driving decision large language model.
[0179] Specifically, the electronic device can identify the vehicle's forward video data within the current driving scene data to determine the current driving scene type. Furthermore, based on historical driving decision records, the electronic device can determine the accuracy of the current driving decision output corresponding to a preset dedicated driving decision model and a preset large language model.
[0180] Step c4: generating dynamic mapping parameters based on the current driving scenario type and the current driving decision output accuracies corresponding to the preset driving decision-specific model and the preset driving decision-making large language model.
[0181] Specifically, the electronic device can determine the dynamic mapping relationship between the preset driving decision-specific model and the preset driving decision-making large language model based on the accuracy difference between them under different driving scenario types. Specifically, the electronic device can map the accuracy of the preset driving decision-specific model and the preset driving decision-making large language model to a weight value or parameter value. For example, a linear function, an exponential function or other suitable mathematical functions can be used to describe this mapping relationship. Assuming that the accuracy of the preset driving decision-specific model in a certain scenario is A1, and the accuracy of the preset driving decision-making large language model is A2, a mapping function f(A1, A2) can be defined so that the output value of f is between 0 and 1, which is used to represent the relative importance or weight of the preset driving decision-specific model and the preset driving decision-making large language model in the scenario.
[0182] The output values of the mapping function for different driving scenario types are then used as dynamic mapping parameters. These parameters can be used to dynamically combine or adjust the outputs of the preset dedicated driving decision model and the preset large language model during the actual driving decision-making process, depending on the current driving scenario type. For example, in one scenario, if the dedicated model is more accurate, the dynamic mapping parameters may cause the final driving decision to rely more on the output of the dedicated model; while in another scenario, if the large language model performs better, the parameters may tend to adopt the recommendations of the large language model.
[0183] Step c5: Mapping the spliced features based on the dynamic mapping parameters to generate low-dimensional features.
[0184] Specifically, the electronic device may utilize a preset mapping processing method to perform mapping processing on the splicing features based on dynamic mapping parameters to generate low-dimensional features.
[0185] The preset mapping processing method may be linear mapping or nonlinear mapping.
[0186] Take the linear mapping method as an example. Assume that the concatenated feature is an n-dimensional vector X=[x1,x2,⋯,x n ], the dynamic mapping parameter is an n-dimensional weight vector W=[w1,w2,⋯,w n ], through the linear transformation Y=W T X= A low-dimensional scalar or vector Y is obtained. Here, Y is the portion of the low-dimensional feature after the mapping process. This process can be repeated to generate multiple low-dimensional features using different weight vectors W, which are then combined into the final low-dimensional feature vector.
[0187] Nonlinear mapping methods, such as neural networks, use concatenated features as neural network inputs, and dynamic mapping parameters can serve as neural network parameters or influence parameter adjustments. Neural networks map high-dimensional concatenated features to a low-dimensional space through multiple layers of nonlinear transformations. For example, a multi-layer perceptron (MLP) can be used. During training, network weights are adjusted based on the dynamic mapping parameters, enabling the network to learn feature representations that are more appropriate for the current driving scenario and generate representative low-dimensional features.
[0188] In step c6, the low-dimensional features are input into a preset hidden layer and feature processing is performed on the low-dimensional features.
[0189] Specifically, when the low-dimensional features are input to the neurons of the preset hidden layer, each neuron will perform a weighted summation on the input signal. Suppose the input low-dimensional feature vector is x=[x1,x2,⋯,x n ], the weight matrix between the neuron and the input layer is W=[w ij ] (where i represents the index of the hidden layer neuron and j represents the index of the input layer neuron), the bias vector is b=[b1,b2,⋯,b m ] (m is the number of hidden layer neurons), then the input zi of the i-th hidden layer neuron is zi= Then, we apply an activation function (such as ReLU, Sigmoid, Tanh, etc.) to zi. The role of the activation function is to introduce nonlinear factors so that the neural network can learn more complex functional relationships. For example, when using the ReLU activation function, if z i >0, then output a i =z i If z i ≤0, then output a i=0.
[0190] If the pre-set hidden layer contains multiple layers, the output of the previous hidden layer will serve as the input to the next hidden layer, and the above process of weight calculation and activation function application will be repeated. With each hidden layer, the features are further abstracted and transformed, gradually forming a more advanced feature representation. For example, the first hidden layer may extract a few simple feature combinations, such as the relationship between vehicle speed and the distance to surrounding vehicles. The second hidden layer may build on this and combine it with environmental data features to learn more complex patterns, such as the combined impact of vehicle speed and distance on safe driving in different weather conditions.
[0191] After processing by the preset hidden layer, the final output is a transformed and abstracted feature representation. These feature representations can be used as input to subsequent layers (such as the output layer) to predict or classify driving decisions. For example, in a driving decision model, the features output by the hidden layer can be input to the output layer, which uses these features to determine whether the vehicle should accelerate, decelerate, or steer.
[0192] Step c7: Preset the weight determination model to determine the output layer, and output the weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model respectively.
[0193] Specifically, the output layer calculates the weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model based on the feature representation output by the hidden layer, and outputs the calculated weight information. This weight information reflects the relative importance of the preset driving decision-specific model and the preset driving decision-making large language model in the final driving decision under the current input driving scenario and model performance conditions. For example, the output weight information may indicate that in a congested urban road scenario, the weight of the preset driving decision-specific model is 0.4, and the weight of the preset driving decision-making large language model is 0.6, which means that in this scenario, the decision of the large language model has a greater impact on the final driving decision.
[0194] Step S2033: Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are integrated to determine the target driving decision corresponding to the vehicle.
[0195] Specifically, based on the weight information, the electronic device performs a weighted calculation on a first driving behavior probability density output by a preset dedicated driving decision model and a second driving behavior probability density output by a preset large language model. Specifically, the weighted calculation may be a weighted sum of the probabilities of each driving behavior. For example, for acceleration, the fused probability = the probability of acceleration in the first driving behavior probability density × the dedicated model weight + the probability of acceleration in the second driving behavior probability density × the large language model weight.
[0196] After merging the new driving behavior probability density, the driving behavior with the highest probability is typically selected as the target driving decision. For example, if the probability of maintaining the current state in the fused probability density is 0.6, the highest probability among all driving behaviors, then the target driving decision is to maintain the current state. This method is simple and direct, enabling quick decision-making and is suitable for most common driving scenarios.
[0197] The driving decision determination method provided in the embodiments of the present application obtains an original large language model, thereby leveraging the existing strong foundation of the original large language model. This avoids the large computational resources and time costs required to train the model from scratch. The weight matrix of the original large language model is decomposed into at least two low-rank matrices, thereby reducing the complexity of the original large language model and facilitating model optimization and understanding. The original parameters corresponding to the original large language model are then frozen, preserving the general language knowledge and capabilities learned by the original large language model. This pre-trained knowledge is very helpful for understanding and processing natural language information related to driving. In addition, freezing the original parameters helps stabilize the model training process. A preset driving decision expert library is obtained to introduce professional knowledge and experience to meet actual driving needs. The expert experience, industry rules, and historical cases in the preset driving decision expert library are identified to construct a knowledge graph. This structured knowledge representation facilitates the understanding and utilization of subsequent models, improving the accessibility and operability of knowledge. The knowledge graph is input into the original large language model to enrich the knowledge reserve of the original large language model, enabling it to better understand driving-related concepts, rules, and experience. For example, the original large language model can use the knowledge graph to understand the meaning of different traffic signs and driving strategies under various road conditions. This allows it to reason and make decisions based on more comprehensive knowledge when handling driving-related tasks. The original large language model recognizes the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments. This allows the original large language model to rehearse and analyze a variety of possible driving situations in a virtual environment, ranging from simple everyday driving scenarios to complex special situations such as inclement weather and traffic accident scenes. Through these simulations, the model can pre-learn and master the optimal driving decisions in different scenarios, improving its ability to cope with real-world driving scenarios. Rewards are evaluated for each simulated driving decision based on a preset reward function, generating reward feedback data. The preset reward function evaluates simulated driving decisions across multiple dimensions, including safety, efficiency, comfort, and social compatibility, providing a comprehensive assessment of the quality of the decision. This multi-dimensional evaluation approach avoids focusing on a single metric (such as safety) while ignoring other important aspects (such as comfort and social compatibility). Reward feedback data determines the direction for updating the parameters of each low-rank matrix, making the parameter updates targeted and accurate. Based on reward feedback, the model can focus on adjusting parameters that have a significant impact on achieving high rewards, avoiding blind parameter updates that could lead to performance degradation. The parameters of each low-rank matrix are updated based on the update direction, resulting in updated parameters corresponding to each low-rank matrix. These updated parameters enable the original large speech model to better handle driving decision-making tasks, improving decision accuracy and reliability. For example, the updated model may be able to more accurately judge the driving intentions of other vehicles in complex road conditions, leading to safer and more reasonable driving decisions.The original parameters are then fused with the updated parameters to obtain the target parameters. This leverages the strengths of both. The original parameters retain the general language understanding capabilities and extensive knowledge base of the large language model, while the updated parameters imbue the model with specific capabilities and expertise for driving decision-making tasks. By fusing these two parameters, the model can leverage its general language processing capabilities to understand and process a variety of driving-related natural language information while also making accurate driving decisions based on the knowledge in the expert database, achieving an organic combination of general and specialized capabilities. Based on the target parameters, a pre-set large language model for driving decisions is obtained. This makes the pre-set large language model for driving decisions a model specifically optimized for driving decision-making tasks. Combining the strong foundation of the original large language model with the specialized driving decision-making capabilities gained through parameter updates from the expert database, it can efficiently process a variety of driving-related information and generate accurate and reasonable driving decisions. This model can better meet the real-time, accuracy, and reliability requirements of autonomous or assisted driving systems, providing strong support for safe vehicle operation.
[0198] In addition, feature extraction is performed on the video data ahead of the vehicle in the current driving scene data, extracting visual features of various objects. By extracting visual features of various objects, such as vehicles, pedestrians, road facilities, and obstacles, the system accurately perceives various objects in the environment ahead of the vehicle. The system identifies the vehicle's operating information, the motion information of surrounding vehicles, and the surrounding road data in the current driving scene data. Using vehicles, pedestrians, road facilities, and obstacles as nodes and the spatial relationships and interactions between them as edges, an initial graph structure is constructed, clearly presenting the relationships between various elements in the current driving scene. This graph structure facilitates analysis of complex driving situations. The system can use this graph structure to quickly analyze the interactions between multiple objects. The system identifies the video data ahead of the vehicle in the current driving scene data and determines the current driving scene type. Different driving scenarios have different characteristics and requirements. Once the scene type is determined, targeted handling can be achieved, thereby improving the accuracy and effectiveness of decision-making. Based on the current driving scene type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure. The information contained in the graph structure is further enriched, and the target graph structure with semantic information provides a more optimized basis for the system to make decisions. The environmental data and time data are feature encoded to generate environmental data features and time data features, so that these data of different types and formats can be converted into a unified feature representation that can be processed by the model. The visual features, target graph structure, environmental data features and time data features are input into the preset cross-modal attention interaction network to achieve the integration of multi-source information. The preset cross-modal attention interaction network calculates the attention scores between the visual features, target graph structure, environmental data features and time data features, so that it can automatically focus on the most important information for understanding and making decisions in the current driving scene. Based on the attention scores between the visual features, target graph structure, environmental data features and time data features, the visual features, target graph structure, environmental data features and time data features are fused to generate target features. The generated target features can more comprehensively and accurately describe the driving scene. The fused target features combine the information advantages of different modalities, encompassing both object appearance and relationship information while also accounting for environmental and temporal factors. This provides a high-quality feature representation for determining the complexity of the driving scene and making driving decisions, helping to improve the accuracy and reliability of decisions. Based on the target features, the complexity of the current driving scene corresponding to the vehicle is determined, thereby accurately assessing the complexity of the current driving scene.
[0199] A difference measure is calculated between the first driving behavior probability density and the second driving behavior probability density. This difference measure provides a clear understanding of the degree of difference between the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model. This helps identify differences between the two models in their understanding of the current driving scenario and decision-making tendencies, providing a basis for subsequent comprehensive consideration of their outputs. The difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scenario are concatenated to generate a concatenated feature. This multi-dimensional information more comprehensively describes the relevant factors of the current driving decision, providing a more robust data foundation for subsequent model analysis. The accuracy of the current driving decision output corresponding to the current driving scenario type, the preset dedicated driving decision model, and the preset large language model is obtained. Based on the current driving scenario type and the accuracy of the current driving decision output corresponding to the preset dedicated driving decision model and the preset large language model, dynamic mapping parameters are generated. This allows for dynamic adaptation to different driving scenarios. Model performance varies in different scenarios, and by dynamically generating mapping parameters, the processing of the concatenated features can be optimized for each specific scenario. The concatenated features are mapped based on dynamic mapping parameters to generate low-dimensional features, effectively reducing the dimensionality of the data. The low-dimensional features are then fed into a pre-set hidden layer for feature processing, revealing deeper information and patterns within them. The output layer of the model, which uses pre-set weights, outputs weight information corresponding to the pre-set dedicated driving decision model and the pre-set large language model for driving decisions. This ensures a reasonable weight allocation between the two models for the current driving scenario, improving the quality of driving decisions. Based on this weight information, the first and second driving behavior probability densities are fused to determine the target driving decision for the vehicle. By fusing the driving behavior probability densities output by the two models using this weight information, the system leverages the domain-specific expertise of the pre-set dedicated driving decision model and the general understanding and reasoning capabilities of the pre-set large language model for driving decisions. Furthermore, this fusion process reduces the errors and uncertainties that may arise from a single model, improving the reliability and stability of driving decisions.
[0200] For example, in order to better introduce the driving decision determination method provided in the embodiment of the present application, as shown in FIG. Figure 3 As shown, the embodiment of the present application provides a timing framework diagram of a driving decision determination method, which is described in detail below:
[0201] 1) Data Input: Visual and Environmental Data: The left side displays various data sources, including images and semantic interpretation of images (e.g., describing vehicle driving conditions and lane occupancy on highways); video data in front of the vehicle, motion information of surrounding vehicles, vehicle motion information, surrounding road data, environmental data (e.g., weather), and time data. This data comprehensively describes the vehicle's driving scenario from various perspectives.
[0202] Prompt word guidance: Through the prompt word "Please give the current decision and the corresponding probability", the model is guided to make decisions based on the input data and give probabilistic judgments.
[0203] 2) Cloud Processing: Data Collection and Knowledge Update: Various data are transmitted to the cloud-based data collection system, which collects and integrates the data and uses it to update the lane-changing decision-making expertise. This is the knowledge-retention phase of the entire decision-making system, ensuring that the system is always based on the latest data.
[0204] Model fine-tuning: Low-rank adaptive algorithm (LoRA): This algorithm is used to optimize and adjust the model. LoRA uses low-rank matrix approximation to efficiently fine-tune the model to adapt to specific lane-changing decision tasks without changing too many model parameters.
[0205] Fine-tuned Large Language Model - Lane Change Decision: The large language model fine-tuned by the LoRA algorithm is specifically used for lane change decision tasks. It can understand and process various input data information and preliminarily provide results such as probability distribution related to lane change decisions.
[0206] 3) Decision Analysis: Scenario Complexity Assessment: The model collaborative decision-making framework assesses the complexity of the scenario, including environmental conditions, road conditions, traffic elements, and time factors. This step comprehensively determines the complexity of the current driving scenario and provides a basis for subsequent decision-making.
[0207] Dynamic Weight Update Network: Based on the results of the scene complexity assessment, the dynamic weight update network adjusts the weights of different models in decision-making. For example, in complex scenes, the weights of certain models that are better at handling complex situations may be increased.
[0208] Dynamic weighted fusion strategy: Using the adjusted weights, the output probability distributions of different models (such as the fine-tuned large language model and the traditional decision model) are fused through a dynamic weighted fusion strategy to obtain the final decision probability, which is used to determine whether to change lanes left, right, or go straight.
[0209] 4) Feedback and Update: Decision results are updated and fed back to the cloud-based data collection system to further optimize and update the lane-changing decision-making expertise base. They are also fed back to the model collaborative decision-making framework to continuously adjust model parameters and decision-making strategies, thereby continuously improving the accuracy and reliability of lane-changing decisions.
[0210] 5) Traditional Decision Model: The figure also mentions the traditional decision model as a research topic. Although not elaborated on in detail, it is part of the overall decision-making system and works in conjunction with the fine-tuned large language model to complete the lane change decision task.
[0211] This embodiment also provides a driving decision determination device for implementing the aforementioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. While the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0212] This embodiment provides a driving decision determination device, such as Figure 4 Shown, including:
[0213] An acquisition module 301 is configured to acquire current driving scene data corresponding to the vehicle, the current driving scene data including at least one of the following: video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data;
[0214] An input module 302 is used to input current driving scene data into a preset driving decision-specific model and a preset driving decision-making language model;
[0215] The determination module 303 is used to determine the target driving decision corresponding to the vehicle based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model.
[0216] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0217] The driving decision determination device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0218] An embodiment of the present invention further provides an electronic device having the above Figure 4The driving decision determination device shown.
[0219] See also Figure 5 , Figure 5 is a structural diagram of an electronic device provided by an optional embodiment of the present invention, such as Figure 5 As shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. The electronic device also includes a communication interface 30 for the electronic device to communicate with other devices or a communication network.
[0220] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A driving decision determination method, characterized in that: Methods include: Obtaining current driving scene data corresponding to the vehicle, the current driving scene data including at least one of video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data; Inputting the current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model, wherein the preset driving decision-specific model outputs a first driving behavior probability density, and the preset driving decision-making large language model outputs a second driving behavior probability density; Inputting current driving scene data into a preset scene complexity determination model; Perform feature extraction on the video data in front of the vehicle in the current driving scene data to extract visual features of multiple objects; the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles; Identify the vehicle's operating information, surrounding vehicle motion information, and surrounding road data in the current driving scene data, and construct an initial graph structure using vehicles, pedestrians, road facilities, and obstacles as nodes and the spatial relationships and interactions between them as edges. Identify the front video data of the vehicle in the current driving scene data to determine the current driving scene type; Based on the current driving scenario type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure. Fuse visual features, target graph structure, environmental data, and time data to generate target features; Based on the target features, determine the complexity of the current driving scene corresponding to the vehicle; Inputting the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene into a preset weight determination model, and outputting weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model respectively; Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are fused to determine the target driving decision corresponding to the vehicle.
2. The method according to claim 1, characterized in that The training process of the preset driving decision-making large language model includes: Get the original large language model; Decomposing the weight matrix of the original large language model into at least two low-rank matrices; Freezing original parameters corresponding to the original large language model; Obtaining a preset driving decision expert database; the preset driving decision expert database includes expert experience, industry rules, and historical cases; Based on the preset driving decision expert database, the parameters of each low-rank matrix are updated to obtain updated parameters corresponding to each low-rank matrix; Fusing the original parameters with the updated parameters to obtain target parameters; Based on the target parameters, the preset driving decision language model is obtained.
3. The method according to claim 2, characterized in that The updating of the parameters of each low-rank matrix based on the preset driving decision expert database to obtain the updated parameters corresponding to each low-rank matrix includes: Identify the expert experience, industry rules, and historical cases in the preset driving decision expert database to construct a knowledge graph; Inputting the knowledge graph into the original large language model; The original large language model recognizes the knowledge graph and outputs corresponding simulated driving decisions under various simulated driving environments; Performing a reward evaluation on each of the simulated driving decisions based on a preset reward function to obtain reward feedback data; the preset reward function includes a safety reward, an efficiency reward, a comfort reward, and a social compatibility reward; Based on the reward feedback data, determining the direction of updating the parameters of each low-rank matrix; The parameters of each of the low-rank matrices are updated based on the update direction to obtain the updated parameters corresponding to each of the low-rank matrices.
4. The method according to claim 3, characterized in that The preset reward function formula is: ; Where N is the number of simulations, For the safety reward, For the efficiency reward, For the said comfort reward, Reward for said social compatibility; ; Where C is the number of collisions, α is the safety penalty coefficient, and α>0; ; Among them, T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency bonus coefficient, and β>0; ; The smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration rate of change, which is expressed as ; n is the number of time intervals during driving, Δa i is the acceleration change in the i-th time interval, and the reasonable following distance is maintained by the actual following distance d actual and ideal following distance d ideal The deviation is measured as , m is the number of times of following the car; γ is the comfort reward coefficient, S a and S d are the normalized constants of the sum of squares of acceleration change rate and following distance deviation, respectively, and γ>0; when the acceleration changes smoothly and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative; ; The degree of traffic flow can be measured by the average speed change rate of surrounding vehicles, which is expressed as , when simulated driving decisions promote smooth traffic flow, , let the number of lane changes that are reasonable and completed successfully be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ >0, when the simulated driving decision helps to improve the average speed of surrounding vehicles and the lane change is completed reasonably, the social compatibility reward is positive.
5. The method according to claim 1, characterized in that Fusing the visual features, the target graph structure, the environmental data, and the time data to generate target features includes: Performing feature encoding on the environmental data and the time data to generate environmental data features and time data features; Inputting the visual features, the target graph structure, the environmental data features, and the time data features into a preset cross-modal attention interaction network; The preset cross-modal attention interaction network calculates an attention score between the visual feature, the target graph structure, the environmental data feature, and the temporal data feature; Based on the attention scores among the visual features, the target graph structure, the environmental data features and the temporal data features, the visual features, the target graph structure, the environmental data features and the temporal data features are fused to generate the target features.
6. The method according to claim 1, characterized in that Inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into a preset weight determination model, and outputting weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model, respectively, including: Calculating a difference measure between the first driving behavior probability density and the second driving behavior probability density; Splicing the difference measure, the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity to generate a splicing feature; Obtaining the current driving decision output accuracy corresponding to the current driving scene type, the preset driving decision-specific model, and the preset driving decision large language model respectively; generating dynamic mapping parameters based on the current driving scene type and the current driving decision output accuracies corresponding to the preset driving decision-specific model and the preset driving decision-making large language model, respectively; Performing mapping processing on the splicing features based on the dynamic mapping parameters to generate low-dimensional features; Inputting the low-dimensional features into a preset hidden layer and performing feature processing on the low-dimensional features; The output layer in the preset weight determination model outputs weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model respectively.
7. A driving decision determination device, characterized in that: The device comprises: an acquisition module, configured to acquire current driving scene data corresponding to the vehicle, the current driving scene data including at least one of video data in front of the vehicle, vehicle operation information, motion information of surrounding vehicles, surrounding road data, environmental data, and time data; An input module, configured to input current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model, wherein the preset driving decision-specific model outputs a first driving behavior probability density, and the preset driving decision-making large language model outputs a second driving behavior probability density; A determination module is used to input the current driving scene data into a preset scene complexity determination model; perform feature extraction on the video data in front of the vehicle in the current driving scene data, and extract visual features of multiple objects; the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles; identify the vehicle operation information, surrounding vehicle movement information, and surrounding road data in the current driving scene data, and regard vehicles, pedestrians, road facilities, and obstacles as nodes, and the spatial position relationship and interaction between vehicles, pedestrians, road facilities, and obstacles as edges to construct an initial graph structure; identify the video data in front of the vehicle in the current driving scene data, and determine the current driving scene type; based on the current driving scene type, add semantic information to at least one edge in the initial graph structure to generate a target graph structure; fuse the visual features, target graph structure, environmental data and time data to generate target features; based on the target features, determine the complexity of the current driving scene corresponding to the vehicle; input the first driving behavior probability density, the second driving behavior probability density and the complexity of the current driving scene into the preset weight determination model, and output the weight information corresponding to the preset driving decision-specific model and the preset driving decision large language model respectively; based on the weight information, fuse the first driving behavior probability density and the second driving behavior probability density to determine the target driving decision corresponding to the vehicle.
8. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the driving decision determination method according to any one of claims 1 to 6 by executing the computer instructions.
Citation Information
Patent Citations
Lane change prediction method and device for target vehicle
CN113147766A
Intelligent emergency rescue strategy output and emergency rescue vehicle control system and method
CN118505168A
Automatic driving three-dimensional scene data preprocessing method and system based on large language model
CN118781267A
Automatic driving vehicle trajectory planning method
CN119283896A