Driving decision determination method and device and electronic equipment
By adopting driving decision determination methods in autonomous vehicles and combining the output of special models and large language models, the problems of weak humanization and poor generalization capabilities of decision-making and planning strategies in hybrid environments are solved, and more accurate and reliable driving decisions are achieved.
Patent Information
- Application Number
- CN202510542429.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In a mixed environment, autonomous vehicles have weak human-like decision-making planning strategies and poor generalization capabilities, resulting in inaccurate decision-making and affecting the operational safety and efficiency of the transportation system.
A driving decision determination method is adopted. By obtaining the current driving scenario data, inputting it into the preset driving decision-making special model and the large language model, combining the outputs of the two models to determine the target driving decision. This method combines the fast decision-making ability of a dedicated model and the language understanding and knowledge reasoning ability of a large language model, improving the accuracy and reliability of decision-making.
By integrating the decision results of dedicated models and large language models, the possibility of driving behavior can be more accurately evaluated, the accuracy and reliability of decisions can be improved, and the decision-making ability of autonomous vehicles in a mixed-traffic environment can be enhanced.
Smart Images

Figure CN120057041A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent driving, and particularly to a method, device, and electronic device for determining driving decisions. Background Art
[0002] From a technical perspective, for an autonomous vehicle to operate safely and efficiently on the road, three modules, namely perception, decision-making and planning, and control execution, need to work together. Decision-making and planning is regarded as the intelligent brain of an autonomous vehicle. It coordinates the overall situation and plans short-term avoidance and long-term routes based on road conditions, traffic rules, vehicle status, etc. It is a crucial link for the safe and efficient operation of an autonomous vehicle.
[0003] Lane change is generally considered to be one of the most difficult-to-predict driving tasks, and it is also an important cause of traffic congestion and vehicle collisions. Therefore, there is an urgent need to study lane change decision-making and planning strategies, especially when we are gradually entering an era of mixed operation of autonomous driving and manual driving.
[0004] In a mixed traffic environment, an important challenge for autonomous vehicles to be compatible with society is the weak anthropomorphic degree and poor generalization ability of decision-making and planning strategies. Whether the decision-making and planning strategy has anthropomorphism and generalization is the key to determining whether the vehicle can accurately complete various driving tasks, which significantly affects the operational safety and efficiency of the traffic system. However, current research on autonomous driving decision-making and planning usually focuses on rule-based or end-to-end modeling, with strong professional reasoning ability but insufficient common sense reasoning.
[0005] Therefore, how to improve the accuracy of autonomous driving decision-making and planning has become an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a method, device, and electronic device for determining driving decisions to solve the problem of how to improve the accuracy of autonomous driving decision-making and planning.
[0007] In a first aspect, the present invention provides a method for determining a driving decision, the method comprising: Obtaining current driving scenario data corresponding to the vehicle itself, the current driving scenario data including at least one of video data in front of the vehicle itself, running information of the vehicle itself, movement information of surrounding vehicles, surrounding road data, environmental data, and time data; Inputting the current driving scenario data into a preset dedicated driving decision model and a preset large language driving decision model; Determining a target driving decision corresponding to the vehicle itself based on a first driving behavior probability density output by the preset dedicated driving decision model and a second driving behavior probability density output by the preset large language driving decision model.
[0008] The driving decision determination method provided by the embodiment of this application obtains the current driving scenario data corresponding to the vehicle of this application. The current driving scenario data includes at least one of the video data in front of the vehicle itself, the running information of the vehicle itself, the movement information of surrounding vehicles, the surrounding road data, the environmental data, and the time data. By collecting the current driving scenario data, the vehicle of this application can comprehensively perceive the driving environment it is in, providing a rich and accurate information basis for subsequent decisions. Input the current driving scenario data into a preset dedicated driving decision model and a preset large language model for driving decisions. The respective advantages of these two models can be fully utilized. The dedicated model is usually optimized for specific driving scenarios and tasks, and is good at handling some common and patterned driving situations, and can quickly and accurately give corresponding decision suggestions. The large language model has powerful language understanding and knowledge reasoning capabilities, can handle complex semantic information and rare driving scenarios, and has good processing capabilities for some situations that require comprehensive judgment and knowledge application. The combined use of the two models enables the vehicle to have a more flexible decision-making method when facing various different types of driving scenarios. Based on the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model for driving decisions, determine the target driving decision corresponding to the vehicle of this application. This way of fusing decisions can integrate the results of the two models and avoid the limitations that may exist in a single model. By analyzing and fusing the two probability densities, the possibility of various driving behaviors can be evaluated more accurately, so as to obtain a target driving decision that more conforms to the actual driving scenario, improving the accuracy and reliability of the decision.
[0009] In an optional implementation manner, the training process of the preset large language model for driving decisions includes: Obtain the original large language model; Decompose the weight matrix of the original large language model into at least two low-rank matrices; Freeze the original parameters corresponding to the original large language model; Obtain a preset driving decision expert library; the preset driving decision expert library includes expert experience, industry rules, and historical cases; Based on the preset driving decision expert library, update the parameters of each low-rank matrix to obtain the updated parameters corresponding to each low-rank matrix respectively; Fuse the original parameters with each updated parameter to obtain the target parameters; Based on the target parameters, obtain the preset large language model for driving decisions.
[0010] The driving decision-making determination method provided by the embodiments of the present application obtains the original large language model, so that the existing powerful foundation of the original large language model can be utilized. This avoids the large amount of computing resources and time costs required for training the model from scratch. The weight matrix of the original large language model is decomposed into at least two low-rank matrices, which can reduce the complexity of the original large language model and facilitate model optimization and understanding. Then, the original parameters corresponding to the original large language model are frozen, and the general language knowledge and capabilities learned by the original large language model will not be damaged. This pre-trained knowledge is very helpful for understanding and processing natural language information related to driving. In addition, freezing the original parameters helps to stabilize the training process of the model. A preset driving decision expert library is obtained to introduce professional knowledge and experience and meet the actual driving needs. Based on the preset driving decision expert library, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix respectively, so that the model can be customized according to the specific requirements of driving decisions. The information in the expert library can guide the preset driving decision large language model to learn specific patterns and rules related to driving, enabling the preset driving decision large language model to gradually adapt to various factors in the driving scenario, such as traffic flow and road condition changes. In this way, the preset driving decision large language model can better capture the relationship between driving decisions and various factors, and thus generate more decisions that conform to the actual driving situation. The original parameters and each updated parameter are fused to obtain the target parameters, thereby being able to give full play to the advantages of both. The original parameters retain the general language understanding ability and extensive knowledge base of the large language model, while the updated parameters enable the model to possess specific capabilities and professional knowledge for driving decision-making tasks. By fusing these two parts of parameters, the model can not only utilize its general language processing ability to understand and process various natural language information related to driving, but also make accurate driving decisions based on the knowledge in the expert library, realizing the organic combination of general capabilities and professional capabilities. Based on the target parameters, a preset driving decision large language model is obtained, making the preset driving decision large language model a model optimized specifically for driving decision-making tasks. It combines the powerful foundation of the original large language model and the professional driving decision-making capabilities obtained through parameter update by the expert library, can efficiently process various information related to driving, and generate accurate and reasonable driving decisions. This model can better meet the requirements of real-time performance, accuracy, and reliability of autonomous driving or assisted driving systems, providing strong support for the safe driving of vehicles.
[0011] In an alternative embodiment, based on the preset driving decision expert library, updating the parameters of each low-rank matrix to obtain the updated parameters corresponding to each low-rank matrix respectively includes: Identifying the expert experience, industry rules, and historical cases in the preset driving decision expert library to construct a knowledge graph; Inputting the knowledge graph into the original large language model; The original large language model identifies the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments; Based on a preset reward function, reward evaluation is performed on each simulated driving decision to obtain reward feedback data; the preset reward function includes safety rewards, efficiency rewards, comfort rewards, and social compatibility rewards; Based on the reward feedback data, the update directions of the parameters of each low-rank matrix are determined; Based on the update directions, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix respectively.
[0012] The driving decision determination method provided by the embodiments of the present application identifies expert experience, industry rules, and historical cases in a preset driving decision expert library to construct a knowledge graph. This structured knowledge representation facilitates the subsequent understanding and utilization of the model, improving the accessibility and operability of knowledge. Inputting the knowledge graph into the original large language model can enrich the knowledge reserve of the original large language model, enabling it to better understand driving-related concepts, rules, and experiences. For example, the original large language model can understand the meanings of different traffic signs and driving strategies under various road conditions through the knowledge graph, and thus, when dealing with driving-related tasks, it can reason and make decisions based on more comprehensive knowledge. The original large language model identifies the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments. This enables the original large language model to preview and analyze various possible driving situations in a virtual environment, covering from simple daily driving scenarios to complex special situations, such as bad weather and traffic accident scenes. Through this simulation, the model can learn and master the best driving decisions in different scenarios in advance, improving its ability to handle actual driving scenarios. Based on a preset reward function, reward evaluation is performed on each simulated driving decision to obtain reward feedback data. The preset reward function evaluates the simulated driving decisions from multiple dimensions such as safety, efficiency, comfort, and social compatibility, and can comprehensively measure the pros and cons of the decisions. This multi-dimensional evaluation method avoids the problem of only focusing on a single indicator (such as safety) while ignoring other important aspects (such as comfort and social compatibility). Based on the reward feedback data, the update directions of the parameters of each low-rank matrix are determined, making the parameter update targeted and accurate. The model can focus on adjusting the parameters that have a greater impact on obtaining high rewards according to the reward feedback, avoiding blindly updating parameters and causing a decline in model performance. Based on the update directions, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix respectively. The updated parameters enable the original large speech model to better handle driving decision tasks, improving the accuracy and reliability of the decisions. For example, the updated model may be able to more accurately judge the driving intentions of other vehicles under complex road conditions, thus making safer and more reasonable driving decisions.
[0013] In an alternative embodiment, the preset reward function formula is: ; where N is the number of simulations, is the safety reward, is the efficiency reward, is the comfort reward, is the social compatibility reward; ; where C is the number of collisions, α is the safety penalty coefficient, and α > 0; ; where T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency reward coefficient, and β > 0; ; where the smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration change rates, denoted as ; n is the number of time intervals during driving, Δa i is the acceleration change amount in the i-th time interval, and the reasonable following distance is maintained by measuring the deviation between the actual following distance d actual and the ideal following distance d ideal , denoted as , m is the number of times involving following situations; γ is the comfort reward coefficient, S a and S d are the normalization constants of the sum of the squares of the acceleration change rates and the sum of the squares of the following distance deviations respectively, and γ > 0; when the acceleration change is smooth and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative; ; where the traffic flow smoothness can be measured by the average speed change rate of surrounding vehicles, denoted as , when the simulated driving decision promotes traffic flow, , let the number of times of reasonable and successful lane changes be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ > 0, when the simulated driving decision helps to increase the average speed of surrounding vehicles and the lane changes are completed reasonably, the social compatibility reward is positive.
[0014] The driving decision determination method provided by the embodiment of the present application, where the preset reward function evaluates the simulated driving decision from multiple dimensions such as safety, efficiency, comfort, and social compatibility, can comprehensively measure the advantages and disadvantages of the decision. This multi-dimensional evaluation method avoids the problem of only focusing on a single indicator (such as safety) while ignoring other important aspects (such as comfort and social compatibility).
[0015] In an alternative implementation manner, based on the first driving behavior probability density output by the preset driving decision dedicated model and the second driving behavior probability density output by the preset driving decision large language model, determining the target driving decision corresponding to the vehicle of this vehicle includes: Inputting the current driving scene data into the preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle of this vehicle; Inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into the preset weight determination model, and outputting the weight information corresponding to the preset driving decision dedicated model and the preset driving decision large language model respectively; Based on the weight information, performing a fusion process on the first driving behavior probability density and the second driving behavior probability density to determine the target driving decision corresponding to the vehicle of this vehicle.
[0016] The driving decision determination method provided by the embodiment of the present application inputs the current driving scene data into the preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle of this vehicle. Thus, the complexity of the driving scene can be accurately evaluated according to the input current driving scene data. Inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into the preset weight determination model, and outputting the weight information corresponding to the preset driving decision dedicated model and the preset driving decision large language model respectively. This way of flexibly adjusting the model weights according to the scene complexity can optimize the entire driving decision-making process and give full play to the advantages of the two models. It avoids fixedly using a certain model or adopting a fixed weight allocation method in all scenarios, thereby improving the accuracy and adaptability of the decision. Based on the weight information, performing a fusion process on the first driving behavior probability density and the second driving behavior probability density to determine the target driving decision corresponding to the vehicle of this vehicle. By fusing the driving behavior probability densities output by the two models through the weight information, the professional knowledge of the preset driving decision dedicated model in a specific field and the general understanding and reasoning ability of the preset driving decision large language model can be fully utilized. In addition, the fusion process can reduce the errors and uncertainties that may occur in a single model, and improve the reliability and stability of the driving decision.
[0017] In an alternative implementation manner, inputting the current driving scene data into the preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle of this vehicle includes: Extract the visual features of multiple objects from the front vehicle video data in the current driving scenario data; the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles; Identify the running information of the host vehicle, the movement information of surrounding vehicles, and the surrounding road data in the current driving scenario data. Regarding vehicles, pedestrians, road facilities, and obstacles as nodes, and the spatial position relationships and interactions between vehicles, pedestrians, road facilities, and obstacles as edges, construct an initial graph structure; Identify the front vehicle video data in the current driving scenario data to determine the type of the current driving scenario; Based on the type of the current driving scenario, add semantic information to at least one edge in the initial graph structure to generate a target graph structure; Fuse the visual features, the target graph structure, the environmental data, and the time data to generate target features; Based on the target features, determine the complexity of the current driving scenario corresponding to the host vehicle.
[0018] The driving decision determination method provided in the embodiment of the present application performs feature extraction on the video data in front of the vehicle in the current driving scene data, and extracts the visual features of multiple objects. By extracting the visual features of multiple objects such as vehicles, pedestrians, road facilities, obstacles, etc., the system can accurately perceive various objects in the environment in front of the vehicle. The vehicle operation information, surrounding vehicle movement information and surrounding road data in the current driving scene data are identified, and the vehicles, pedestrians, road facilities and obstacles are regarded as nodes. The spatial position relationship and interaction between vehicles, pedestrians, road facilities and obstacles are used as edges to construct an initial graph structure, which can clearly present the relationship between the various elements in the current driving scene. This graph structure provides convenience for analyzing complex driving situations. The system can quickly analyze the mutual influence between multiple objects through the graph structure. The video data in front of the vehicle in the current driving scene data is identified to determine the current driving scene type. Different driving scenes have different characteristics and requirements. After clarifying the scene type, various situations can be handled in a targeted manner, thereby improving the accuracy and effectiveness of decision-making. Based on the current driving scene type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure. The information contained in the graph structure is further enriched, and the target graph structure with semantic information provides a more optimized basis for the system to make decisions. The visual features, target graph structure, environmental data and time data are fused to generate target features. The fused target features have stronger representation capabilities and can more accurately describe the complexity and characteristics of the current driving scene. It can capture information that cannot be reflected by a single data source and provide a more reliable basis for subsequent decision-making. Based on the target features, the complexity of the current driving scene corresponding to the vehicle is determined. In this way, the complexity of the current driving scene can be accurately evaluated.
[0019] In an optional implementation, the visual features, the target graph structure, the environmental data and the time data are fused to generate the target features, including: Perform feature encoding on the environmental data and the time data to generate environmental data features and time data features; Inputting visual features, target graph structure, environmental data features and temporal data features into a preset cross-modal attention interaction network; A preset cross-modal attention interaction network calculates the attention scores between visual features, target graph structure, environmental data features, and temporal data features; Based on the attention scores among the visual features, the target graph structure, the environmental data features and the temporal data features, the visual features, the target graph structure, the environmental data features and the temporal data features are fused to generate the target features.
[0020] The driving decision determination method provided by the embodiments of the present application encodes the environmental data and time data to generate environmental data features and time data features, so as to convert these data of different types and formats into a unified feature representation that can be processed by the model. The visual features, target graph structure, environmental data features and time data features are input into a preset cross-modal attention interaction network to achieve the integration of multi-source information. The preset cross-modal attention interaction network calculates the attention scores between the visual features, target graph structure, environmental data features and time data features, enabling automatic focusing on the information that is most important for understanding and making decisions in the current driving scenario. Based on the attention scores between the visual features, target graph structure, environmental data features and time data features, the visual features, target graph structure, environmental data features and time data features are fused to generate target features. The generated target features can describe the driving scenario more comprehensively and accurately. The fused target features integrate the information advantages of different modalities, including both the appearance and relationship information of objects, and also consider the influence of environmental and time factors, providing a high-quality feature representation for subsequent determination of the driving scenario complexity and making driving decisions, which helps to improve the accuracy and reliability of the decisions.
[0021] In an alternative embodiment, the first driving behavior probability density, the second driving behavior probability density, and the current driving scenario complexity are input into a preset weight determination model, and the weight information corresponding to the preset driving decision specific model and the preset driving decision large language model is output, including: Calculate the difference metric between the first driving behavior probability density and the second driving behavior probability density; Concatenate the difference metric, the first driving behavior probability density, the second driving behavior probability density, and the current driving scenario complexity to generate a concatenated feature; Obtain the current driving decision output accuracies corresponding to the current driving scenario type, the preset driving decision specific model, and the preset driving decision large language model respectively; Generate dynamic mapping parameters based on the current driving scenario type and the current driving decision output accuracies corresponding to the preset driving decision specific model and the preset driving decision large language model respectively; Perform mapping processing on the concatenated feature based on the dynamic mapping parameters to generate a low-dimensional feature; Input the low-dimensional feature into a preset hidden layer to perform feature processing on the low-dimensional feature; The output layer in the preset weight determination model outputs the weight information corresponding to the preset driving decision specific model and the preset driving decision large language model respectively.
[0022] The driving decision determination method provided in the embodiment of the present application calculates the difference measure between the first driving behavior probability density and the second driving behavior probability density; by calculating the difference measure, the difference between the first driving behavior probability density output by the preset driving decision-making dedicated model and the second driving behavior probability density output by the preset driving decision-making large language model can be clearly understood. This helps to find the difference between the two models in the understanding of the current driving scene and the decision-making tendency, and provides a basis for the subsequent comprehensive consideration of the output of the two models. The difference measure, the first driving behavior probability density, the second driving behavior probability density and the complexity of the current driving scene are spliced to generate splicing features. These multi-dimensional information can more comprehensively describe the relevant factors of the current driving decision and provide a more sufficient data basis for the analysis of subsequent models. The current driving scene type, the preset driving decision-making dedicated model and the preset driving decision-making large language model are obtained. The current driving decision output accuracy corresponding to each other is obtained; based on the current driving scene type and the current driving decision output accuracy corresponding to the preset driving decision-making dedicated model and the preset driving decision-making large language model, dynamic mapping parameters are generated. It enables dynamic adaptation to different driving scenes. The performance of the model is different in different scenes. By dynamically generating mapping parameters, the processing method of splicing features can be optimized for each specific scene. Based on the dynamic mapping parameters, the concatenated features are mapped to generate low-dimensional features, which effectively reduces the dimensional complexity of the data. The low-dimensional features are input into the preset hidden layer and feature processed to dig out deeper information and patterns in the low-dimensional features. The output layer in the preset weight determination model outputs the weight information corresponding to the preset driving decision-making dedicated model and the preset driving decision-making large language model. The reasonable weight allocation of the two models in the current driving scenario is achieved, which improves the quality of driving decisions.
[0023] In a second aspect, the present invention provides a driving decision determination device, the device comprising: An acquisition module is used to acquire current driving scene data corresponding to the vehicle, the current driving scene data including at least one of the video data in front of the vehicle, the vehicle operation information, the movement information of surrounding vehicles, the surrounding road data, the environment data and the time data; An input module, used to input current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model; The determination module is used to determine the target driving decision corresponding to the vehicle based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model.
[0024] The driving decision-making determination device provided by the embodiment of the present application acquires the current driving scenario data corresponding to the vehicle itself. The current driving scenario data includes at least one of the video data in front of the vehicle itself, the running information of the vehicle itself, the movement information of surrounding vehicles, the surrounding road data, the environmental data, and the time data. By collecting the current driving scenario data, it provides a rich and accurate information basis for subsequent decisions. The current driving scenario data is input into a preset dedicated driving decision model and a preset large language model for driving decisions. The respective advantages of these two models can be fully utilized. The dedicated model is usually optimized for specific driving scenarios and tasks and is good at handling some common and patterned driving situations, and can quickly and accurately give corresponding decision suggestions. The large language model has strong language understanding and knowledge reasoning capabilities, can handle complex semantic information and rare driving scenarios, and has good processing capabilities for some situations that require comprehensive judgment and knowledge application. The combined use of the two models enables the vehicle to have a more flexible decision-making method when facing various different types of driving scenarios. Based on the first driving behavior probability density output by the preset dedicated driving decision model and the second driving behavior probability density output by the preset large language model for driving decisions, the target driving decision corresponding to the vehicle itself is determined. This way of fusing decisions can integrate the results of the two models and avoid the limitations that may exist in a single model. By analyzing and fusing the two probability densities, the possibility of various driving behaviors can be more accurately evaluated, so as to obtain a target driving decision that more conforms to the actual driving scenario, improving the accuracy and reliability of the decision.
[0025] In a third aspect, the present invention provides an electronic device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the driving decision-making determination method according to the first aspect or any corresponding implementation manner thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the specific implementation manners of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific implementation manners or the prior art. Obviously, the following drawings are some implementation manners of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a flowchart of the driving decision-making determination method according to the embodiment of the present invention; Figure 2 It is a flowchart of another driving decision-making determination method according to the embodiment of the present invention; Figure 3 It is a timing framework diagram of the driving decision-making determination method according to the embodiment of the present invention; Figure 4 is a structural block diagram of a driving decision determination device according to an embodiment of the present invention; Figure 5 is a schematic diagram of the hardware structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners
[0028] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that for the method for determining a driving decision provided in the embodiments of the present application, the execution subject may be a device for determining a driving decision. The device for determining a driving decision may be implemented as part or all of an electronic device through software, hardware, or a combination of software and hardware. Among them, the electronic device may be a control unit in the vehicle of the present vehicle. In the following method embodiments, the execution subject is taken as an electronic device for illustration.
[0030] According to an embodiment of the present invention, an embodiment of a method for determining a driving decision is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings may be executed in a computer system such as a set of computer executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0031] In this embodiment, a method for determining a driving decision is provided, which can be used for the above-mentioned electronic device. Figure 1 is a flowchart of a method for determining a driving decision according to an embodiment of the present invention. As Figure 1 shown, the process includes the following steps: Step S101, obtain the current driving scenario data corresponding to the vehicle of the present vehicle.
[0032] Among them, the current driving scenario data includes at least one of the video data in front of the vehicle itself, the running information of the vehicle itself, the movement information of surrounding vehicles, the surrounding road data, the environmental data, and the time data.
[0033] Specifically, the electronic device can collect video data in front of the vehicle based on a camera installed in front of the vehicle. The electronic device obtains the running information of its own vehicle through the vehicle's own sensors, such as a speed sensor, an acceleration sensor, a steering wheel angle sensor, a gyroscope, etc. to obtain the corresponding running information of its own vehicle. The electronic device mainly relies on on-vehicle radars (such as millimeter-wave radars, lidar) and vehicle networking technology to obtain the movement information of surrounding vehicles. For example, a millimeter-wave radar can measure information such as the distance, relative speed, and angle between surrounding vehicles and the vehicle itself, with high accuracy and real-time performance; a lidar can generate three-dimensional point cloud data of the surrounding environment, more accurately depicting the position and shape of surrounding vehicles. Vehicle networking technology can obtain the driving state information of surrounding vehicles, such as vehicle speed, driving direction, braking state, etc., through communication between vehicles (such as V2V communication).
[0034] The electronic device can obtain surrounding road data through devices such as an in-vehicle map system, road sensors (such as geomagnetic sensors, road profile sensors), and cameras. Exemplarily, the in-vehicle map system provides basic information about the road, such as the road grade, number of lanes, speed limit signs, intersection information, etc. Road sensors can monitor the road conditions in real time, such as the flatness of the road surface and the presence of obstacles. Cameras can identify information such as road markings and lane lines on the road through image recognition technology.
[0035] In addition, the electronic device can also collect through environmental sensors installed on the vehicle, such as temperature sensors, humidity sensors, light sensors, rain sensors, etc. In addition, more comprehensive environmental information can be obtained by interacting with an external environmental monitoring system (such as a weather station). The electronic device mainly obtains time information through the vehicle's clock system, and can combine with the navigation system to obtain the specific date and time.
[0036] Step S102, input the current driving scenario data into a preset dedicated driving decision-making model and a preset large language model for driving decision-making.
[0037] Specifically, a preset dedicated driving decision-making model is usually specifically designed for driving decision-making tasks, and its architecture may be based on traditional machine learning algorithms (such as decision trees, support vector machines) or deep learning models (such as convolutional neural networks, recurrent neural networks, etc.). Taking deep learning models as an example, a preset dedicated driving decision-making model often takes into account the characteristics of driving scenario data. For the video data in front of the host vehicle, a convolutional neural network (CNN) is used to extract visual features in the image, such as the features of objects like vehicles, pedestrians, road signs, etc. Through multiple layers of convolution and pooling operations, CNN can automatically learn different levels of feature representations in the image, from simple edge and texture information to complex object shape and structure information. For numerical data such as the running information of the host vehicle and the movement information of surrounding vehicles, it is processed through fully connected layers, and these data are fused with visual features to comprehensively understand the driving scenario. After the calculation and processing of the model, the preset dedicated driving decision-making model outputs the probability density of the first driving behavior. This probability density represents the probability distribution of various possible driving behaviors (such as accelerating, decelerating, changing lanes to the left, changing lanes to the right, maintaining the current state, etc.) in the current driving scenario. For example, the output probability density may indicate that in the current scenario, the probability of accelerating is 0.1, the probability of decelerating is 0.3, the probability of changing lanes to the left is 0.1, the probability of changing lanes to the right is 0.1, the probability of maintaining the current state is 0.4, etc. The probability density output by the preset dedicated driving decision-making model is based on its learning and understanding of the input data, reflecting the model's judgment and prediction of the current driving scenario.
[0038] The preset large language model for driving decision-making is improved and optimized based on large language models (such as GPT series, BERT, etc.). In the application of driving decision-making, the preset large language model for driving decision-making needs to convert driving scenario data into language-form inputs. For numerical data, it needs to be converted into appropriate text descriptions. For example, the vehicle speed of "60 kilometers per hour" is described as "the current vehicle speed is sixty kilometers per hour". For image data, it may be necessary to use an image description generation algorithm to convert the key information in the image into text statements. During the data conversion process, it is necessary to ensure the accuracy and integrity of the information so that the model can correctly understand the driving scenario. In addition, operations such as word segmentation and word embedding may also be required on the input text to convert the text into a vector form that the model can process.
[0039] The preset driving decision large language model outputs the second driving behavior probability density based on the understanding and reasoning of the input text. Similar to the preset driving decision-specific model, this probability density represents the probability distribution of various driving behaviors in the current driving scenario. With its powerful language understanding and reasoning capabilities, the large language model can process some complex driving scene descriptions and semantic information, and output a more logical and reasonable driving behavior probability density. For example, when encountering a complex description of a traffic sign or traffic rules, the large language model can accurately determine the driving behavior that should be taken based on its understanding of the language, and output the corresponding probability density.
[0040] Step S103, determining a target driving decision corresponding to the vehicle based on a first driving behavior probability density output by a preset driving decision-specific model and a second driving behavior probability density output by a preset driving decision large language model.
[0041] Specifically, the electronic device can determine the weight information corresponding to the preset driving decision-specific model and the preset driving decision-making large language model. Then, the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision-making large language model are weighted and calculated. The specific weighted calculation can be a weighted sum of the probabilities of each driving behavior. For example, for acceleration behavior, the fused probability = the probability of acceleration in the first driving behavior probability density × the weight of the dedicated model + the probability of acceleration in the second driving behavior probability density × the weight of the large language model.
[0042] After the new driving behavior probability density is obtained through fusion, the driving behavior with the highest probability is usually selected as the target driving decision. For example, in the fused probability density, the probability of maintaining the current state is 0.6, which is the highest probability among all driving behaviors. Then the target driving decision is to maintain the current driving state. This method is simple and direct, can make decisions quickly, and is suitable for most common driving scenarios.
[0043] The driving decision-making determination method provided by the embodiments of the present application obtains the current driving scenario data corresponding to the vehicle of this vehicle. The current driving scenario data includes at least one of the video data in front of the vehicle itself, the running information of the vehicle itself, the movement information of surrounding vehicles, the surrounding road data, the environmental data, and the time data. By collecting the current driving scenario data, the vehicle of this vehicle can comprehensively perceive the driving environment it is in, including the road conditions ahead, its own driving state, the dynamics of surrounding vehicles, road conditions, environmental factors, and time factors, etc., providing a rich and accurate information basis for subsequent decision-making. Input the current driving scenario data into a preset dedicated driving decision-making model and a preset large language model for driving decision-making. The respective advantages of these two models can be fully utilized. The dedicated model is usually optimized for specific driving scenarios and tasks, and is good at dealing with some common and patterned driving situations, and can quickly and accurately give corresponding decision-making suggestions. The large language model has strong language understanding and knowledge reasoning capabilities, can handle complex semantic information and rare driving scenarios, and has good processing capabilities for some situations that require comprehensive judgment and knowledge application. The combined use of the two models enables the vehicle to have a more flexible decision-making method when facing various different types of driving scenarios. Whether it is a simple daily driving scenario or a complex and special situation, appropriate driving decisions can be generated through the collaborative work of these two models, improving the vehicle's ability to cope with various complex environments. Based on the first driving behavior probability density output by the preset dedicated driving decision-making model and the second driving behavior probability density output by the preset large language model for driving decision-making, determine the target driving decision corresponding to the vehicle of this vehicle. This way of fusing decisions can integrate the results of the two models and avoid the limitations that may exist in a single model. By analyzing and fusing the two probability densities, the possibility of various driving behaviors can be more accurately evaluated, so as to obtain a target driving decision that more conforms to the actual driving scenario, improving the accuracy and reliability of the decision-making.
[0044] In this embodiment, a driving decision-making determination method is provided, which can be used for the above-mentioned electronic device. Figure 2 It is a flowchart of the driving decision-making determination method according to the embodiments of the present invention, as Figure 2 shown, and this process includes the following steps: Step S201, obtain the current driving scenario data corresponding to the vehicle of this vehicle.
[0045] Among them, the current driving scenario data includes at least one of the video data in front of the vehicle itself, the running information of the vehicle itself, the movement information of surrounding vehicles, the surrounding road data, the environmental data, and the time data.
[0046] For this step, please refer to the introduction of step S101 above, and details will not be repeated here.
[0047] Step S202: Input the current driving scenario data into a preset dedicated driving decision-making model and a preset large language model for driving decision-making.
[0048] Among them, the training process of the preset large language model for driving decision-making may include the following steps: Step a1: Obtain the original large language model.
[0049] Specifically, the electronic device can receive the original large language model input by the user or the original large language model sent by other devices. The embodiments of the present application do not specifically limit the manner in which the electronic device obtains the original large language model.
[0050] Step a2: Decompose the weight matrix of the original large language model into at least two low-rank matrices.
[0051] Specifically, the electronic device uses the low-rank decomposition method to decompose the weight matrix of the original large language model into two low-rank matrices.
[0052] Among them, the low-rank decomposition method can be any one of low-rank decomposition methods such as the singular value decomposition method and the non-negative matrix decomposition method.
[0053] Step a3: Freeze the original parameters corresponding to the original large language model.
[0054] Specifically, the electronic device can freeze the original parameters corresponding to the original large language model by setting the requires_grad attribute (in PyTorch) or the trainable attribute (in TensorFlow) of the parameters.
[0055] Step a4: Obtain a preset driving decision-making expert library.
[0056] Among them, the preset driving decision-making expert library includes expert experience, industry rules, and historical cases.
[0057] Specifically, the electronic device can receive the preset driving decision-making expert library input by the user or the preset driving decision-making expert library sent by other devices. The electronic device can also search for the preset driving decision-making expert library in the storage space.
[0058] Step a5: Based on the preset driving decision-making expert library, update the parameters of each low-rank matrix to obtain the updated parameters corresponding to each low-rank matrix respectively.
[0059] Specifically, the above step a5 may include the following steps: Step a51: Identify the expert experience, industry rules, and historical cases in the preset driving decision-making expert library and construct a knowledge graph.
[0060] Specifically, the electronic device conducts in-depth analysis of the textual knowledge in the preset driving decision expert database. Using natural language processing technology, textual content such as expert experience, industry rules, and historical cases are converted into structured knowledge representation, and key entities are identified from the knowledge representation. Such as "driver", "vehicle", "road", "traffic sign", "weather", etc. For example, in the sentence "The driver drove at 80 kilometers per hour on the highway on a rainy day", "rainy day" is the weather entity, "driver" is the person entity, "highway" is the road entity, and "80 kilometers per hour" is the speed entity.
[0061] Then, the electronic device determines the relationship between entities, such as "the driver operates the vehicle", "the vehicle is traveling on the road", "the traffic sign indicates the driver's behavior", etc. Taking "the red light ahead, the driver steps on the brake" as an example, the "indication" relationship between "red light (traffic sign)" and "driver" and the "execution" relationship between "driver" and "brake (operation behavior)" can be extracted. Next, the electronic device assigns corresponding attributes to the entity, such as the color, model, and speed of the vehicle, the age and driving experience of the driver, the grade and width of the road, etc. For example, for "a black car is traveling at a speed of 50 kilometers per hour on a city road", "black" can be extracted as the color attribute of the vehicle, "car" as the model attribute of the vehicle, "city road" as the type attribute of the road, and "50 kilometers per hour" as the speed attribute of the vehicle.
[0062] Then, the electronic device uses the extracted entities as nodes of the knowledge graph, and each node represents a specific concept or object. For example, "driver", "vehicle", "road", "traffic sign", etc. all exist as independent nodes in the knowledge graph. Edges are created based on the relationship between entities, and the direction and type of the edge represent the direction and nature of the relationship. For example, an "indication" edge is created from the "traffic sign" node to the "driver" node, indicating the indication relationship of the traffic sign to the driver; an "operation" edge is created from the "driver" node to the "vehicle" node, indicating the operation relationship of the driver to the vehicle.
[0063] Finally, the electronic device adds the extracted attributes to the corresponding nodes to enrich the node information and generate a knowledge graph. For example, attributes such as color, model, and speed are added to the "vehicle" node, and attributes such as age and driving experience are added to the "driver" node. In this way, each node in the knowledge graph has detailed description information and can more accurately represent the knowledge related to driving decisions.
[0064] Step a52, input the knowledge graph into the original large language model.
[0065] Specifically, the electronic device converts the knowledge graph into a graph vector representation, and then inputs the graph vector representation into the original large language model.
[0066] Step a53, the original large language model identifies the knowledge graph and outputs corresponding simulated driving decisions in various simulated driving environments.
[0067] Specifically, the original large language model, such as a model based on the Transformer architecture, processes the input data through multiple layers of self-attention mechanisms and feed-forward neural networks. When processing inputs related to the knowledge graph, the self-attention mechanism can capture the relationships between different parts of the input data, thereby understanding the semantic information in the knowledge graph. For example, the original large language model can, through the self-attention mechanism, pay attention to the relationship between "traffic signs" and "driver behavior" in the knowledge graph, as well as their importance in the current simulated driving environment. The multi-layer structure of the original large language model enables it to extract features of the input data layer by layer, from simple word semantic understanding to complex knowledge reasoning.
[0068] Based on the understanding of the knowledge graph and other input information, the original large language model conducts reasoning and logical judgments. It can infer reasonable driving decisions according to the knowledge such as expert experience, industry rules, and historical cases stored in the knowledge graph, combined with the characteristics of the current simulated driving environment. For example, if the knowledge graph records the expert experience that when driving on urban roads in rainy days, the vehicle speed should be reduced and a safe distance should be maintained from the vehicle in front, when the model identifies that the current simulated driving environment is "driving on urban roads in rainy days", it will reason based on this knowledge and judge the appropriate driving behavior in this situation, such as suggesting that the driver reduce the vehicle speed to an appropriate range and keep a certain safe distance from the vehicle in front.
[0069] Next, the original large language model generates simulated driving decisions based on the reasoning results. This may involve extracting relevant decision-making suggestions from the knowledge graph and adjusting and optimizing them according to the specific situation of the current simulated driving environment. For example, under different vehicle speeds and road conditions, the original large language model will generate specific operation instructions according to the rules and experience in the knowledge graph, such as "the current vehicle speed is 60 km / h, approaching an intersection ahead, please reduce the speed to 30 km / h and pay attention to observing pedestrians". The original large language model may also consider the comprehensive influence of multiple factors, such as weather, traffic flow, vehicle status, etc., to generate comprehensive driving decisions.
[0070] Step a54, based on a preset reward function, reward evaluation is carried out on each simulated driving decision to obtain reward feedback data.
[0071] The preset reward function includes safety rewards, efficiency rewards, comfort rewards, and social compatibility rewards.
[0072] In an optional implementation manner of the present application, the formula of the preset reward function is: ; where N is the number of simulation times, is the safety reward, is the efficiency reward, is the comfort reward, is the social compatibility reward; ; where C is the number of collisions, α is the safety penalty coefficient, and α > 0; for example, each time a collision occurs, α points are deducted. If no collision occurs, C = 0, then R safety = 0.
[0073] ; where T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency reward coefficient, and β > 0; for example, when the actual completion time is equal to the expected time, the efficiency reward is 0; when the actual completion time is less than the expected time, the reward is positive, otherwise it is negative.
[0074] ; where the smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration change rates, denoted as ; n is the number of time intervals during driving, Δa i is the acceleration change amount in the i-th time interval, and the reasonable following distance is maintained by measuring the deviation between the actual following distance d actual and the ideal following distance d ideal , denoted as , m is the number of times related to following situations; γ is the comfort reward coefficient, S a and S d are the normalization constants of the sum of the squares of the acceleration change rates and the sum of the squares of the following distance deviations respectively, and γ > 0; when the acceleration change is smooth and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative; ; where the traffic flow smoothness can be measured by the average speed change rate of surrounding vehicles, denoted as , when the simulated driving decision promotes traffic flow smoothness, , let the number of times of reasonable and successful lane changes be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ > 0, when the simulated driving decision helps to increase the average speed of surrounding vehicles and the lane changes are reasonably completed, the social compatibility reward is positive.
[0075] Specifically, the electronic device performs reward evaluation on each simulated driving decision based on a preset reward function to obtain reward feedback data.
[0076] Step a55, based on the reward feedback data, the update directions of the low-rank matrix parameters are determined.
[0077] Specifically, to determine the relationship between the reward feedback data and the low-rank matrix parameters, the electronic device can establish an association model. By using machine learning methods, the low-rank matrix parameters can be used as input features, and the reward feedback data can be used as output labels to train a model to learn the mapping relationship between them. For example, a neural network model can be used. The parameter values of the low-rank matrix are input into the network, and after multiple layers of calculations and transformations, the predicted reward scores are output. The parameters of the model are adjusted through a large amount of training data so that the model can predict the relationship between the reward scores and the low-rank matrix parameters as accurately as possible.
[0078] After establishing the association model, the feature importance of the low-rank matrix parameters is analyzed. Some feature selection methods, such as gradient-based methods and information gain-based methods, can be used to determine which low-rank matrix parameters have a greater impact on the reward feedback data. For example, by calculating the magnitude of the gradient of each parameter, the larger the gradient, the more significant the impact of the parameter on the reward score. In this way, the key parameters that have an important impact on the reward feedback data can be found, providing a basis for determining the parameter update direction. Then, the electronic device can adopt a gradient-based optimization method to determine the parameter update direction according to the gradient calculated by the association model between the reward feedback data and the low-rank matrix parameters. For the parameters that increase the reward score, increase their values (along the direction of gradient ascent); for the parameters that decrease the reward score, decrease their values (along the direction of gradient descent). For example, if the gradient of a certain low-rank matrix parameter is positive, it means that increasing the value of this parameter may increase the reward score, so the value is appropriately increased when updating the parameter. By continuously updating the parameters along the gradient direction, the low-rank matrix is gradually optimized to improve the reward score of the simulated driving decision.
[0079] Step a56, based on the update directions, the parameters of each low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix respectively.
[0080] Specifically, the electronic device determines the update direction and amplitude of the parameters by calculating the gradient of the reward feedback data with respect to the low-rank matrix parameters. Then, the parameters of the low-rank matrix are updated to obtain the updated parameters corresponding to each low-rank matrix. In the gradient-based update method, the learning rate η is a key factor. If the learning rate is too large, the step size of parameter update will be too large, which may cause the parameters to skip the optimal value and even make the model non-convergent; if the learning rate is too small, the parameter update speed will be very slow and the training time will be very long. Therefore, it is necessary to dynamically adjust the learning rate according to the training process. A common method is learning rate decay, where the learning rate is gradually decreased as the training progresses. For example, the initial learning rate is set to 0.1, and after a certain number of training steps (such as 1000 steps), the learning rate is multiplied by a decay factor (such as 0.9), that is, the learning rate becomes 0.1×0.9. This can quickly adjust the parameters in the initial stage of training, reduce the step size when approaching the optimal value, and improve the accuracy of parameter update.
[0081] Step a6, fuse the original parameters with each updated parameter to obtain the target parameters.
[0082] Specifically, the electronic device can obtain the weight information corresponding to the original parameters and each updated parameter, and then, based on the weight information corresponding to the original parameters and each updated parameter, fuse the original parameters with each updated parameter to obtain the target parameters.
[0083] Step a7, obtain a preset driving decision large language model based on the target parameters.
[0084] Specifically, the electronic device generates a preset driving decision large language model based on the target parameters.
[0085] Step S203, determine the target driving decision corresponding to the vehicle based on the first driving behavior probability density output by the preset driving decision dedicated model and the second driving behavior probability density output by the preset driving decision large language model.
[0086] Specifically, the above step S203 may include the following steps: Step S2031, input the current driving scene data into the preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle.
[0087] Specifically, the above step S2031 may include the following steps: Step b1, extract features from the front vehicle video data in the current driving scene data to extract visual features of multiple objects.
[0088] Among them, the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles.
[0089] Specifically, the electronic device can use a deep learning model to detect and locate at least one of vehicles, pedestrians, road facilities, and obstacles in the video data in front of the vehicle. The deep learning model can be Faster R-CNN, the YOLO (You Only Look Once) series, SSD (Single Shot MultiBox Detector), etc. Then, for the detected objects, the electronic device can use a CNN to extract their visual features. The CNN consists of multiple convolutional layers, pooling layers, and fully connected layers. The convolutional layer slides a convolutional kernel over the image to extract local features of the image, such as edges and textures; the pooling layer downsamples the output of the convolutional layer to reduce the amount of data while retaining important feature information; the fully connected layer integrates the extracted features and outputs the feature vector of the object. For example, based on classic CNN models such as ResNet and VGG, feature extraction can be performed on the image regions of detected objects such as vehicles and pedestrians, so as to obtain the visual features of various objects. After being pre-trained on a large-scale image dataset, these models can learn rich visual feature representations and can effectively extract features such as the shape, color, and texture of objects.
[0090] Step b2: Identify the running information of the host vehicle, the movement information of surrounding vehicles, and the surrounding road data in the current driving scenario data. Using vehicles, pedestrians, road facilities, and obstacles as nodes, and the spatial position relationships and interactions between vehicles, pedestrians, road facilities, and obstacles as edges, construct an initial graph structure.
[0091] Specifically, the electronic device identifies the running information of the host vehicle, the movement information of surrounding vehicles, and the surrounding road data in the current driving scenario data. Both the host vehicle and surrounding vehicles are regarded as nodes in the graph structure. Each vehicle node contains relevant information about the vehicle, such as the type of the vehicle (sedan, truck, bus, etc.), speed, position, driving direction, etc. Pedestrians, as important participants on the road, are also determined as nodes in the graph structure. The pedestrian node contains information such as the position, speed, walking direction, and posture of the pedestrian. Road facilities include traffic lights, traffic signs, street lights, guardrails, etc. These facilities play an important guiding and safety guarantee role in the traffic system. Regarding them as nodes can clearly show the relationships between road facilities and vehicles and pedestrians. Obstacles on the road, such as fences in construction areas and fallen objects, are also determined as nodes. The obstacle node contains information such as the position, size, and shape of the obstacle.
[0092] Then, based on the spatial position information among vehicles, pedestrians, road facilities, and obstacles, edges representing the spatial position relationships are constructed. For example, if there is a certain distance and angle relationship between a vehicle and a traffic signal ahead, an edge can be established between the vehicle node and the traffic signal node, and the attributes of the edge can include information such as the distance and angle between the two. Similarly, the spatial position relationships between a vehicle and a pedestrian, and between a vehicle and an obstacle can also be represented by edges. These edges can intuitively display the relative position relationships among the nodes in the graph structure, providing a basis for subsequent path planning and obstacle avoidance decisions. In addition to the spatial position relationships, considering the interactions among vehicles, pedestrians, road facilities, and obstacles, edges representing the interactions are constructed. For example, when the traffic signal turns green, the vehicle will start or accelerate according to the signal indication, and this control relationship of the signal on the vehicle can be reflected by establishing an edge representing the interaction between the signal node and the vehicle node. Another example is that when a pedestrian crosses the road, the vehicle needs to decelerate or stop to give way, and this interaction between the vehicle and the pedestrian can also be represented by an edge. These interaction edges can reflect the dynamic relationships among the nodes in the graph structure, helping to analyze and predict behaviors and events in traffic scenarios.
[0093] Through the above steps, vehicles, pedestrians, road facilities, and obstacles are determined as nodes, and edges are constructed based on their spatial position relationships and interactions, finally forming an initial graph structure.
[0094] Step b3: Identify the video data in front of the ego vehicle in the current driving scenario data to determine the current driving scenario type.
[0095] Specifically, identify the video data in front of the ego vehicle in the current driving scenario data to determine the current driving scenario type. Among them, the driving scenario type can be a highway driving scenario, an urban road driving scenario, a school area driving scenario, a rural road driving scenario, etc. Among them, different current driving scenario types have different levels of attention.
[0096] Step b4: Based on the current driving scenario type, add semantic information to at least one edge in the initial graph structure to generate a target graph structure.
[0097] Specifically, the electronic device determines the content of the semantic information to be added according to the current driving scenario type and the characteristics of the edge. For example, for the edge between a vehicle and a pedestrian, semantic information such as "Pedestrians may suddenly cross the road, slow down and give way" can be added; for the edge between a vehicle and a traffic signal, semantic information such as "Stop and wait at a red light and proceed only when the light is green" can be added; for the edge between a vehicle and a vehicle in an adjacent lane, semantic information such as "Maintain a safe distance and appropriate speed when overtaking" can be added, etc. These semantic information should have clear guiding significance and can help understand and process various relationships in the driving scenario.
[0098] Then, the determined semantic information is added to the edge of the initial graph structure in a suitable way. A semantic information field can be added to the attributes of the edge to store the corresponding semantic description; or in a graphical way, the semantic information can be marked beside or near the edge to make it more intuitively displayed. For example, use a text label to indicate the semantic information beside the edge, or add a special "Semantic Description" item to the attribute list of the edge to record the detailed semantic content.
[0099] Step b5, fuse the visual features, the target graph structure, the environmental data, and the time data to generate the target features.
[0100] Specifically, the above step b5 may include the following steps: Step b51, perform feature encoding on the environmental data and the time data to generate environmental data features and time data features.
[0101] Specifically, for the environmental data, for some classified environmental data, such as weather conditions, one-hot encoding or label encoding is often used. For numerical environmental data, such as temperature, humidity, light intensity, etc., normalization or standardization processing may be required. To more comprehensively represent the characteristics of the environmental data, feature combination and extraction can also be performed. For example, combine temperature and humidity into a new feature "humid heat index", or extract the main components of the data through methods such as principal component analysis (PCA) to reduce the data dimension while retaining most of the information.
[0102] For time data, time data generally includes information such as specific time points (e.g., year, month, day, hour, minute, second) or time intervals. First, the time data needs to be converted into a suitable format for encoding. For example, date and time data is converted into a timestamp (the number of seconds or milliseconds from a fixed starting time point to the current time), which facilitates subsequent numerical calculations. Time data has obvious periodicity, such as 24 hours in a day, 7 days in a week, 12 months in a year, etc. These periodic features can be extracted. For example, the time of a day can be divided into different time periods (morning, noon, afternoon, evening) and represented using one-hot encoding; or the number of days in a week can be encoded numerically, with Monday being 1, Tuesday being 2, and so on. By extracting periodic features, the model can better capture time-related patterns.
[0103] Integrate the environmental data features and time data features processed as above to form the final feature vector. Different types of features can be arranged in a certain order. For example, environmental data features are arranged first, followed by time data features, to form a complete feature vector as the input for the subsequent model.
[0104] Step b52: Input the visual features, target graph structure, environmental data features, and time data features into a preset cross-modal attention interaction network.
[0105] Specifically, the electronic device inputs the visual features, target graph structure, environmental data features, and time data features into a preset cross-modal attention interaction network.
[0106] Step b53: The preset cross-modal attention interaction network calculates the attention scores between the visual features, target graph structure, environmental data features, and time data features.
[0107] Specifically, to enable the interaction of features in different modalities in the same space, the preset cross-modal attention interaction network performs a linear transformation on the features of each modality. By multiplying by the corresponding weight matrix, the features of each modality are mapped into a new feature space. Then, the attention scores between the visual features, target graph structure, environmental data features, and time data features are calculated.
[0108] Taking the calculation of the attention score between visual features and features of other modalities as an example, the dot product or other similarity measurement methods are usually used. Suppose we want to calculate the attention score V ′ between the visual feature G ′ and the target graph structure A VG . It can be calculated through the dot product operation A VG = ( V ′) T G′to initially obtain a similarity value. To make the attention score more interpretable and numerically stable, a scaling operation may be performed on it. Similarly, the attention scores A VE and A VT can be calculated between the visual features and the environmental data features, the time data features, as well as the attention scores between other modal features, such as A GE 、A GT 、A ET and so on.
[0109] Step b54, based on the attention scores between the visual features, the target map structure, the environmental data features, and the time data features, fuse the visual features, the target map structure, the environmental data features, and the time data features to generate target features.
[0110] Specifically, multiply the visual features, the target map structure, the environmental data features, and the time data features by their corresponding attention scores and then add them up to obtain the target features.
[0111] Step b6, based on the target features, determine the complexity of the current driving scenario corresponding to the vehicle itself.
[0112] Specifically, a preset scenario complexity determination model outputs the complexity of the current driving scenario corresponding to the vehicle itself based on the target features.
[0113] Step S2032, input the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scenario into a preset weight determination model, and output the weight information corresponding to the preset driving decision dedicated model and the preset driving decision large language model respectively.
[0114] Specifically, the above step S2032 may include the following steps: Step c1, calculate the difference metric between the first driving behavior probability density and the second driving behavior probability density.
[0115] Specifically, the electronic device can calculate the difference metric between the first driving behavior probability density and the second driving behavior probability density based on a preset metric method.
[0116] Among them, the preset metric method can be methods such as KL divergence metric method, JS divergence metric method, Wasserstein distance (earth mover's distance) metric method, Euclidean distance, etc. The embodiments of the present application do not make specific limitations on the preset metric method.
[0117] Step c2, splice the difference metric, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scenario to generate a spliced feature.
[0118] Specifically, the electronic device can splice the difference metric, the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity to generate a spliced feature.
[0119] Step c3: Obtain the current driving decision output accuracies corresponding to the current driving scene type, the preset driving decision specific model, and the preset driving decision large language model respectively.
[0120] Specifically, the electronic device can identify the video data in front of the vehicle in the current driving scene data to determine the current driving scene type. In addition, the electronic device can also determine the current driving decision output accuracies corresponding to the preset driving decision specific model and the preset driving decision large language model respectively based on the historical driving decision records.
[0121] Step c4: Generate dynamic mapping parameters based on the current driving scene type and the current driving decision output accuracies corresponding to the preset driving decision specific model and the preset driving decision large language model respectively.
[0122] Specifically, the electronic device can determine the dynamic mapping relationship between them according to the accuracy differences between the preset driving decision specific model and the preset driving decision large language model under different driving scene types. Specifically, the electronic device can map the accuracies of the preset driving decision specific model and the preset driving decision large language model to a weight value or a parameter value. For example, a linear function, an exponential function, or other suitable mathematical functions can be used to describe this mapping relationship. Suppose the accuracy of the preset driving decision specific model in a certain scene is A1, and the accuracy of the preset driving decision large language model is A2. A mapping function f(A1, A2) can be defined so that the output value of f is between 0 and 1, which is used to represent the relative importance or weight of the preset driving decision specific model and the preset driving decision large language model in this scene.
[0123] Then, the output values of the mapping function under different driving scene types are used as dynamic mapping parameters. These parameters can be used to dynamically combine or adjust the output results of the preset driving decision specific model and the preset driving decision large language model according to the current driving scene type during the actual driving decision process. For example, in a certain scene, if the accuracy of the specific model is high, the dynamic mapping parameters may cause the final driving decision to rely more on the output of the specific model; while in another scene, if the large language model performs better, the parameters will tend to adopt the suggestions of the large language model.
[0124] Step c5: Perform mapping processing on the spliced feature based on the dynamic mapping parameters to generate a low-dimensional feature.
[0125] Specifically, the electronic device can use a preset mapping processing method to map and process the splicing features based on the dynamic mapping parameters to generate low-dimensional features.
[0126] Among them, the preset mapping processing method can be a linear mapping or a non-linear mapping.
[0127] Taking the linear mapping method as an example. Suppose the splicing feature is an n-dimensional vector X = [x 1 , x 2 , ⋯, x n , and the dynamic mapping parameter is an n-dimensional weight vector W = [w 1 , w 2 , ⋯, w n . Through the linear transformation Y = W T X = a low-dimensional scalar or vector Y is obtained. Here, Y is a part of the low-dimensional features after mapping processing. This process can be repeated to generate multiple low-dimensional features through different weight vectors W and combined into the final low-dimensional feature vector.
[0128] For the non-linear mapping method, such as a neural network. The splicing features are used as the input of the neural network, and the dynamic mapping parameters can be used as the parameters of the neural network or affect the adjustment of the parameters. The neural network maps the high-dimensional splicing features to a low-dimensional space through multiple layers of non-linear transformations. For example, a multi-layer perceptron (MLP) can be used. During the training process, the weights of the network are adjusted according to the dynamic mapping parameters, so that the network can learn a feature representation more suitable for the current driving scenario and generate representative low-dimensional features.
[0129] Step c6, input the low-dimensional features into a preset hidden layer to process the low-dimensional features.
[0130] Specifically, when the low-dimensional features are input into the neurons of the preset hidden layer, each neuron will perform a weighted sum on the input signal. Let the input low-dimensional feature vector be x = [x 1 , x 2 , ⋯, x n , the weight matrix between the neuron and the input layer is W = [w ij (where i represents the index of the hidden layer neuron and j represents the index of the input layer neuron), and the bias vector is b = [b 1 , b 2 , ⋯, b m (m is the number of hidden layer neurons), then the input zi of the i-th hidden layer neuron is zi = 。Then, an activation function (such as ReLU, Sigmoid, Tanh, etc.) is applied to zi. The role of the activation function is to introduce non - linear factors, enabling the neural network to learn more complex functional relationships. For example, when using the ReLU activation function, if z i > 0, then the output a i = z i ; if z i ≤ 0, then the output a i = 0.
[0131] If the preset hidden layer contains multiple layers, then the output of the previous hidden layer will be used as the input of the next hidden layer, repeating the above - mentioned process of weight calculation and activation function application. With each passing through a hidden layer, the features will be further abstracted and transformed, gradually forming a more advanced feature representation. For example, the first hidden layer may extract some simple feature combinations, such as the relationship between vehicle speed and the distance to surrounding vehicles; the second hidden layer may, on this basis, combine environmental data features and learn more complex patterns, such as the comprehensive impact of vehicle speed and distance on safe driving under different weather conditions.
[0132] After being processed by the preset hidden layer, the final output is a transformed and abstracted feature representation. These feature representations can be used as the input for subsequent layers (such as the output layer) for predicting or classifying driving decisions. For example, in a driving decision - making model, the features output by the hidden layer can be input into the output layer, and the output layer determines what driving behaviors such as accelerating, decelerating, or steering the vehicle should take based on these features.
[0133] Step c7, determine the output layer in the preset weight determination model, and output the weight information corresponding to the preset driving decision - making dedicated model and the preset driving decision - making large - language model respectively.
[0134] Specifically, the output layer calculates the weight information corresponding to the preset driving decision - making dedicated model and the preset driving decision - making large - language model respectively based on the feature representation output by the hidden layer, and the calculated weight information will be output by the output layer. These weight information reflect the relative importance of the preset driving decision - making dedicated model and the preset driving decision - making large - language model in the final driving decision under the current input driving scenario and model performance conditions. For example, the output weight information may indicate that in the urban congested road scenario, the weight of the preset driving decision - making dedicated model is 0.4, and the weight of the preset driving decision - making large - language model is 0.6, which means that in this scenario, the decision of the large - language model has a greater impact on the final driving decision.
[0135] Step S2033, based on the weight information, perform a fusion process on the first driving behavior probability density and the second driving behavior probability density to determine the target driving decision corresponding to the vehicle of this vehicle.
[0136] Specifically, the electronic device performs a weighted calculation on the first driving behavior probability density output by the preset dedicated driving decision-making model and the second driving behavior probability density output by the preset large language model for driving decision-making based on the weight information. The specific weighted calculation can be a weighted sum of the probabilities of each driving behavior. For example, for the acceleration behavior, the fused probability = the probability of acceleration in the first driving behavior probability density × the weight of the dedicated model + the probability of acceleration in the second driving behavior probability density × the weight of the large language model.
[0137] After obtaining the new driving behavior probability density through fusion, usually the driving behavior with the highest probability is selected as the target driving decision. For example, in the fused probability density, the probability of maintaining the current state is 0.6, which is the highest among all driving behaviors. Then the target driving decision is to maintain the current driving state. This method is simple and direct, can make decisions quickly, and is applicable to most conventional driving scenarios.
[0138] The driving decision-making method provided by the embodiments of this application obtains the original large language model, so that the existing powerful foundation of the original large language model can be utilized, avoiding the large amount of computing resources and time costs required for training the model from scratch. The weight matrix of the original large language model is decomposed into at least two low-rank matrices, which can reduce the complexity of the original large language model and facilitate model optimization and understanding. Then, the original parameters corresponding to the original large language model are frozen, and the general language knowledge and capabilities learned by the original large language model will not be damaged. This pre-trained knowledge is very helpful for understanding and processing natural language information related to driving. In addition, freezing the original parameters helps to stabilize the model training process. A preset driving decision expert library is obtained, introducing professional knowledge and experience to meet actual driving needs. The expert experience, industry rules, and historical cases in the preset driving decision expert library are identified to construct a knowledge graph. This structured knowledge representation is convenient for subsequent model understanding and utilization, improving the accessibility and operability of knowledge. The knowledge graph is input into the original large language model, which can enrich the knowledge reserve of the original large language model and enable it to better understand concepts, rules, and experiences related to driving. For example, the original large language model can understand the meanings of different traffic signs and driving strategies in various road conditions through the knowledge graph, so that when dealing with driving-related tasks, it can reason and make decisions based on more comprehensive knowledge. The original large language model identifies the knowledge graph and outputs simulated driving decisions corresponding to various simulated driving environments. This enables the original large language model to rehearse and analyze various possible driving situations in a virtual environment, covering from simple daily driving scenarios to complex special situations, such as bad weather and traffic accident scenes. Through this simulation, the model can learn and master the best driving decisions in different scenarios in advance, improving its ability to handle actual driving scenarios. The reward evaluation of each simulated driving decision is carried out based on a preset reward function to obtain reward feedback data. The preset reward function evaluates the simulated driving decisions from multiple dimensions such as safety, efficiency, comfort, and social compatibility, and can comprehensively measure the advantages and disadvantages of the decisions. This multi-dimensional evaluation method avoids the problem of only focusing on a single indicator (such as safety) while ignoring other important aspects (such as comfort and social compatibility). Based on the reward feedback data, the update directions of the parameters of each low-rank matrix are determined, making the parameter update targeted and accurate. The model can, according to the reward feedback, focus on adjusting the parameters that have a greater impact on obtaining high rewards, avoiding blindly updating parameters and causing the model performance to decline. The parameters of each low-rank matrix are updated based on the update directions to obtain the updated parameters corresponding to each low-rank matrix. The updated parameters enable the original large language model to better handle driving decision-making tasks, improving the accuracy and reliability of the decisions. For example, the updated model may be able to more accurately judge the driving intentions of other vehicles in complex road conditions, thus making safer and more reasonable driving decisions.Then, the original parameters are fused with each updated parameter to obtain the target parameters, thus enabling the full play of the advantages of both. The original parameters retain the general language understanding ability and extensive knowledge base of the large language model, while the updated parameters endow the model with specific capabilities and expertise for driving decision-making tasks. By fusing these two parts of parameters, the model can not only utilize its general language processing ability to understand and process various natural language information related to driving, but also make accurate driving decisions based on the knowledge in the expert library, achieving an organic combination of general capabilities and professional capabilities. Based on the target parameters, a preset large language model for driving decision-making is obtained, making the preset large language model for driving decision-making a model optimized specifically for driving decision-making tasks. It combines the powerful foundation of the original large language model and the professional driving decision-making capabilities obtained through parameter updates via the expert library, and can efficiently process various information related to driving and generate accurate and reasonable driving decisions. This model can better meet the requirements of real-time performance, accuracy, and reliability for autonomous driving or assisted driving systems, providing strong support for the safe driving of vehicles.
[0139] In addition, feature extraction is performed on the video data in front of the vehicle in the current driving scene data to extract visual features of multiple objects. By extracting visual features of multiple objects such as vehicles, pedestrians, road facilities, obstacles, etc., the system can accurately perceive various objects in the environment in front of the vehicle. The vehicle operation information, surrounding vehicle movement information and surrounding road data in the current driving scene data are identified, and the vehicles, pedestrians, road facilities and obstacles are regarded as nodes. The spatial position relationship and interaction between vehicles, pedestrians, road facilities and obstacles are used as edges to construct an initial graph structure, which can clearly present the relationship between various elements in the current driving scene. This graph structure provides convenience for analyzing complex driving situations. The system can quickly analyze the mutual influence between multiple objects through the graph structure. The video data in front of the vehicle in the current driving scene data is identified to determine the current driving scene type. Different driving scenes have different characteristics and requirements. After clarifying the scene type, various situations can be handled in a targeted manner, thereby improving the accuracy and effectiveness of decision-making. Based on the current driving scene type, semantic information is added to at least one edge in the initial graph structure to generate a target graph structure. The information contained in the graph structure is further enriched, and the target graph structure with semantic information provides a more optimized basis for the system to make decisions. The environmental data and time data are feature encoded to generate environmental data features and time data features, so that these data of different types and formats can be converted into a unified feature representation that can be processed by the model. The visual features, target graph structure, environmental data features and time data features are input into the preset cross-modal attention interaction network to achieve the integration of multi-source information. The preset cross-modal attention interaction network calculates the attention scores between the visual features, target graph structure, environmental data features and time data features, so that it can automatically focus on the most important information for understanding and making decisions in the current driving scene. Based on the attention scores between the visual features, target graph structure, environmental data features and time data features, the visual features, target graph structure, environmental data features and time data features are fused to generate target features. The generated target features can describe the driving scene more comprehensively and accurately. The fused target features combine the information advantages of different modalities, including the appearance and relationship information of objects, and taking into account the influence of environmental and time factors. They provide high-quality feature representation for the subsequent determination of the complexity of the driving scene and making driving decisions, which helps to improve the accuracy and reliability of decision-making. Based on the target features, the complexity of the current driving scene corresponding to the vehicle is determined. In this way, the complexity of the current driving scene can be accurately evaluated.
[0140] Calculate the difference measure between the first driving behavior probability density and the second driving behavior probability density; by calculating the difference measure, the difference between the first driving behavior probability density output by the preset driving decision-making dedicated model and the second driving behavior probability density output by the preset driving decision-making large language model can be clearly understood. This helps to discover the differences between the two models in terms of understanding the current driving scene and decision-making tendencies, and provides a basis for the subsequent comprehensive consideration of the outputs of the two models. The difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene are spliced to generate splicing features. These multi-dimensional information can more comprehensively describe the relevant factors of the current driving decision, and provide a more sufficient data basis for the analysis of subsequent models. Obtain the current driving decision output accuracy corresponding to the current driving scene type, the preset driving decision-making dedicated model, and the preset driving decision-making large language model; based on the current driving scene type and the current driving decision output accuracy corresponding to the preset driving decision-making dedicated model and the preset driving decision-making large language model, generate dynamic mapping parameters. It enables dynamic adaptation to different driving scenes. The performance of the model is different in different scenes. By dynamically generating mapping parameters, the processing method of splicing features can be optimized for each specific scene. Based on the dynamic mapping parameters, the concatenated features are mapped to generate low-dimensional features, which effectively reduces the dimensional complexity of the data. The low-dimensional features are input into the preset hidden layer and feature processed, which can dig out deeper information and patterns in the low-dimensional features. The output layer in the preset weight determination model outputs the weight information corresponding to the preset driving decision-making dedicated model and the preset driving decision-making large language model. The reasonable weight allocation of the two models in the current driving scenario is realized, and the quality of driving decisions is improved. Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are fused to determine the target driving decision corresponding to the vehicle. By fusing the driving behavior probability densities output by the two models through the weight information, the professional knowledge of the preset driving decision-making dedicated model in a specific field and the general understanding and reasoning ability of the preset driving decision-making large language model can be fully utilized. In addition, the fusion processing can reduce the errors and uncertainties that may occur in a single model and improve the reliability and stability of driving decisions.
[0141] For example, in order to better introduce the driving decision determination method provided in the embodiment of the present application, Figure 3 As shown, the embodiment of the present application provides a timing framework diagram of a driving decision determination method, which is described in detail below: 1) Data input section: Visual and environmental data: Multiple data sources are shown on the left, including pictures and semantic parsing of pictures (such as describing the driving conditions of vehicles on the highway, lane occupancy, etc.); front vehicle video data, surrounding vehicle movement information, own vehicle movement information, surrounding road data, environmental data (such as weather, etc.), time data, etc. These data comprehensively describe the driving scenario where the vehicle is located from different perspectives.
[0142] Prompt guidance: Through the prompt "Please give the current decision and the corresponding probability", the model is guided to make a decision based on the input data and give a probability judgment.
[0143] 2) Cloud processing section: Data collection and knowledge update: Various types of data are transmitted to the cloud data collection system, which is responsible for collecting and integrating the data and using it to update the professional knowledge base for lane change decisions. This is the knowledge reserve update link of the entire decision-making system to ensure that the system can make decisions based on the latest data.
[0144] Model fine-tuning: Low-Rank Adaptation (LoRA) algorithm: This algorithm is used to optimize and adjust the model. LoRA efficiently fine-tunes the model in a way of low-rank matrix approximation without changing too many model parameters to adapt to specific lane change decision tasks.
[0145] Fine-tuned large language model - lane change decision: The large language model fine-tuned by the LoRA algorithm is specifically used for lane change decision tasks. It can understand and process various input data information and initially give results such as probability distributions related to lane change decisions.
[0146] 3) Decision analysis section: Scene complexity assessment: The model collaborative decision-making framework assesses the scene complexity of environmental conditions, road conditions, traffic elements, time factors, etc. This step is to comprehensively judge the complexity of the current driving scene and provide a basis for subsequent decisions.
[0147] Dynamic weight update network: According to the results of the scene complexity assessment, the dynamic weight update network adjusts the weights of different models in the decision-making process. For example, in complex scenarios, the weights of certain models that are more proficient in handling complex situations may be increased.
[0148] Dynamic weighted fusion strategy: Using the adjusted weights, the output probability distributions of different models (such as the fine-tuned large language model and traditional decision models) are fused through the dynamic weighted fusion strategy to obtain the final decision probability, thereby determining whether to change lanes to the left, change lanes to the right, or go straight.
[0149] 4) Feedback update part: The decision results will be fed back and updated. On the one hand, it is fed back to the cloud data collection system for further optimizing and updating the professional knowledge base of lane-changing decisions. On the other hand, it is fed back to the model collaborative decision-making framework to continuously adjust the model parameters and decision-making strategies, and continuously improve the accuracy and reliability of lane-changing decisions.
[0150] 5) Traditional decision-making model: The traditional decision-making model is also mentioned in the figure as the research content. Although it is not elaborated in detail, it is a part of the entire decision-making system and works in collaboration with the fine-tuned large language model, etc., to jointly complete the lane-changing decision-making task.
[0151] In this embodiment, a driving decision determination device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" may be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0152] This embodiment provides a driving decision determination device, as Figure 4 shown, including: An acquisition module 301, configured to acquire current driving scenario data corresponding to the vehicle itself, where the current driving scenario data includes at least one of front vehicle video data, own vehicle operation information, surrounding vehicle movement information, surrounding road data, environmental data, and time data; An input module 302, configured to input the current driving scenario data into a preset dedicated driving decision-making model and a preset large language driving decision-making model; A determination module 303, configured to determine a target driving decision corresponding to the vehicle itself based on a first driving behavior probability density output by the preset dedicated driving decision-making model and a second driving behavior probability density output by the preset large language driving decision-making model.
[0153] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0154] The driving decision determination device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0155] This embodiment of the present invention also provides an electronic device having the above Figure 4 shown driving decision determination device.
[0156] Please refer toFigure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by an alternative embodiment of the present invention. As Figure 5 shown, the electronic device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. The electronic device further includes a communication interface 30 for communicating between the electronic device and other devices or communication networks.
[0157] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations fall within the scope defined by the appended claims.
Claims
1. A driving decision determination method, characterized in that: The method comprises: Acquire current driving scene data corresponding to the vehicle, wherein the current driving scene data includes at least one of the video data in front of the vehicle, the vehicle operation information, the motion information of surrounding vehicles, the surrounding road data, the environment data and the time data; Inputting the current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model; Based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model, the target driving decision corresponding to the host vehicle is determined.
2. The method according to claim 1, characterized in that: The training process of the preset driving decision-making large language model includes: Get the original large language model; Decomposing the weight matrix of the original large language model into at least two low-rank matrices; Freezing original parameters corresponding to the original large language model; Obtaining a preset driving decision expert database; the preset driving decision expert database includes expert experience, industry rules, and historical cases; Based on the preset driving decision expert database, the parameters of each of the low-rank matrices are updated to obtain update parameters corresponding to each of the low-rank matrices; The original parameters are combined with the updated parameters to obtain target parameters; Based on the target parameters, the preset driving decision-making large language model is obtained.
3. The method according to claim 2, characterized in that The updating of the parameters of each low-rank matrix based on the preset driving decision expert database to obtain the updated parameters corresponding to each low-rank matrix includes: Identify the expert experience, industry rules, and historical cases in the preset driving decision expert database to construct a knowledge graph; Inputting the knowledge graph into the original large language model; The original large language model recognizes the knowledge graph and outputs corresponding simulated driving decisions under various simulated driving environments; Performing reward evaluation on each of the simulated driving decisions based on a preset reward function to obtain reward feedback data; the preset reward function includes a safety reward, an efficiency reward, a comfort reward, and a social compatibility reward; Based on the reward feedback data, determining the direction of updating each of the low-rank matrix parameters; The parameters of each of the low-rank matrices are updated based on the update direction to obtain the update parameters corresponding to each of the low-rank matrices.
4. The method according to claim 3, characterized in that The preset reward function formula is: ; Where N is the number of simulations, For the safety reward, For the efficiency reward, For the said comfort reward, Reward for said social compatibility; Where C is the number of collisions, α is the safety penalty coefficient, and α>0; ; Among them, T expected is the expected time for the vehicle to complete the task, T actual is the actual completion time, β is the efficiency bonus coefficient, and β>0; ; The smoothness of acceleration and deceleration is measured by the sum of the squares of the acceleration rate of change, denoted as ; n is the number of time intervals during driving, Δa i is the acceleration change in the i-th time interval, and the reasonable following distance is maintained by the actual following distance d actual The ideal following distance d ideal The deviation is measured by , m is the number of times the following vehicle is involved; γ is the comfort reward coefficient, S a and S d are the normalized constants of the sum of squares of the acceleration change rate and the sum of squares of the following distance deviation, and γ>0; when the acceleration changes smoothly and the following distance is reasonable, the comfort reward approaches γ; otherwise, the reward decreases or even becomes negative; ; The degree of traffic flow can be measured by the average speed change rate of surrounding vehicles, denoted as , when simulated driving decisions promote traffic flow, , let the number of lane changes that are reasonable and completed successfully be , δ is the social compatibility reward coefficient, N total-lane-change is the total number of lane changes, and δ >0, when the simulated driving decision helps to improve the average speed of surrounding vehicles and the lane change is completed reasonably, the social compatibility reward is positive.
5. The method according to claim 1, characterized in that The determining of the target driving decision corresponding to the host vehicle based on the first driving behavior probability density output by the preset driving decision-specific model and the second driving behavior probability density output by the preset driving decision large language model includes: Inputting the current driving scene data into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the vehicle; Inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into a preset weight determination model, and outputting weight information corresponding to the preset driving decision-specific model and the preset driving decision large language model respectively; Based on the weight information, the first driving behavior probability density and the second driving behavior probability density are fused to determine a target driving decision corresponding to the host vehicle.
6. The method according to claim 5, characterized in that The inputting the current driving scene data into a preset scene complexity determination model to determine the current driving scene complexity corresponding to the host vehicle includes: Performing feature extraction on the video data in front of the vehicle in the current driving scene data to extract visual features of multiple objects; the multiple objects include at least one of vehicles, pedestrians, road facilities, and obstacles; Identify the vehicle operation information, the surrounding vehicle movement information and the surrounding road data in the current driving scene data, and construct an initial graph structure by taking vehicles, pedestrians, road facilities and obstacles as nodes and the spatial position relationships and interactions between vehicles, pedestrians, road facilities and obstacles as edges; Identify the front video data of the vehicle in the current driving scene data to determine the current driving scene type; Based on the current driving scene type, adding semantic information to at least one edge in the initial graph structure to generate a target graph structure; Fusing the visual features, the target graph structure, the environmental data, and the time data to generate target features; Based on the target feature, determine the complexity of the current driving scene corresponding to the host vehicle.
7. The method according to claim 6, characterized in that The step of fusing the visual features, the target graph structure, the environmental data, and the time data to generate target features includes: Performing feature encoding on the environment data and the time data to generate environment data features and time data features; Inputting the visual features, the target graph structure, the environmental data features and the time data features into a preset cross-modal attention interaction network; The preset cross-modal attention interaction network calculates the attention score between the visual feature, the target graph structure, the environmental data feature and the temporal data feature; Based on the attention scores among the visual features, the target graph structure, the environmental data features and the temporal data features, the visual features, the target graph structure, the environmental data features and the temporal data features are fused to generate the target features.
8. The method according to claim 5, characterized in that The step of inputting the first driving behavior probability density, the second driving behavior probability density, and the current driving scene complexity into a preset weight determination model, and outputting weight information corresponding to the preset driving decision-specific model and the preset driving decision large language model, respectively, includes: Calculating a difference measure between the first driving behavior probability density and the second driving behavior probability density; Splicing the difference measure, the first driving behavior probability density, the second driving behavior probability density, and the complexity of the current driving scene to generate a splicing feature; Obtaining the current driving decision output accuracy corresponding to the current driving scene type, the preset driving decision-specific model, and the preset driving decision large language model respectively; Generate dynamic mapping parameters based on the current driving scene type and the current driving decision output accuracy corresponding to the preset driving decision-specific model and the preset driving decision large language model respectively; Performing mapping processing on the splicing features based on the dynamic mapping parameters to generate low-dimensional features; Inputting the low-dimensional features into a preset hidden layer, and performing feature processing on the low-dimensional features; The output layer in the preset weight determination model outputs weight information corresponding to the preset driving decision-specific model and the preset driving decision large language model respectively.
9. A driving decision determination device, characterized in that: The device comprises: An acquisition module is used to acquire current driving scene data corresponding to the vehicle, wherein the current driving scene data includes at least one of the video data in front of the vehicle, the vehicle operation information, the movement information of surrounding vehicles, the surrounding road data, the environment data and the time data; An input module, used for inputting the current driving scene data into a preset driving decision-specific model and a preset driving decision-making large language model; A determination module is used to determine a target driving decision corresponding to the vehicle based on a first driving behavior probability density output by the preset driving decision-specific model and a second driving behavior probability density output by the preset driving decision large language model.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the driving decision determination method according to any one of claims 1 to 8 by executing the computer instructions.
Citation Information
Patent Citations
Lane change prediction method and device for target vehicle
CN113147766A
Vehicle lane changing prediction method based on attention mechanism
CN114889608A
Intelligent emergency rescue strategy output and emergency rescue vehicle control system and method
CN118505168A
Automatic driving three-dimensional scene data preprocessing method and system based on large language model
CN118781267A
Vehicle driving decision-making method and device, electronic equipment and storage medium
CN119058744A
Cited By
Automatic driving strategy generation method and device and vehicle-mounted large model deployment method and device
CN120671782A
Multi-vehicle collaborative dynamic negotiation scheduling method and related device
CN121011098A
Multi-vehicle cooperative dynamic negotiation scheduling method and related device
CN121011098B