Driving decision method and device, computer device, readable storage medium and program product

By combining environmental information and physiological signal characteristics and utilizing preset decision-making models and risk perception models, the problem of mismatch between autonomous driving decisions and human safety decisions is solved, achieving more accurate and safe autonomous driving decisions.

CN119773812BActive Publication Date: 2025-10-24TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510108104.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-10-24
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

In traditional autonomous driving technology, the vehicle's autonomous driving decision-making results do not match the human-based safety decision-making requirements, resulting in insufficient safety.

Method used

By combining environmental information and physiological signal characteristics, using preset decision models and risk perception models, the driving action decision results are determined, and risk perception parameters are used as judgment conditions to select appropriate driving decisions to match human safety needs.

Benefits of technology

It improves the accuracy and safety of autonomous driving decisions, ensures that vehicles take safety measures in a timely manner in high-risk situations, and protects the safety of people and vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119773812B_ABST
    Figure CN119773812B_ABST
Patent Text Reader

Abstract

The application relates to a driving decision method and device, computer equipment, a readable storage medium and a program product. The method comprises the following steps: determining a first driving action decision result and a second driving action decision result based on environment information and a preset decision model; determining a risk perception parameter based on a preset risk perception model and a physiological signal feature; and determining a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result. The method can match automatic driving decisions with actual vehicle scenes, thereby ensuring the safety of automatic driving.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of automatic driving, in particular to a driving decision method and device, computer equipment, readable storage medium and program product. BACKGROUND

[0002] With the development of automatic driving technology, automatic driving vehicles gradually move towards practicality. Safe decision-making for automatic driving is one of the key technologies of automatic driving vehicles. The automatic driving in the traditional technology often adopts a guided reinforcement learning method. These methods only focus on reinforcement learning decision-making based on vehicle information, and are limited by algorithms and vehicle information, resulting in that the output result of the automatic driving decision of the vehicle does not match the safety decision-making demand based on human factors. SUMMARY

[0003] Therefore, it is necessary to provide a driving decision method, device, computer equipment, readable storage medium and program product to match the automatic driving decision with the safety decision-making based on human factors, so as to ensure the safety of automatic driving.

[0004] In a first aspect, the present application provides a driving decision method, comprising:

[0005] determining a first driving action decision result and a second driving action decision result based on environment information and a preset decision model;

[0006] determining a risk perception parameter based on a preset risk perception model and a physiological signal feature;

[0007] determining a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0008] In one embodiment, the determining a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result comprises:

[0009] if the risk perception parameter is a first parameter, determining the second driving action decision result as the driving decision, the first parameter representing that the vehicle is in a high-risk state;

[0010] if the risk perception parameter is a second parameter, determining the first driving action decision result as the driving decision, the second parameter representing that the vehicle is in a low-risk state.

[0011] In one of the embodiments, the preset decision model includes a first preset decision model and a second preset decision model, the first preset decision model is trained based on a reinforcement learning model, and the second preset decision model includes one of a vehicle lateral control model, a vehicle longitudinal control model, and a vehicle lateral-longitudinal comprehensive control model. The first driving action decision result and the second driving action decision result are determined based on the environmental information and the preset decision model, including:

[0012] The first driving action decision result is obtained based on the environmental information and the first preset decision model.

[0013] The second driving action decision result is obtained based on the environmental information and the second preset decision model. In one of the embodiments, the method further includes:

[0014] The difference data between the driving decision and the first driving action decision result is calculated.

[0015] The target optimization parameter of the first preset decision model is determined based on the difference data.

[0016] The first preset decision model is updated based on the target optimization parameter to obtain an updated first preset decision model.

[0017] In one of the embodiments, the target optimization parameter of the first preset decision model is determined based on the difference data, including:

[0018] A product value of a risk perception parameter and the difference data is calculated, and the product value is determined as decision difference data of the first driving action decision result and the driving decision.

[0019] A state action value function under the environmental information and the driving decision is calculated.

[0020] A target function is determined based on the state action value function and the decision difference data.

[0021] The parameter gradient of the first preset decision model is updated according to the target function and data information in an experience replay pool to obtain the target optimization parameter of the first preset decision model.

[0022] In one of the embodiments, the method further includes:

[0023] Sample physiological signals under a preset dangerous scene are collected.

[0024] The sample physiological signals are feature-extracted to obtain sample physiological signal features.

[0025] The preset risk perception model is obtained by training based on the sample physiological signal features and the human subjective risk label, and represents a mapping relationship between the physiological signal features and the human subjective risk label.

[0026] In a second aspect, the present application further provides a driving decision device, comprising:

[0027] A first determination module is configured to determine a first driving action decision result and a second driving action decision result based on the environment information and a preset decision model.

[0028] A second determination module is configured to determine a risk perception parameter based on the preset risk perception model and the physiological signal features.

[0029] A third determination module is configured to determine a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0030] In a third aspect, the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0031] The first driving action decision result and the second driving action decision result are determined based on the environment information and a preset decision model.

[0032] The risk perception parameter is determined based on the preset risk perception model and the physiological signal features.

[0033] The driving decision is determined based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0034] In a fourth aspect, the present application further provides a computer readable storage medium, which stores a computer program, and the computer program implements the following steps when executed by a processor:

[0035] The first driving action decision result and the second driving action decision result are determined based on the environment information and a preset decision model.

[0036] The risk perception parameter is determined based on the preset risk perception model and the physiological signal features.

[0037] The driving decision is determined based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0038] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program implements the following steps when executed by a processor:

[0039] determine a first driving action decision result and a second driving action decision result based on the environmental information and a preset decision model;

[0040] determine a risk perception parameter based on the preset risk perception model and the physiological signal feature;

[0041] determine a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0042] The driving decision method, device, computer device, readable storage medium and program product can determine the first driving action decision result and the second driving action decision result based on the environmental information, determine the risk perception parameter based on the physiological signal feature, and determine the first driving action decision result or the second driving action decision result as the driving decision based on the risk perception parameter as the judgment condition, so as to make the decision result of the automatic driving vehicle match the environment in which the actual vehicle is located, improve the accuracy of the driving decision, and further protect the safety of the personnel and the vehicle. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related art, the drawings needed to be used in the description of the embodiments of the present application or the related art will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other related drawings can be obtained by those skilled in the art without creative labor.

[0044] Figure 1 a flowchart of the driving decision method in an embodiment;

[0045] Figure 2 a schematic diagram of the driving decision method in an embodiment;

[0046] Figure 3 a structural block diagram of the driving decision device in an embodiment;

[0047] Figure 4 an internal structure diagram of the computer device in an embodiment. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.

[0049] In an exemplary embodiment, as Figure 1As shown, a driving decision method is provided, and the embodiment is exemplified by applying the method to a terminal. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction of the terminal and the server. In the embodiment, the method includes the following steps:

[0050] Step 101, determining a first driving action decision result and a second driving action decision result based on environment information and a preset decision model.

[0051] The environment information can represent information of an environment in which the vehicle is located, for example, the environment information includes but is not limited to one or more of environment picture information, point cloud information, self-vehicle position, self-vehicle speed, and position and speed of surrounding obstacles. The first driving action decision result and the second driving action decision result can represent behaviors that the vehicle should take in combination with the current environment information.

[0052] Specifically, the terminal can receive environment information collected by a sensor, and can receive environment information obtained by processing initial environment data collected by a roadside device or a self-vehicle algorithm. The terminal can input the environment information into a preset decision model to output the first driving action decision result and the second driving action decision result.

[0053] Exemplarily, the terminal can obtain environment picture information through a visual sensor, obtain point cloud information through an active radar or a millimeter wave radar, and obtain self-vehicle position, speed, and position and speed of surrounding obstacles by processing initial environment data such as running time and running distance collected by a roadside device or a self-vehicle algorithm.

[0054] Step 102, determining a risk perception parameter based on a preset risk perception model and a physiological signal feature.

[0055] The physiological signal feature can reflect the physiological response of the occupant's body to the external environment change, and can be obtained by feature extraction on the physiological signal. The risk perception parameter can reflect whether the occupant perceives that the vehicle is currently in a risk state.

[0056] Specifically, the terminal can obtain a physiological signal collected by a sensing device. The terminal can pre-process the physiological signal to obtain a pre-processed physiological signal, and perform feature extraction on the pre-processed physiological signal to obtain a physiological signal feature. The terminal can input the physiological signal feature into a preset risk perception model to output a risk perception parameter.

[0057] The specific implementation process of the terminal pre-processing the physiological signal can include: the terminal can filter the physiological signal by using a filter to obtain a filtered physiological signal; the terminal can collect the physiological signal in a resting state, and perform baseline correction on the filtered physiological signal by using the physiological signal in the resting state to obtain a baseline-corrected physiological signal. The terminal can perform motion artifact removal processing on the baseline-corrected physiological signal by using a motion artifact removal algorithm to obtain a pre-processed physiological signal. The motion artifact removal algorithm can be a temporal derivative distribution repair (TDDR) algorithm or the like.

[0058] The specific implementation process of the terminal extracting features from the pre-processed physiological signal can include: the terminal can extract time domain features of the pre-processed physiological signal by using a time domain extraction tool; convert the time domain features into frequency domain features by using a Fourier transform algorithm; extract time sequence features of the pre-processed physiological signal based on a deep learning network, and determine the time domain features, the frequency domain features, and the time sequence features as physiological signal features. In addition, the physiological signal features can be processed by dimension reduction.

[0059] The specific implementation process of the terminal inputting the physiological signal features into the preset risk perception model to output a risk perception parameter can include: inputting the physiological signal features into a machine learning model, classifying different risk levels by using a machine learning method, and determining the risk level as the risk perception parameter.

[0060] For example, the terminal can collect an occupant's blood oxygen signal by using a blood oxygen detection device, and determine the blood oxygen signal as the physiological signal. The blood oxygen detection device can be a functional near-infrared spectroscopy (fNIRS). In addition, the terminal can also collect an electroencephalogram (EEG) signal collected by an EEG device on a brain cap, and determine the EEG signal as the physiological signal, and the like. It should be understood that this is only for illustration, and the physiological signal is not specifically limited.

[0061] Step 103, determining a driving decision based on the risk perception parameter, the first driving action decision result, and the second driving action decision result.

[0062] Specifically, the terminal can determine the first driving action decision result as the driving decision based on the risk perception parameter, or determine the second driving action decision result as the driving decision.

[0063] The driving decision method determines the first driving action decision result and the second driving action decision result based on the environment information, determines the risk perception parameter based on the physiological signal feature, and determines the first driving action decision result or the second driving action decision result as the driving decision based on the risk perception parameter as the judgment condition. The decision result of the autonomous vehicle is matched with the actual environment of the vehicle, so as to improve the accuracy of the driving decision and further protect the safety of the personnel and the vehicle.

[0064] In an exemplary embodiment, the specific implementation process of the step 103 "determining the driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result" can include:

[0065] If the risk perception parameter is the first parameter, the second driving action decision result is determined as the driving decision; if the risk perception parameter is the second parameter, the first driving action decision result is determined as the driving decision.

[0066] The first parameter can reflect that the vehicle is in a high-risk state, and the second parameter can reflect that the vehicle is in a low-risk state. The high-risk state can represent that the current state of the vehicle can cause harm to the vehicle and the personnel in the vehicle. The low-risk state can represent that the current state of the vehicle will not cause harm to the vehicle and the personnel in the vehicle. The second driving action decision result can be a conservative driving action decision result. The first driving action decision result can be a driving action decision result output by a driving decision model to be trained.

[0067] Specifically, if the risk perception parameter is the first parameter, the terminal determines that the vehicle is currently in a high-risk state, and determines the second driving action decision result as the driving decision. If the risk perception parameter is the second parameter, the terminal determines that the vehicle is in a normal state or a low-risk state, and determines the first driving action decision result as the driving decision. The terminal can control the driving state of the vehicle based on the driving decision.

[0068] Exemplarily, the first parameter can be 1, and the second parameter can be 0. When the risk perception parameter λ = 1, i.e., the occupant perceives that the vehicle is in a high-risk state, the second driving action decision result is determined as the driving decision. When the risk perception parameter λ = 1, i.e., the occupant perceives that the vehicle is in a normal state or a low-risk state, the first driving action decision result is determined as the driving decision. The specific formula of the driving decision can be:

[0069]

[0070] wherein a is the driving decision; λ is the risk perception parameter; a RL is the first driving action decision result; a IDM is the second driving action decision result.

[0071] In the embodiment, by taking the physiological signal features as the basis for judging driving risk, the automatic driving is more safe and humanized, the first driving action decision result can be monitored in real time, and when the first driving action decision result output does not conform to the actual scene, the second driving action decision result is switched in time to help the exploration of the safe driving strategy of reinforcement learning.

[0072] In an exemplary embodiment, the specific implementation process of step 101 "determining the first driving action decision result and the second driving action decision result based on the environment information and the preset decision model" can include:

[0073] Based on the environment information and the first preset decision model, the first driving action decision result is obtained; based on the environment information and the second preset decision model, the second driving action decision result is obtained. The preset decision model includes the first preset decision model and the second preset decision model, the first preset strategy model can be based on the learned driving decision model, and the first preset decision model can be obtained based on the reinforcement learning algorithm, wherein the reinforcement learning algorithm can be a Twin Delayed Deep Deterministic Policy Gradient (TD3) algorithm or a Deep Deterministic Policy Gradient (DDPG) reinforcement learning algorithm. The second preset decision model can be one of a vehicle lateral control model, a vehicle longitudinal control model, and a vehicle lateral and longitudinal comprehensive control model. For example, the vehicle lateral control model can be a preview angle model, the vehicle longitudinal control model can be an Intelligent Driver Model (IDM) or a proportional integral model, and the vehicle lateral and longitudinal comprehensive control model can be a Model Predictive Control (MPC) model. For example, for the longitudinal control task, the Intelligent Driver Model is selected as the second preset decision model. The selection result of the vehicle lateral and longitudinal control is related to the current control task of the vehicle.

[0074] Specifically, the terminal can obtain a first driving action decision result based on the environment information and the first preset decision model. The terminal can select a corresponding driving decision model according to a lateral and longitudinal control state of the vehicle. For a vehicle with separate longitudinal control, the terminal can input the environment data into a vehicle longitudinal control model to obtain a second driving action decision result. For a vehicle with separate lateral control, the terminal can input the environment data into a vehicle lateral control model to obtain a second driving action decision result. For a vehicle with comprehensive lateral and longitudinal control, the terminal can input the environment data into a vehicle lateral and longitudinal comprehensive control model to obtain a second driving action decision result. In addition, the terminal can also use deceleration, lane changing and other command operations output by other decision algorithms.

[0075] Exemplarily, when the vehicle performs a longitudinal control task, the terminal obtains a first driving action decision result by means of reinforcement learning, and converts the environment information into state input to obtain the first driving action decision result in the current policy of reinforcement learning. The terminal obtains a second driving action by inputting the environment information into a vehicle longitudinal control model (such as an intelligent driver model).

[0076] In this embodiment, different decision models are selected for different scenarios, which can improve the accuracy of the driving action decision result and further protect the safety of passengers in dangerous states.

[0077] In an exemplary embodiment, the driving decision method further comprises:

[0078] calculating difference data of the driving decision and the first driving action decision result; determining a target optimization parameter of the first preset decision model based on the difference data; updating the first preset decision model based on the target optimization parameter, to obtain an updated first preset decision model.

[0079] The difference data can represent the difference between the action instructions of the driving decision and the first driving action decision result, such as the difference between deceleration, acceleration or turning, and the like. The target optimization parameter can be an optimization parameter of the first preset decision model for the next state.

[0080] Specifically, the terminal can subtract the first driving action decision result from the driving decision to obtain a subtraction result, take the absolute value of the subtraction result to obtain the difference data. The terminal can calculate the target optimization parameter of the first preset decision model based on the difference data. The terminal can update the model parameters of the first preset decision model based on the target optimization parameter to obtain an updated first preset decision model.

[0081] In this embodiment, the difference data between the driving decision and the first driving action decision result is used to determine the target optimization parameter of the first preset decision model, which guides the first preset decision model to learn to make the optimal action decision in the actual scenario and accelerates the learning rate of reinforcement learning.

[0082] In an example embodiment, determining the target optimization parameter of the first preset decision model based on the difference data includes:

[0083] calculating a product value of the risk perception parameter and the difference data, determining the product value as a decision difference value of the first driving action decision and the driving decision; calculating a state-action value function of the driving decision under the environment information; determining a target function based on the state-action value function and the decision difference data; updating a parameter gradient of the first preset decision model according to the target function and data information in the experience replay pool, to obtain the target optimization parameter of the first preset decision model.

[0084] The value range of the risk perception parameter is related to the specific application scenario. The state-action value function can reflect the predicted return of the action made according to the driving decision under the current environment information. The experience replay pool can be used to store a series of data results generated by the vehicle in the process of interacting with the environment. Each experience usually includes the environment information of the vehicle, the decision action of the vehicle, the reward feedback obtained by the vehicle from the environment after executing the decision action, the next state of the vehicle, and the termination flag.

[0085] Specifically, the terminal can determine the risk perception parameter based on the current application scenario, and calculate a product value of the risk perception parameter and the difference data, and determine the product value as a decision difference value of the first driving action decision and the driving decision. The terminal can calculate a state-action value function based on a preset action value function (Q function), the current environment information, and the driving decision adopted under the current environment information. The terminal can obtain a total target function by solving the optimization target of maximizing the state-action value function and minimizing the difference data based on the risk perception parameter. The terminal can calculate a vector of all partial derivatives of the target function with respect to the initial model parameters by using a gradient calculation tool, and determine the vector as a gradient value of the initial model parameters. The terminal can calculate the mean value of the target function for the selected state-action pairs according to the batch size of the batch-selected state-action pairs. The terminal can calculate the gradient value with respect to the initial model parameters according to the mean value of the target function, multiply the gradient value by the learning rate, and perform gradient update on the initial model parameters to obtain the target optimization parameter of the first preset decision model. The specific formula of the target optimization parameter can be:

[0086]

[0087] wherein θ is the parameter of the first preset decision model; β is the learning rate; ∇θ is the gradient value of the initial model parameter; and B is the number of state-action pairs in the experience replay pool. ​is a state-action value function of the driving decision under the environmental information; s is a current state (environmental information), which is specifically defined as a multi-dimensional vector containing environmental information, and π θ is a current decision (driving decision) being executed, which is specifically defined as a mapping relationship from the state s multi-dimensional vector to an action, and λ is a risk perception parameter; φ1 is a reinforcement learning action network parameter.

[0088] The state-action value function can represent a predicted return of the selected action under the state s according to the policy π θ .

[0089] When the risk perception parameter is 1, corresponding to a high-risk state of the vehicle, the objective function is to maximize the state value function + minimize the decision difference data, and the objective function is an optimization target combined with both; when the risk perception parameter is 0, corresponding to a low-risk state of the vehicle, the objective function is to maximize the action value function.

[0090] When a = a RL , the specific formula of the target optimization parameter can be:

[0091]

[0092] At this time, it can be understood that the first preset decision model is in a normal reinforcement learning update process.

[0093] In the embodiment, the difference data is used to calculate the target optimization parameter, which guides the first preset decision model to learn to make optimal action decisions in actual scenarios, and accelerates the learning rate of reinforcement learning.

[0094] In an exemplary embodiment, the driving decision method further comprises:

[0095] Collecting sample physiological signals under a preset dangerous scene; performing feature extraction on the sample physiological signals to obtain sample physiological signal features; training based on the sample physiological signal features and a human subjective risk label to obtain a preset risk perception model.

[0096] The preset dangerous scene can be a simulated known risk scene. The human subjective risk label can be a high-risk label and a low-risk label. The preset risk perception model can be a classification model and other supervised learning models. The preset risk perception model can represent the mapping relationship between the physiological signal features and the human subjective risk label.

[0097] Specifically, the terminal can collect preset dangerous scenes and test scenes based on a driving simulator and a real car, and obtain sample physiological signals collected by a sensing device. The terminal can extract time domain features of the sample physiological signals using a time domain extraction tool, convert the time domain features into frequency domain features using a Fourier transform algorithm, extract time sequence features of the sample physiological signals based on a deep learning network, and determine the time domain features, the frequency domain features, and the time sequence features as sample physiological signal features. The terminal can input the sample physiological signal features into a human factor subjective risk label, in which, based on the collected sample physiological signal features and the human factor subjective risk label data, a dataset of a mapping relationship between the sample physiological signal features and the human factor subjective risk label is constructed, a machine learning model is trained, classification analysis is performed, and a preset risk perception model is obtained.

[0098] In addition, the terminal can preprocess the sample physiological signals to obtain preprocessed sample physiological signals. The specific preprocessing process can include: the terminal filters the sample physiological signals using a filter to obtain filtered sample physiological signals; the terminal can collect sample physiological signals in a resting state, and use the sample physiological signals in the resting state to perform baseline correction on the filtered sample physiological signals to obtain baseline corrected sample physiological signals. The terminal can use a motion artifact removal algorithm to remove motion artifacts from the baseline corrected sample physiological signals to obtain preprocessed sample physiological signals. The terminal can perform feature extraction on the preprocessed sample physiological signals.

[0099] In this embodiment, the preset risk perception model is trained by the sample physiological signals, so as to perceive the risk situation and provide a guide direction for determining an accurate driving decision.

[0100] In one exemplary embodiment, as shown in Figure 2 A driving decision method is provided, which can include two parts: a risk perception module based on the state of the occupant and a guide reinforcement learning module based on the risk perception of the occupant. The risk perception module based on the state of the occupant includes: collecting blood oxygen signals based on a blood oxygen monitoring device, blood oxygen signal preprocessing, feature extraction, and training a preset risk perception model. The reinforcement learning module based on the risk perception of the occupant includes: collecting blood oxygen signals, blood oxygen signal preprocessing, feature extraction, risk perception, collecting environmental information, reinforcement learning, and monitoring correction. The monitoring correction includes two sub-steps of correcting the decision result and guiding the reinforcement learning.

[0101] The risk perception module based on the occupant state can specifically include: in a few common high-level automatic driving scenes with risks, the terminal acquires a sample blood oxygen signal of the occupant collected by a blood oxygen detection device; the terminal pre-processes the sample blood oxygen signal to obtain a pre-processed sample blood oxygen signal; the terminal extracts features from the pre-processed sample blood oxygen signal to obtain sample blood oxygen signal features, pairs the sample blood oxygen signal features with human subjective risk labels to construct a data set, and trains a machine learning model based on the data set to obtain the human subjective risk labels.

[0102] The reinforcement learning module based on occupant risk perception can specifically include: in an unknown high-level automatic driving scene with risks, the terminal can acquire a blood oxygen signal, pre-process the blood oxygen signal to obtain a pre-processed blood oxygen signal, and extract features from the pre-processed blood oxygen signal to obtain blood oxygen signal features; the terminal inputs the blood oxygen signal features into a preset risk perception model to obtain risk perception parameters; the terminal can acquire environmental information, inputs the environmental information into an intelligent driver model (a second preset decision model) to obtain a decision result (a second driving action decision result) of the intelligent driver model, and inputs the environmental information into a reinforcement learning model (a first preset decision model) to obtain a decision result (a first driving action decision result) of the reinforcement learning model; when the risk perception parameters represent a high risk, the decision result of the intelligent driver model is used as a driving decision; otherwise, the decision result of the reinforcement learning model is used as the driving decision. The terminal can introduce a deviation of the driving decision from the decision result of the reinforcement learning model into an optimization objective of the reinforcement learning model to perform reinforcement learning and correction on the reinforcement learning model, and correct the reinforcement learning model.

[0103] To remove noise such as breathing, heart rate, blood pressure fluctuations, Mayer wave noise, and factors such as motion that affect the result noise, the sample signal pre-processing process can specifically include: filtering the data using a filter, baseline correction based on resting data, and motion artifact removal processing using a time derivative distribution repair TDDR algorithm. In addition, the original data dimension is relatively large, and feature extraction and dimension reduction operations are required.

[0104] In one example, the risk perception of the occupant is extracted based on the blood oxygen signal, and the subjective risk perception result of the occupant is used as a judgment condition, that is, when the occupant feels a high risk, the automatic driving strategy (driving decision) is selected as the first driving action decision result; when the occupant feels a low risk, the action of reinforcement learning is not intervened, and in order to effectively utilize the decision result of the intelligent driver model when the occupant feels a high risk, timely guide the correction of reinforcement learning in this scenario, and accelerate the learning rate of reinforcement learning, the deviation of the decision result of the intelligent driver model from the output result of reinforcement learning is introduced into the optimization objective of the reinforcement learning action network to accelerate the learning rate of reinforcement learning.

[0105] In combination with the fNIRS technology, physiological information of the occupant is taken as a risk judgment basis for the automatic driving, and a guided reinforcement learning algorithm is combined to realize a safer, more human, fast and reliable automatic driving training method. The fNIRS technology is further applied to the driving decision of the automatic driving. The embodiments of the present application have high reference and practical value in the intelligent decision field of the high-level automatic driving system.

[0106] It should be understood that, although each step in the flowchart involved in each embodiment as described above is displayed in sequence according to the arrow, these steps are not necessarily executed in sequence according to the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other sequences. Moreover, at least part of the steps in the flowchart involved in each embodiment as described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or steps or stages in other steps.

[0107] Based on the same inventive concept, the embodiments of the present application also provide a driving decision device for implementing the driving decision method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more driving decision device embodiments provided below can refer to the limitations of the driving decision method in the above text, and will not be repeated here.

[0108] In one exemplary embodiment, as shown in Figure 3 A driving decision device 30 is provided, comprising: a first determination module 31, a second determination module 32 and a third determination module 33, wherein:

[0109] The first determination module 31 is configured to determine a first driving action decision result and a second driving action decision result based on the environment information and the preset decision model;

[0110] The second determination module 32 is configured to determine a risk perception parameter based on the preset risk perception model and the physiological signal feature;

[0111] The third determination module 33 is configured to determine a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result.

[0112] In an example embodiment, the third determining module 33 is configured to determine the second driving action decision result as the driving decision if the risk perception parameter is the first parameter, and determine the first driving action decision result as the driving decision if the risk perception parameter is the second parameter.

[0113] In an example embodiment, the first determining module 31 is configured to obtain a first driving action decision result based on the environment information and a first preset decision model, and obtain a second driving action decision result based on the environment information and a second preset decision model.

[0114] In an example embodiment, the first determining module 31 is further configured to calculate difference data of the driving decision and the first driving action decision result, determine a target optimization parameter of the first preset decision model based on the difference data, and update the first preset decision model based on the target optimization parameter to obtain an updated first preset decision model.

[0115] In an example embodiment, the first determining module 31 is further configured to calculate a product value of the risk perception parameter and the difference data, determine the product value as decision difference data of the first driving action decision result and the driving decision, calculate a state action value function of the driving decision under the environment information, determine a target function based on the state action value function and the decision difference data, and update a parameter gradient of the first preset decision model according to the target function and data information in an experience replay pool to obtain the target optimization parameter of the first preset decision model.

[0116] In an example embodiment, the training module is configured to collect sample physiological signals under a preset dangerous scene, extract features of the sample physiological signals to obtain sample physiological signal features, and train based on the sample physiological signal features and a human subjective risk label to obtain a preset risk perception model, the preset risk perception model representing a mapping relationship between the physiological signal features and the human subjective risk label.

[0117] The above-described various modules in the driving decision device can be realized in whole or in part by software, hardware, and a combination thereof. The above-described various modules can be embedded in or independent of a processor in a computer device in a hardware form, or stored in a memory in the computer device in a software form, so as to be called and executed by a processor to perform operations corresponding to the above-described various modules.

[0118] In an example embodiment, a computer device is provided, which can be a terminal, and an internal structure diagram of the computer device can be as shown in Figure 4As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through the system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with the external terminal in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, near field communication (Near Field Communication, NFC) or other technologies. The computer program is executed by the processor to realize a driving decision method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.

[0119] Those skilled in the art can understand that, Figure 4 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0120] In one exemplary embodiment, a computer device is provided, comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the following steps:

[0121] Based on the environmental information and the preset decision model, determine the first driving action decision result and the second driving action decision result;

[0122] Based on the preset risk perception model and the physiological signal feature, determine the risk perception parameter;

[0123] Based on the risk perception parameter, the first driving action decision result and the second driving action decision result, determine the driving decision.

[0124] In one embodiment, the processor executing the computer program further implements the following steps:

[0125] The first driving action decision result and the second driving action decision result determine a driving decision, including:

[0126] If the risk perception parameter is the first parameter, the second driving action decision result is determined as the driving decision;

[0127] If the risk perception parameter is the second parameter, the first driving action decision result is determined as the driving decision.

[0128] In one embodiment, the processor, when executing the computer program, also implements the following steps:

[0129] Based on the environment information and the first preset decision model, a first driving action decision result is obtained;

[0130] Based on the environment information and the second preset decision model, a second driving action decision result is obtained.

[0131] In one embodiment, the processor, when executing the computer program, also implements the following steps:

[0132] The difference data of the driving decision and the first driving action decision result is calculated;

[0133] Based on the difference data, a target optimization parameter of the first preset decision model is determined;

[0134] Based on the target optimization parameter, the first preset decision model is updated to obtain an updated first preset decision model.

[0135] In one embodiment, the processor, when executing the computer program, also implements the following steps:

[0136] The product value of the risk perception parameter and the difference data is calculated, and the product value is determined as the decision difference data of the first driving action decision result and the driving decision;

[0137] The state action value function under the environment information using the driving decision is calculated;

[0138] Based on the state action value function and the decision difference data, a target function is determined; according to the target function and the data information in the experience replay pool, the parameter gradient update of the first preset decision model is obtained, and the target optimization parameter of the first preset decision model is obtained.

[0139] In one embodiment, the processor, when executing the computer program, also implements the following steps:

[0140] Sample physiological signals under a preset dangerous scene are collected;

[0141] The sample physiological signal is subjected to feature extraction to obtain a sample physiological signal feature;

[0142] The sample physiological signal feature and the human factor subjective risk label are used for training to obtain a preset risk perception model, and the preset risk perception model represents a mapping relationship between the physiological signal feature and the human factor subjective risk label.

[0143] In one embodiment, a computer readable storage medium is provided, and the computer readable storage medium stores a computer program. The computer program is executed by a processor to implement the following steps:

[0144] Based on the environmental information and the preset decision model, a first driving action decision result and a second driving action decision result are determined;

[0145] Based on the preset risk perception model and the physiological signal feature, a risk perception parameter is determined;

[0146] Based on the risk perception parameter, the first driving action decision result and the second driving action decision result, a driving decision is determined.

[0147] In one embodiment, a computer program product is provided, and the computer program product includes a computer program. The computer program is executed by a processor to implement the following steps:

[0148] Based on the environmental information and the preset decision model, a first driving action decision result and a second driving action decision result are determined;

[0149] Based on the preset risk perception model and the physiological signal feature, a risk perception parameter is determined;

[0150] Based on the risk perception parameter, the first driving action decision result and the second driving action decision result, a driving decision is determined.

[0151] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of the related data need to comply with relevant regulations.

[0152] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. The non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. The volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, the RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., without being limited thereto.

[0153] The technical features of the above embodiments can be combined in any manner. To make the description concise, all possible combinations of the technical features in the above embodiments are not described, but as long as the combinations of the technical features do not exist contradictions, they should be considered as the scope of the present application.

[0154] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.

Claims

1. A driving decision method characterized by, The method comprises: determining a first driving action decision result and a second driving action decision result based on environmental information and a preset decision model; determining a risk perception parameter based on a preset risk perception model and physiological signal features; determining a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result; wherein the determination of the driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result comprises: if the risk perception parameter is a first parameter, determining the second driving action decision result as the driving decision, the first parameter representing that the vehicle is in a high-risk state; if the risk perception parameter is a second parameter, determining the first driving action decision result as the driving decision, the second parameter representing that the vehicle is in a low-risk state; the preset decision model comprises a first preset decision model and a second preset decision model, the first preset decision model being trained based on a reinforcement learning model, and the second preset decision model comprising one of a vehicle lateral control model, a vehicle longitudinal control model and a vehicle lateral-longitudinal comprehensive control model, the determination of the first driving action decision result and the second driving action decision result based on the environmental information and the preset decision model comprising: obtaining the first driving action decision result based on the environmental information and the first preset decision model; obtaining the second driving action decision result based on the environmental information and the second preset decision model; the method further comprises: collecting sample physiological signals in a preset dangerous scene; extracting features from the sample physiological signals to obtain sample physiological signal features; training based on the sample physiological signal features and a human subjective risk label to obtain a preset risk perception model, the preset risk perception model representing a mapping relationship between physiological signal features and a human subjective risk label.

2. The method of claim 1, wherein, The method further comprises: calculating difference data of the driving decision and the first driving action decision result; determining a target optimization parameter of the first preset decision model based on the difference data; updating the first preset decision model based on the target optimization parameter to obtain an updated first preset decision model.

3. The method of claim 2, wherein, The determination of the target optimization parameter of the first preset decision model based on the difference data comprises: calculating a product value of the risk perception parameter and the difference data, and determining the product value as decision difference data of the first driving action decision result and the driving decision; calculating a state action value function of the driving decision under the environmental information; determining a target function based on the state action value function and the decision difference data; updating a parameter gradient of the first preset decision model according to the target function and data information in an experience replay pool to obtain the target optimization parameter of the first preset decision model.

4. The method of claim 2, wherein, The difference data represents the difference between the driving decision and the action instruction of the first driving action decision result.

5. A driving decision device characterized by comprising: The device comprises: The first determining module is configured to determine a first driving action decision result and a second driving action decision result based on the environment information and a preset decision model; The second determining module is configured to determine a risk perception parameter based on a preset risk perception model and the physiological signal feature; The third determining module is configured to determine a driving decision based on the risk perception parameter, the first driving action decision result and the second driving action decision result; The third determining module is specifically configured to determine the second driving action decision result as the driving decision if the risk perception parameter is a first parameter, the first parameter indicating that the vehicle is in a high-risk state; The third determining module is specifically configured to determine the first driving action decision result as the driving decision if the risk perception parameter is a second parameter, the second parameter indicating that the vehicle is in a low-risk state; The preset decision model includes a first preset decision model and a second preset decision model, the first preset decision model being trained based on a reinforcement learning model, and the second preset decision model including one of a vehicle lateral control model, a vehicle longitudinal control model and a vehicle lateral-longitudinal comprehensive control model; The first determining module is specifically configured to obtain the first driving action decision result based on the environment information and the first preset decision model; obtain the second driving action decision result based on the environment information and the second preset decision model; The device further includes a training module, which is configured to collect sample physiological signals in a preset dangerous scene; extract features from the sample physiological signals to obtain sample physiological signal features; 6. The apparatus of claim 5, wherein, train based on the sample physiological signal features and a human subjective risk label to obtain a preset risk perception model, the preset risk perception model being a mapping relationship between the physiological signal features and the human subjective risk label. The first determining module is further configured to calculate difference data of the driving decision and the first driving action decision result; determine target optimization parameters of the first preset decision model based on the difference data; 7. The apparatus of claim 6, wherein, update the first preset decision model based on the target optimization parameters to obtain an updated first preset decision model. The first determining module is further configured to calculate a product value of the risk perception parameter and the difference data, and determine the product value as decision difference data of the first driving action decision result and the driving decision; calculate a state action value function of the driving decision under the environment information; determine a target function based on the state action value function and the decision difference data; 8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, update a parameter gradient of the first preset decision model according to the target function and data information in an experience replay pool to obtain target optimization parameters of the first preset decision model.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 4.

10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4. The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Intelligent vehicle integrated decision-making system based on man-machine co-driving idea

    CN114802306A

  • Autonomous driving method and device, and vehicle

    WO2024138453A1