Vehicle Behavior Decision Method, Device, Equipment and Readable Storage Medium
The vehicle behavior decision-making method constructs an expert rule library and trains an estimation network to address decision-making challenges in complex tunnel environments, enhancing vehicle decision-making accuracy and rationality.
Patent Information
- Application Number
- CN202211729358.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-30
AI Technical Summary
In the complex underground space environment of the mining area, it is difficult for vehicles to make reasonable decisions and control, resulting in behavioral decision-making errors and trajectory planning that are not feasible, and tracking errors increase.
An expert rule base is built based on the vehicle's historical driving environment and decision-making instructions, and the valuation network is trained through a deep reinforcement learning algorithm until the training is completed, and the optimal decision-making instructions at the current moment are output.
By strengthening the training mechanism, reasonable decision-making control of vehicles in complex environments is achieved, and the accuracy of decision-making and feasibility of trajectory planning are improved.
Smart Images

Figure CN116061967B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and particularly to a vehicle behavior decision-making method, device, equipment and readable storage medium. Background Art
[0002] Due to the mixed vehicles, narrow space and multi-slope sharp turns in the underground mine area, the driving of vehicles in the underground mine area is severely restricted. Using the traditional decision control method, due to the lack of the deduction mechanism between the environmental variables and the decision control variables, it is difficult to make a reasonable decision control in the complex underground mine space environment, and it is very easy to cause problems such as vehicle behavior decision-making errors, infeasible trajectory planning and increased tracking error. Summary of the Invention
[0003] The main purpose of the present invention is to provide a vehicle behavior decision-making method, device, equipment and readable storage medium, aiming to solve the problem that it is difficult for current vehicles to make reasonable decision control in a complex space environment.
[0004] In a first aspect, the present invention provides a vehicle behavior decision-making method, and the vehicle behavior decision-making method includes:
[0005] Construct an expert rule base based on the historical driving environment and historical decision instructions of the vehicle, wherein the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane;
[0006] Train a valuation network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches a preset number of times, and obtain a trained valuation network;
[0007] Input the driving environment of the vehicle at the current moment into the trained valuation network to obtain the decision instruction output by the trained valuation network.
[0008] Optionally, the step of constructing an expert rule base based on the historical driving environment and historical decision instructions of the vehicle includes:
[0009] Obtain the correspondence between the historical driving environment and historical decision instructions of the vehicle;
[0010] Construct an expert rule base based on the correspondence between the historical driving environment and historical decision instructions of the vehicle, wherein the historical driving environment of the vehicle includes the lane where the vehicle is located and the obstacle conditions in each lane.
[0011] Optionally, the preset number of times includes a first preset number of times and a second preset number of times, and the step of training a valuation network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches the preset number of times and obtaining a trained valuation network includes:
[0012] Obtain the driving environment of the vehicle at time t during the training process;
[0013] Input the driving environment at time t into the estimation network and the expert rule base respectively;
[0014] Detect whether there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t;
[0015] If there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision instruction corresponding to the target historical driving environment and the first decision instruction output by the estimation network;
[0016] If there is no target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the second decision instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm;
[0017] Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction, where time t + 1 is the next moment of time t;
[0018] Update the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value;
[0019] Use the driving environment at time t, the first decision instruction or the second decision instruction, the new reward value, and the driving environment at time t + 1 as a set of training data, and return to execute the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process until N sets of training data are obtained, where N is a positive integer;
[0020] Arbitrarily select a set of training data from the N sets of training data, and update the weights of the estimation network through the loss function;
[0021] Use the driving environment at time t + 1 as the driving environment at time t, and return to execute the step of inputting the driving environment at time t into the estimation network and the expert rule base respectively until the number of times of updating the weights of the estimation network is greater than or equal to the first preset number of times, and obtain the trained estimation network, where the weights of the estimation network updated every second preset number of times are assigned to the target network, and the second preset number of times is less than the first preset number of times.
[0022] Optionally, the step of updating the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value includes:
[0023] Detect whether the first decision instruction is the same as the target historical decision instruction;
[0024] If the detection result is that the first decision instruction is the same as the target historical decision instruction, then calculate the sum of the reward value and the additional reward value to obtain a new reward value;
[0025] If the detection result is that the first decision instruction is different from the target historical decision instruction, then calculate the difference between the reward value and the additional reward value to obtain a new reward value.
[0026] Optionally, the step of arbitrarily selecting a set of training data from N sets of training data and updating the weights of the evaluation network through a loss function includes:
[0027] Input the training data into the evaluation network and the target network respectively to obtain the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data;
[0028] Detect whether the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data are the same;
[0029] If the predicted Q value and the target Q value are the same, the weights of the evaluation network remain unchanged;
[0030] If the predicted Q value and the target Q value are the same, use the gradient descent method to solve the loss function to obtain the new weights of the evaluation network.
[0031] In a second aspect, the present invention further provides a vehicle behavior decision-making device, where the vehicle behavior decision-making device includes:
[0032] A construction module, configured to construct an expert rule base based on the historical driving environment of the vehicle and the historical decision instruction, where the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane;
[0033] A training module, configured to train the evaluation network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches a preset number of times, and obtain a trained evaluation network;
[0034] A decision instruction acquisition module, configured to input the driving environment of the vehicle at the current moment into the trained evaluation network to obtain the decision instruction output by the trained evaluation network.
[0035] Optionally, the preset number of times includes a first preset number of times and a second preset number of times, and the training module is specifically configured to:
[0036] Obtain the driving environment of the vehicle at time t during the training process;
[0037] Input the driving environment at time t into the estimation network and the expert rule base respectively;
[0038] Detect whether there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t;
[0039] If there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision instruction corresponding to the target historical driving environment and the first decision instruction output by the estimation network;
[0040] If there is no target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the second decision instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm;
[0041] Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction, where time t + 1 is the next moment of time t;
[0042] Update the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value;
[0043] Use the driving environment at time t, the first decision instruction or the second decision instruction, the new reward value, and the driving environment at time t + 1 as a set of training data, and return to the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process until N sets of training data are obtained, where N is a positive integer;
[0044] Arbitrarily select a set of training data from the N sets of training data and update the weights of the estimation network through the loss function;
[0045] Use the driving environment at time t + 1 as the driving environment at time t, and return to the step of inputting the driving environment at time t into the estimation network and the expert rule base respectively until the number of times of updating the weights of the estimation network is greater than or equal to the first preset number of times to obtain a trained estimation network, where the weights of the estimation network updated every second preset number of times are assigned to the target network, and the second preset number of times is less than the first preset number of times.
[0046] Optionally, the training module is further configured to:
[0047] Input the training data into the estimation network and the target network respectively to obtain the predicted Q value output by the estimation network and the target Q value output by the target network corresponding to each set of training data;
[0048] Check whether the predicted Q-value output by the evaluation network corresponding to each set of training data is the same as the target Q-value output by the target network;
[0049] If the predicted Q-value is the same as the target Q-value, the weights of the evaluation network remain unchanged;
[0050] If the predicted Q-value is the same as the target Q-value, use the gradient descent method to solve the loss function to obtain the new weights of the evaluation network.
[0051] In a third aspect, the present invention further provides a vehicle behavior decision-making device, which includes a processor, a memory, and a vehicle behavior decision-making program stored on the memory and executable by the processor. When the vehicle behavior decision-making program is executed by the processor, the steps of the vehicle behavior decision-making method described above are implemented.
[0052] In a fourth aspect, the present invention further provides a readable storage medium, on which a vehicle behavior decision-making program is stored. When the vehicle behavior decision-making program is executed by a processor, the steps of the vehicle behavior decision-making method described above are implemented.
[0053] In the present invention, an expert rule base is constructed based on the historical driving environment and historical decision-making instructions of the vehicle. Among them, the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane; the evaluation network is trained based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches a preset number of times, and a trained evaluation network is obtained; the driving environment of the vehicle at the current moment is input into the trained evaluation network, and a decision-making instruction output by the trained evaluation network is obtained. Through the present invention, during the process of continuously strengthening the training of the evaluation network based on the expert rule base and the driving environment of the vehicle during the training of the evaluation network, the deduction mechanism between the environmental variables and the decision control variables is continuously carried out. Therefore, when the driving environment of the vehicle at the current moment is input into the trained evaluation network, the decision-making instruction output by the trained evaluation network is the optimal decision-making instruction corresponding to the environment where the vehicle is located at the current moment, solving the problem that it is difficult for current vehicles to make reasonable decision control in a complex spatial environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a schematic flowchart of an embodiment of the vehicle behavior decision-making method of the present invention;
[0055] Figure 2 is Figure 1 a detailed flowchart of step S20 in;
[0056] Figure 3 It is an architecture diagram for training the evaluation network of the present invention;
[0057] Figure 4 For Figure 2 Refined process schematic diagram of step S209 in
[0058] Figure 5 Functional module schematic diagram of an embodiment of the vehicle behavior decision-making device of the present invention;
[0059] Figure 6 Hardware structure schematic diagram of the vehicle behavior decision-making device involved in the solution of the embodiment of the present invention.
[0060] The realization, functional characteristics and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the accompanying drawings. Specific embodiments
[0061] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0062] In a first aspect, an embodiment of the present invention provides a vehicle behavior decision-making method.
[0063] In one embodiment, with reference to Figure 1 , Figure 1 Is a process schematic diagram of an embodiment of the vehicle behavior decision-making method of the present invention. As Figure 1 Shown, the vehicle behavior decision-making method includes:
[0064] Step S10, constructing an expert rule base based on the historical driving environment and historical decision-making instructions of the vehicle, wherein the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane;
[0065] In this embodiment, an expert rule base is constructed based on various historical driving environments of multiple vehicles or various historical driving environments of a single vehicle, and the historical decision-making instructions corresponding to each historical driving environment. Among them, the historical driving environment of the vehicle includes the lane where the vehicle is located and the obstacles in each lane, and the obstacles include vehicles, pedestrians and / or fences, etc.
[0066] Further, in one embodiment, step S10 includes:
[0067] Obtain the correspondence between the historical driving environment and historical decision-making instructions of the vehicle;
[0068] Construct an expert rule base based on the correspondence between the historical driving environment and historical decision-making instructions of the vehicle, wherein the historical driving environment of the vehicle includes the lane where the vehicle is located and the obstacle conditions in each lane.
[0069] In this embodiment, multiple historical driving environments of multiple vehicles or multiple historical driving environments of a single vehicle, and the historical decision-making instructions corresponding to each historical driving environment are obtained. An expert rule base is constructed according to the correspondence between the historical driving environment of the vehicle and the historical decision-making instructions. Specifically, the expert rule base is shown in Table 1 below.
[0070] Table 1
[0071] Serial number Expert rule base 1 IF(C = 0) AND (W0 = 1) AND (W1 = 0) Then A = 2 2 IF(C = 0) AND (W0 = 0) Then A = 3 3 IF(C = 1) AND (W0 = 1) AND (W1 = 1) AND (W2 = 0) Then A = 2 4 IF(C = 1) AND (W0 = 0) AND (W1 = 1) AND (W2 = 1) Then A = 0 5 IF(C = 1) AND (W1 = 0) Then A = 3 6 IF(C = 2) AND (W1 = 0) AND (W2 = 1) Then A = 0 7 IF(C = 2) AND (W2 = 0) Then A = 3 8 IF(C = 2) AND (W1 = 1) AND (W2 = 1) Then A = 4
[0072] C represents the lane where the vehicle is located. Among them, C = 0 means the vehicle is in the left lane, C = 1 means the vehicle is in the middle lane, and C = 2 means the vehicle is in the right lane. W0 = 0 means there is no obstacle in the left lane, and W0 = 1 means there is an obstacle in the left lane. W1 = 0 means there is no obstacle in the middle lane, and W1 = 1 means there is an obstacle in the middle lane. W2 = 0 means there is no obstacle in the right lane, and W2 = 1 means there is an obstacle in the right lane. A represents the historical decision-making instruction corresponding to each historical driving environment. Among them, A = 0 means the vehicle changes lanes to the left, A = 1 means the vehicle goes straight at the current speed, A = 2 means the vehicle changes lanes to the right, A = 3 means the vehicle accelerates, and A = 4 means the vehicle decelerates.
[0073] Taking the obstacle as a dangerous vehicle that may cause a collision in front of the vehicle as an example, as shown in serial number 1, when the vehicle is in the left lane and there is a dangerous vehicle in the left lane and no dangerous vehicle in the middle lane, the optimal decision-making instruction for the vehicle is to change lanes to the right. As shown in serial number 2, when the vehicle is in the left lane and there is no dangerous vehicle in the left lane, the optimal decision-making instruction for the vehicle is to accelerate. As shown in serial number 3, when the vehicle is in the middle lane and there is a dangerous vehicle in the left lane, a dangerous vehicle in the middle lane, and no dangerous vehicle in the right lane, the optimal decision-making instruction for the vehicle is to change lanes to the right. As shown in serial number 4, when the vehicle is in the middle lane and there is no dangerous vehicle in the left lane, a dangerous vehicle in the middle lane, and a dangerous vehicle in the right lane, the optimal decision-making instruction for the vehicle is to change lanes to the left. As shown in serial number 5, when the vehicle is in the middle lane and there is no dangerous vehicle in the middle lane, the optimal decision-making instruction for the vehicle is to accelerate. As shown in serial number 6, when the vehicle is in the right lane and there is no dangerous vehicle in the middle lane and a dangerous vehicle in the right lane, the optimal decision-making instruction for the vehicle is to change lanes to the left. As shown in serial number 7, when the vehicle is in the right lane and there is no dangerous vehicle in the right lane, the optimal decision-making instruction for the vehicle is to accelerate. As shown in serial number 8, when the vehicle is in the right lane and there is a dangerous vehicle in the middle lane and a dangerous vehicle in the right lane, the optimal decision-making instruction for the vehicle is to decelerate.
[0074] Step S20: Train the value network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches the preset number of times, and a trained value network is obtained.
[0075] In this embodiment, a deep reinforcement learning algorithm guided by an expert rule base is used to train the valuation network, obtain the driving environment of the vehicle during the training of the valuation network, and continuously perform reinforcement training on the valuation network based on the expert rule base and the driving environment of the vehicle during the training of the valuation network until the number of training times reaches a preset number, and then a trained valuation network is obtained.
[0076] Step S30: Input the driving environment of the vehicle at the current moment into the trained valuation network to obtain a decision instruction output by the trained valuation network.
[0077] In this embodiment, a trained valuation network is obtained. After that, by inputting the driving environment of the vehicle at the current moment into the trained valuation network, a decision instruction output by the trained valuation network can be obtained. It is easy to understand that the decision instruction output by the trained valuation network is the optimal decision instruction corresponding to the driving environment of the vehicle at the current moment.
[0078] In this embodiment, an expert rule base is constructed based on the historical driving environment and historical decision instructions of the vehicle. Among them, the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane; the valuation network is trained based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches a preset number, and then a trained valuation network is obtained; the driving environment of the vehicle at the current moment is input into the trained valuation network to obtain a decision instruction output by the trained valuation network. Through this embodiment, in the process of continuously performing reinforcement training on the valuation network based on the expert rule base and the driving environment of the vehicle during the training of the valuation network, the deduction mechanism between the environmental variables and the decision control variables is continuously carried out. Therefore, by inputting the driving environment of the vehicle at the current moment into the trained valuation network, the decision instruction output by the trained valuation network obtained is the optimal decision instruction corresponding to the environment where the vehicle is located at the current moment, solving the problem that it is difficult for current vehicles to make reasonable decision control in a complex spatial environment.
[0079] Further, in an embodiment, refer to Figure 2 , Figure 2 which is Figure 1 a detailed flowchart of step S20 in Figure 2 As shown in
[0080] Step S201: Obtain the driving environment of the vehicle at time t during the training process;
[0081] Step S202: Input the driving environment at time t into the valuation network and the expert rule base respectively;
[0082] Step S203: Detect whether there is a target historical driving environment in the historical driving environments of the vehicles included in the expert rule base that is the same as the driving environment at time t;
[0083] Step S204: If there is a target historical driving environment in the historical driving environments of the vehicles included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision instruction corresponding to the target historical driving environment and the first decision instruction output by the evaluation network;
[0084] Step S205: If there is no target historical driving environment in the historical driving environments of the vehicles included in the expert rule base that is the same as the driving environment at time t, obtain a second decision instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm;
[0085] Step S206: Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction, where time t + 1 is the next moment of time t;
[0086] Step S207: Update the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value;
[0087] Step S208: Use the driving environment at time t, the first decision instruction or the second decision instruction, the new reward value, and the driving environment at time t + 1 as a set of training data, and return to the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process, until N sets of training data are obtained, where N is a positive integer;
[0088] Step S209: Arbitrarily select a set of training data from the N sets of training data, and update the weights of the evaluation network through the loss function;
[0089] Step S210: Use the driving environment at time t + 1 as the driving environment at time t, and return to the step of inputting the driving environment at time t into the evaluation network and the expert rule base respectively, until the number of times of updating the weights of the evaluation network is greater than or equal to the first preset number of times, and obtain a trained evaluation network, where the weights of the evaluation network updated every second preset number of times are assigned to the target network, and the second preset number of times is less than the first preset number of times.
[0090] In this embodiment, refer to Figure 3 , Figure 3 is the architecture diagram of the training evaluation network of the present invention. As Figure 3 shown, obtain the driving environment ψ of the vehicle at time t during the training of the evaluation network t .
[0091] Input the driving environment ψ at time t t into the estimation network and the expert rule base respectively. Detect whether there is a target historical driving environment in the historical driving environments of the vehicle included in the expert rule base that is the same as the driving environment at time t.
[0092] If the detection result is that there is a target historical driving environment in the historical driving environments of the vehicle included in the expert rule base that is the same as the driving environment at time t, then the first decision instruction output by the estimation network is obtained, and the target historical decision instruction corresponding to the target historical driving environment is obtained.
[0093] If the detection result is that there is no target historical driving environment in the historical driving environments of the vehicle included in the expert rule base that is the same as the driving environment at time t, then the second decision instruction corresponding to the driving environment at time t during the training process of the vehicle is obtained based on the greedy algorithm. Among them, when solving the problem by the greedy algorithm, it always makes the best choice at present. That is to say, it does not consider the overall optimality, and only makes a locally optimal solution.
[0094] After obtaining the first decision instruction or the second decision instruction, the vehicle executes the first decision instruction or the second decision instruction. Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction. Among them, and time t + 1 is the next moment of time t.
[0095] After obtaining the first decision instruction, the target historical decision instruction, and the reward value, based on the first decision instruction, the target historical decision instruction, and the additional reward value, the reward value is updated through the reward value function to obtain a new reward value.
[0096] After obtaining the new reward value, with the driving environment ψ at time t t 、the first decision instruction A1 or the second decision instruction A2, the new reward value r′, and the driving environment ψ at time t + 1 t+1 as a set of training data {ψ t , A, r′, ψ t+1} is stored in the experience pool. Among them, when there is a target historical driving environment, A = A1, and when there is no target historical driving environment, A = A2. Then return to execute the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment at time t during the training process of the vehicle, that is, loop to execute steps S203 to S208. It is easy to think that each time it loops, the number of groups of training data will increase by one group. Until N groups of training data are obtained, where N is a positive integer.
[0097] Then, randomly select a set of training data from the experience pool, i.e., N sets of training data, and update the weights of the evaluation network through the loss function.
[0098] Check whether the number of times of updating the weights of the evaluation network is less than the first preset number of times. If the number of times of updating the weights of the evaluation network is less than the first preset number of times, use the driving environment at time t + 1 as the driving environment at time t, and return to execute the step of inputting the driving environment at time t into the evaluation network and the expert rule base respectively, that is, loop to execute step S202 to step S210 until the number of times of updating the weights of the evaluation network is greater than or equal to the first preset number of times. At this time, the evaluation network is the trained evaluation network, and step S202 to step S210 are no longer looped. Among them, assign the weights of the evaluation network updated every second preset number of times to the target network, and the second preset number of times is less than the first preset number of times. Specifically, if the first preset number of times is 100 times and the second preset number of times is 20 times, then every 20 times of updating the weights of the evaluation network, assign the weights of the evaluation network updated 20 times to the target network, so that the weight value of the target network is the same as the weight value of the evaluation network updated 20 times. That is, when the weights of the evaluation network are updated 20 times, assign the weights ω of the evaluation network to the target network, that is, assign the weights ω - of the target network to ω. When the weights of the evaluation network are updated 40 times, assign the weights of the evaluation network to the target network again, and so on. Every time the weights of the evaluation network are updated by the second preset number of times, assign the weights of the evaluation network to the target network once. During the process of updating the weights of the evaluation network, the weights of the target network are updated only once after the weights of the evaluation network are updated by the second preset number of times. Therefore, the target Q value output by the target network is relatively fixed for a period of time. Therefore, the introduction of the target network increases the stability of learning.
[0099] Further, in one embodiment, step S207 includes:
[0100] Check whether the first decision instruction and the target historical decision instruction are the same;
[0101] If the detection result is that the preliminary decision instruction and the target historical decision instruction are the same, calculate the sum of the reward value and the additional reward value to obtain a new reward value;
[0102] If the detection result is that the preliminary decision instruction and the target historical decision instruction are not the same, calculate the difference between the reward value and the additional reward value to obtain a new reward value.
[0103] In this embodiment, after obtaining the first decision instruction, the target historical decision instruction, the additional reward value, and the reward value, check whether the first decision instruction and the target historical decision instruction are the same.
[0104] If the detection result shows that the first decision instruction is the same as the target historical decision instruction, then calculate the sum of the reward value and the additional reward value to obtain a new reward value r′, that is, r′ = r + η, where r represents the reward value and η represents the additional reward value.
[0105] If the detection result shows that the first decision instruction is different from the target historical decision instruction, then calculate the difference between the reward value and the additional reward value to obtain a new reward value r′, that is, r′ = r - η.
[0106] Further, in one embodiment, refer to Figure 4 , Figure 4 For Figure 2 is a detailed flowchart of step S209 in Figure 4 As shown, step S209 includes:
[0107] Step S290: Input the training data into the evaluation network and the target network respectively to obtain the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each group of training data;
[0108] Step S291: Detect whether the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each group of training data are the same;
[0109] Step S292: If the predicted Q value and the target Q value are the same, the weights of the evaluation network remain unchanged;
[0110] Step S293: If the predicted Q value and the target Q value are the same, use the gradient descent method to solve the loss function to obtain the new weights of the evaluation network.
[0111] In this embodiment, taking N as 10 as an example, randomly select a group of training data from 10 groups of training data and input it into the evaluation network and the target network respectively to obtain the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each group of training data.
[0112] Whether the predicted Q value output by the evaluation network and the target Q value output by the target network are the same.
[0113] If the predicted Q value output by the evaluation network corresponding to the first group of training data and the target Q value output by the target network are the same, the weights of the evaluation network remain unchanged.
[0114] If the predicted Q value output by the evaluation network and the target Q value output by the target network are different, use the gradient descent method to solve the loss function to obtain the new weights of the evaluation network, where the loss function L(w) is as follows:
[0115] L(w) = E[(Q target -Q main (ψ t , A; ω))2 , where Q target = r′ + γmax a′ Q main (ψ t+1 , A′; ω),
[0116] Q target represents the target Q value output by the target network, ψ t represents the driving environment at time t, A represents the decision-making instruction at time t, ω represents the weight, r′ represents the new reward value, γ represents the discount factor, and ψ t+1 represents the driving environment at time t + 1, and A′ represents the decision-making instruction at time t + 1.
[0117] In a second aspect, an embodiment of the present invention further provides a vehicle behavior decision-making device.
[0118] In one embodiment, referring to Figure 5 , Figure 5 is a schematic diagram of the function modules of an embodiment of the vehicle behavior decision-making device of the present invention. As Figure 5 shown, the vehicle behavior decision-making device includes:
[0119] A construction module 10, configured to construct an expert rule library based on the historical driving environment and historical decision-making instructions of the vehicle, where the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane;
[0120] A training module 20, configured to train the value network based on the expert rule library and the driving environment of the vehicle during the training process until the number of training times reaches a preset number of times, and obtain a trained value network;
[0121] A decision-making instruction acquisition module 30, configured to input the driving environment of the vehicle at the current moment into the trained value network, and obtain the decision-making instruction output by the trained value network.
[0122] Further, in one embodiment, the construction module 10 is specifically configured to:
[0123] Obtain the correspondence between the historical driving environment and historical decision-making instructions of the vehicle;
[0124] Construct an expert rule library based on the correspondence between the historical driving environment and historical decision-making instructions of the vehicle, where the historical driving environment of the vehicle includes the lane where the vehicle is located and the obstacle conditions in each lane.
[0125] Further, in one embodiment, the preset number of times includes a first preset number of times and a second preset number of times. The training module 20 is configured to:
[0126] Obtain the driving environment of the vehicle at time t during the training process;
[0127] Input the driving environment at time t into the evaluation network and the expert rule base respectively;
[0128] Detect whether there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t;
[0129] If there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision instruction corresponding to the target historical driving environment and the first decision instruction output by the evaluation network;
[0130] If there is no target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the second decision instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm;
[0131] Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction, where time t + 1 is the next moment of time t;
[0132] Update the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value;
[0133] Use the driving environment at time t, the first decision instruction or the second decision instruction, the new reward value, and the driving environment at time t + 1 as a set of training data, and return to the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process until N sets of training data are obtained, where N is a positive integer;
[0134] Arbitrarily select a set of training data from the N sets of training data and update the weights of the evaluation network through the loss function;
[0135] Use the driving environment at time t + 1 as the driving environment at time t, and return to the step of inputting the driving environment at time t into the evaluation network and the expert rule base respectively until the number of times of updating the weights of the evaluation network is greater than or equal to the first preset number of times, and obtain the trained evaluation network. Assign the weights of the evaluation network updated every second preset number of times to the target network, where the second preset number of times is less than the first preset number of times.
[0136] Further, in an embodiment, the training module 20 is further configured to:
[0137] Detect whether the first decision instruction is the same as the target historical decision instruction;
[0138] If the detection result shows that the first decision instruction is the same as the target historical decision instruction, then calculate the sum of the reward value and the additional reward value to obtain a new reward value;
[0139] If the detection result shows that the first decision instruction is different from the target historical decision instruction, then calculate the difference between the reward value and the additional reward value to obtain a new reward value.
[0140] Further, in one embodiment, the training module 20 is further configured to:
[0141] Input the training data into the value network and the target network respectively, and obtain the predicted Q value output by the value network and the target Q value output by the target network corresponding to each group of training data;
[0142] Detect whether the predicted Q value output by the value network and the target Q value output by the target network corresponding to each group of training data are the same;
[0143] If the predicted Q value and the target Q value are the same, the weights of the value network remain unchanged;
[0144] If the predicted Q value and the target Q value are the same, use the gradient descent method to solve the loss function to obtain the new weights of the value network.
[0145] Among them, the function implementation of each module in the above vehicle behavior decision-making device corresponds to each step in the above vehicle behavior decision-making method embodiment, and its function and implementation process will not be elaborated here one by one.
[0146] In a third aspect, an embodiment of the present invention provides a vehicle behavior decision-making device, and this vehicle behavior decision-making device can be a device with data processing functions such as a personal computer (PC), a notebook computer, a server, etc.
[0147] Refer to Figure 6 , Figure 6This is a schematic diagram of the hardware structure of the vehicle behavior decision-making device involved in the solution of the embodiment of the present invention. In the embodiment of the present invention, the vehicle behavior decision-making device may include a processor 1001 (such as a Central Processing Unit, CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components; the user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard); the network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wireless-Fidelity, WI-FI interface); the memory 1005 may be a high-speed random access memory (random access memory, RAM), or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001. Those skilled in the art can understand that Figure 6 the hardware structure shown in
[0148] does not constitute a limitation to the present invention, and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 6 , Figure 6 in the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a vehicle behavior decision-making program. Among them, the processor 1001 may call the vehicle behavior decision-making program stored in the memory 1005 and execute the vehicle behavior decision-making method provided by the embodiment of the present invention.
[0149] In a fourth aspect, the embodiment of the present invention further provides a readable storage medium.
[0150] The vehicle behavior decision-making program is stored on the readable storage medium of the present invention. When the vehicle behavior decision-making program is executed by a processor, the steps of the vehicle behavior decision-making method as described above are realized.
[0151] Among them, the method realized when the vehicle behavior decision-making program is executed can refer to the various embodiments of the vehicle behavior decision-making method of the present invention, and will not be elaborated here.
[0152] It should be noted that in this document, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or system comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or system. Without further limitation, an element defined by the phrase "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or system comprising that element.
[0153] The serial numbers of the above embodiments of the present invention are only for description and do not represent the superiority or inferiority of the embodiments.
[0154] From the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions to enable a terminal device to execute the methods described in the various embodiments of the present invention.
[0155] The above are only the preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the description and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A vehicle behavior decision-making method, characterized in that, The vehicle behavior decision-making method includes: Construct an expert rule base based on the vehicle's historical driving environment and historical decision-making instructions, where the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane; Train the valuation network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches the preset number of times, and obtain the trained valuation network; Obtain the driving environment of the vehicle at time t during the training process; Input the driving environment at time t into the valuation network and the expert rule base respectively; Detect whether there is a target historical driving environment in the vehicle's historical driving environment included in the expert rule base that is the same as the driving environment at time t; If there is a target historical driving environment in the vehicle's historical driving environment included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision-making instruction corresponding to the target historical driving environment and the first decision-making instruction output by the valuation network; If there is no target historical driving environment in the vehicle's historical driving environment included in the expert rule base that is the same as the driving environment at time t, obtain the second decision-making instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm; Obtain the additional reward value, the reward value, and the driving environment at time t+1 after the vehicle executes the first decision-making instruction or the second decision-making instruction, where time t+1 is the next moment of time t; Update the reward value based on the first decision-making instruction, the target historical decision-making instruction, and the additional reward value to obtain a new reward value; Use the driving environment at time t, the first decision-making instruction or the second decision-making instruction, the new reward value, and the driving environment at time t+1 as a set of training data, and return to the step of detecting whether there is a target historical driving environment in the multiple vehicle historical driving environments included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process until N sets of training data are obtained, where N is a positive integer; Arbitrarily select a set of training data from the N sets of training data, and update the weights of the valuation network through the loss function; Use the driving environment at time t+1 as the driving environment at time t, and return to the step of inputting the driving environment at time t into the valuation network and the expert rule base respectively until the number of times of updating the weights of the valuation network is greater than or equal to the first preset number of times, and obtain the trained valuation network. Among them, assign the weights of the valuation network updated every second preset number of times to the target network, and the second preset number of times is less than the first preset number of times; Input the driving environment of the vehicle at the current moment into the trained valuation network to obtain the decision-making instruction output by the trained valuation network.
2. The vehicle behavior decision-making method according to claim 1, wherein The step of constructing an expert rule base based on the vehicle's historical driving environment and historical decision-making instructions includes: Obtain the correspondence between the vehicle's historical driving environment and historical decision-making instructions; Construct an expert rule base based on the correspondence between the vehicle's historical driving environment and historical decision-making instructions, where the vehicle's historical driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane.
3. The vehicle behavior decision-making method according to claim 1, wherein, The step of updating the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value includes: Detect whether the first decision instruction is the same as the target historical decision instruction; If the detection result is that the first decision instruction is the same as the target historical decision instruction, calculate the sum of the reward value and the additional reward value to obtain a new reward value; If the detection result is that the first decision instruction is not the same as the target historical decision instruction, calculate the difference between the reward value and the additional reward value to obtain a new reward value.
4. The vehicle behavior decision-making method according to claim 1, wherein, The step of arbitrarily selecting a set of training data from N sets of training data and updating the weights of the evaluation network through a loss function includes: Input the training data into the evaluation network and the target network respectively to obtain the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data; Detect whether the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data are the same; If the predicted Q value and the target Q value are the same, the weights of the evaluation network remain unchanged; If the predicted Q value and the target Q value are not the same, solve the loss function using the gradient descent method to obtain the new weights of the evaluation network.
5. A vehicle behavior decision-making device, characterized in that The vehicle behavior decision device includes: A construction module for constructing an expert rule base based on the historical driving environment and historical decision instructions of the vehicle, where the driving environment includes the lane where the vehicle is located and the obstacle conditions in each lane; A training module for training the evaluation network based on the expert rule base and the driving environment of the vehicle during the training process until the number of training times reaches a preset number, and obtaining a trained evaluation network; Obtain the driving environment of the vehicle at time t during the training process; Input the driving environment at time t into the evaluation network and the expert rule base respectively; Detect whether there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t; If there is a target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the target historical decision instruction corresponding to the target historical driving environment and the first decision instruction output by the evaluation network; If there is no target historical driving environment in the historical driving environment of the vehicle included in the expert rule base that is the same as the driving environment at time t, obtain the second decision instruction corresponding to the driving environment of the vehicle at time t during the training process based on the greedy algorithm; Obtain the additional reward value, the reward value, and the driving environment at time t + 1 after the vehicle executes the first decision instruction or the second decision instruction, where time t + 1 is the next moment of time t; Update the reward value based on the first decision instruction, the target historical decision instruction, and the additional reward value to obtain a new reward value; Taking the driving environment at time t, the first decision instruction or the second decision instruction, the new reward value, and the driving environment at time t+1 as a set of training data, return the step of detecting whether there is a target historical driving environment in the historical driving environments of multiple vehicles included in the expert rule base that is the same as the driving environment of the vehicle at time t during the training process, until N sets of training data are obtained, where N is a positive integer; Arbitrarily select a set of training data from the N sets of training data, and update the weights of the evaluation network through the loss function; Taking the driving environment at time t+1 as the driving environment at time t, return the step of inputting the driving environment at time t into the evaluation network and the expert rule base respectively, until the number of times of updating the weights of the evaluation network is greater than or equal to the first preset number of times, and obtain the trained evaluation network, where the weights of the evaluation network updated every second preset number of times are assigned to the target network, and the second preset number of times is less than the first preset number of times; The decision instruction acquisition module is used to input the driving environment of the vehicle at the current moment into the trained evaluation network to obtain the decision instruction output by the trained evaluation network.
6. The vehicle behavior decision-making device according to claim 5, wherein The training module is further used for: Inputting the training data into the evaluation network and the target network respectively, and obtaining the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data; Detecting whether the predicted Q value output by the evaluation network and the target Q value output by the target network corresponding to each set of training data are the same; If the predicted Q value and the target Q value are the same, the weights of the evaluation network remain unchanged; If the predicted Q value and the target Q value are different, use the gradient descent method to solve the loss function to obtain the new weights of the evaluation network.
7. A vehicle behavior decision-making device, characterized in that, The vehicle behavior decision device includes a processor, a memory, and a vehicle behavior decision program stored on the memory and executable by the processor. When the vehicle behavior decision program is executed by the processor, the steps of the vehicle behavior decision method according to any one of claims 1 to 4 are implemented.
8. A readable storage medium, characterized in that, The vehicle behavior decision program is stored on the readable storage medium. When the vehicle behavior decision program is executed by the processor, the steps of the vehicle behavior decision method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Lane changing decision model generation method and unmanned vehicle lane changing decision method and device
CN112937564A