An autonomous driving decision-making method based on an adaptive discriminator
By dynamically adjusting the discriminator discrimination ability in autonomous driving decisions, the problem of mismatch between the game difficulty of generator and discriminator is solved, and the balance of the progress speed of generator and discriminator is achieved and the imitation learning effect is improved.
Patent Information
- Application Number
- CN202310229472.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-10
AI Technical Summary
In autonomous driving decisions, the game difficulty of the generator in generative imitation learning does not match the discriminator's game, resulting in the network accuracy of the discriminator is much higher than that of the generator, and it is difficult for the generator to generate interactive trajectors that can deceive the discriminator, thereby affecting the effect of imitation learning.
By dynamically adjusting the discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's discriminator's adaptive ability is gradually weakened when the information entropy volume is too small. The threshold comparison method is used to select the threshold based on statistics to achieve the discriminator's adaptive ability.
The speed of progress between the generator and the discriminator is achieved, avoiding the problem of excessive or weak discriminator discriminator discriminator discriminatorial ability, and improving the generator's learning ability and imitation learning effect.
Smart Images

Figure CN116360429B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technology of vehicle autonomous driving, in particular to an autonomous driving decision-making method based on an adaptive discriminator. Background Art
[0002] Autonomous driving is one of the most cutting-edge research fields at present, which is of great significance for reducing traffic accidents, improving traffic efficiency, reducing environmental pollution and liberating human labor. A complete and mature autonomous driving technology includes three modules: environmental perception, decision-making and planning, and motion control. The environmental perception module collects road traffic information, the decision-making and planning module makes decisions based on the collected environmental information, and the motion control module makes corresponding control actions according to the decision instructions of the upper layer.
[0003] There are various solutions to the autonomous driving decision-making problem involved in the decision-making and planning module. The main research directions include rule-based decision-making, reinforcement learning-based methods, and imitation learning-based methods. Rule-based decision-making cannot be applied to complex and changeable traffic environments, and reinforcement learning-based methods face difficulties when migrating to real environments. Therefore, imitation learning-based methods remain the key content of researching autonomous driving decision-making problems.
[0004] Imitation learning is mainly divided into three methods: behavior cloning, inverse reinforcement learning, and generative adversarial imitation learning. Behavior cloning is a supervised learning method, and the data set is human expert driving data. The environment (traffic scene) is used as the input, and the decisions made by human experts (such as steering wheel angle, pedal angle, etc.) are used as labels. Inverse reinforcement learning needs to first learn the reward function and then perform forward reinforcement learning based on the reward function. Generative adversarial imitation learning can interact with the environment, and the interaction trajectories generated by generative adversarial imitation learning gradually approach the interaction trajectories of human experts, so as to achieve the purpose of imitating the decision-making strategies of human experts.
[0005] Generative adversarial imitation learning consists of a human expert experience pool, a generator, and a discriminator. The generator is the decision-making strategy network to be trained, which interacts with the environment to generate interaction trajectories. The generated interaction trajectories and the interaction trajectories in the human expert experience pool are jointly input into the discriminator, and the discriminator will distinguish whether the interaction trajectories belong to human experts or the generator. The generator is responsible for generating interaction trajectories that are sufficient to deceive the discriminator, and the discriminator is responsible for distinguishing the data source. The network training of both belongs to a game process.
[0006] The problem of autonomous driving decision-making is a complex problem in a high-dimensional space. It is very difficult for the generator to interact with the environment and generate interaction trajectories similar to those of human experts that are sufficient to deceive the discriminator, while it is very simple for the discriminator to distinguish the data sources of the interaction trajectories. The training difficulty of the generator network is much higher than that of the discriminator network, so the accuracy of the discriminator network will be much higher than that of the generator. No matter how the generator is improved, the generated interaction trajectories can be easily distinguished by the recognizer, resulting in the discriminator being unable to learn effective gradient information. The game difficulty between the generator and the discriminator does not match, and it is very difficult for generative adversarial imitation learning to achieve significant results.
[0007] The progress speeds of the generator and the discriminator are inconsistent, which will not only lead to an increase in the training cycle, but may even cause problems of non-convergence. Currently, the main research directions for this problem mainly include: improving the generator network to enhance its learning ability, and reducing the update frequency of the discriminator network to reduce its discrimination ability. Although too high discriminator accuracy will generate gradients with lower information content, which is not conducive to the final effect of imitation learning, a discriminator with too weak discrimination ability will also hinder the learning ability of the generator. Therefore, how to balance the game problem between the generator and the discriminator still needs further research.
[0008] The key network for generative adversarial imitation learning in autonomous driving decision-making is the generative adversarial network. How to solve the balance problem between the generator and the discriminator of this network is related to whether we can quickly and stably learn the decision-making strategies of human experts. Simply increasing the difficulty of the discriminator or delaying the update speed of the discriminator network is unreasonable. Too weak discrimination ability of the discriminator will hinder the progress of the generator; at the same time, too strong discrimination ability of the discriminator will cause the generator to be unable to obtain effective gradient information. Therefore, it is very important to design a reasonable algorithm to dynamically adjust the discrimination ability of the discriminator so that it is always within a suitable range, which can not only correctly distinguish the data sources, but also not feedback ineffective gradient information. Summary of the Invention
[0009] To solve the above problems, the present invention provides an autonomous driving decision-making method based on an adaptive discriminator, which dynamically adjusts the discrimination ability of the discriminator so that the discriminator has an adaptive ability.
[0010] To achieve the above purpose, the basic idea of the present invention is: by restricting the volume of the information entropy flowing into the discriminator, the training difficulty of the discriminator is increased; at the same time, when the volume of the information entropy flowing into the discriminator is too small, the information flow restriction will be gradually weakened, so that the information obtained by the discriminator always remains at a reasonable level. To achieve the purpose of dynamic adjustment, the present invention adopts a method of threshold comparison, and the volume of the information entropy flowing in fluctuates around a certain threshold, and the selection of the threshold is based on statistics.
[0011] The technical route of the present invention is as follows: An autonomous driving decision-making method based on an adaptive discriminator, comprising the following steps:
[0012] A. Collect interaction trajectories
[0013] Human experts drive a test vehicle on a real traffic road, make judgments and decisions based on the information obtained according to their own experience, interact with the road environment, and obtain human expert interaction trajectories; Let the generator network be G(x), that is, the decision-making network to be learned; its input is the information required to make decisions, that is, the road traffic scene pictures taken by the camera; the output is the decision-making quantity, that is, the steering wheel angle and the pedal angle. Use the initialized generator network to interact with the environment to obtain the generator network interaction trajectories;
[0014] B. Determine the data source
[0015] Let the discriminator network be D(x). The network structure of the discriminator network needs to be designed according to the generator network. Its input is the set of human expert interaction trajectories and generator network interaction trajectories, that is, a set of information required to make decisions and the decision-making quantities obtained according to this information; its output is the probability value, that is, the probability that the discriminator determines that the input data comes from the human expert interaction trajectories;
[0016] C. Dilute the information entropy volume
[0017] Dilute the information entropy volume flowing into the discriminator by adding an information entropy regularization term to the objective function. Define the first half of the discriminator network structure as D 1 (x). After the input x passes through D 1 (x), it becomes z. The second half is defined as D 2 (z); The volume of the information entropy flowing into the discriminator is defined as M, which is characterized by the upper bound of the mutual information, that is:
[0018] M = ∫p(x)KL(p(z|x)||r(z))dx
[0019] Among them, p(x) is the distribution of the data input into the discriminator, r(z) is the normal distribution; the information threshold is defined as I, KL is the KL divergence, also called the information divergence, which is a measure to describe the difference between two probability distributions. p(z|x) is the probability distribution of z under the condition that x is determined. || is the symbol in the KL divergence formula:
[0020] KL(P||Q) = ∫P(x)log(P(x) / Q(x))dx
[0021] The objective function obtained therefrom is as follows:
[0022]
[0023] where p G (x) is the distribution of the generator. Writing the above equation in the form of an expectation gives the following equation:
[0024]
[0025] In the above equation, λmax((M - I), 0) is a regularization term for dynamically adjusting the discriminative ability of the discriminator;
[0026] where denotes maximizing with respect to the data terms related to G(x), denotes minimizing with respect to the data terms related to D 1 (x), D 2 (z). p * (x) represents the probability distribution of the interaction trajectories of human experts, and p G (x) represents the probability distribution of the interaction trajectories generated by the generator network. λ is the penalty coefficient, which characterizes the proportion of the loss caused by the inflow of information volume exceeding the threshold.
[0027] D, update network
[0028] Calculate the loss through the above objective function and update the generator network and the discriminator network. Use the updated generator network and discriminator network to repeat the training in step A until the recognition accuracy of the discriminator stabilizes at 49% - 51%, and then end.
[0029] Furthermore, the specific discriminant formula for the accuracy rate in step D is:
[0030]
[0031] where TRUE_TRACK is the number of scenarios in 5 interaction trajectories of human experts, each scenario corresponding to an action, and FAKE_ACTION represents the number of actions that are judged by the discriminator network to be made by the generator network among these actions; FAKE_TRACK is the number of scenarios in 1 interaction trajectory of the generator network, each scenario corresponding to an action, and TRUE_ACTION is the number of actions that are judged by the discriminator to be made by human experts among these actions.
[0032] Compared with the prior art, the present invention has the following advantages:
[0033] 1. In the present invention, when the information flowing into the discriminator is greater than the information threshold, the loss function will increase, and the information volume passing through will be reduced when updating the network; while when the information flowing into the discriminator is less than the threshold, this loss does not work, and the discriminator will automatically increase the information volume passing through in order to more accurately judge the data source. In this way, the purpose of dynamic adjustment can be achieved.
[0034] 2. Instead of simply increasing the learning difficulty of the discriminator, the present invention adopts a method of dynamically adjusting the discrimination ability of the discriminator, which can not only balance the progress speed of the discriminator and the generator, but also prevent the discrimination ability of the discriminator from being too low to hinder the learning of the strategy, thus solving the problem of unbalanced game between the generator and the discriminator.
[0035] 3. The present invention reasonably sets the upper limit of the amount of information entropy flowing into the discriminator, so that the regularization term can screen out the truly discriminative information, and at the same time, this information is sufficient to represent the data distribution characteristics. This not only reduces the process of the network processing and analyzing unnecessary information, but also speeds up the convergence speed of the network.
[0036] 4. The present invention divides the discriminator into two parts and restricts the amount of information entropy flowing into the front and rear parts. Compared with the traditional method of encoding information using an encoder and then passing it into the discriminator, it can reduce the network size, achieve the purpose of network lightweight, and also has a good effect on accelerating the network convergence speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a flowchart of the present invention, and the gradient information in the figure is the gradient generated during the network calculation process.
[0038] Figure 2 is a local network of the generator network, which is used to process image information to obtain a one-dimensional vector.
[0039] Figure 3 is the recurrent network part of the generator network, which can make full use of the historical information during the driving process of the autonomous vehicle, and then analyze the current information for decision-making.
[0040] Figure 4 are the network parameters of the generator network.
[0041] Figure 5 is the overall architecture of the generator network.
[0042] Figure 6 is the overall architecture of the discriminator network. The plus sign in the figure indicates concatenation. DETAILED DESCRIPTION OF THE INVENTION
[0043] The present invention will be further described below with reference to the accompanying drawings.
[0044] Before conducting the experiment, it is necessary to set up the experimental environment first: This invention uses Carla as the experimental platform and collects the interactive trajectories of human experts and the interactive trajectories of the generator network by controlling the test vehicle. The test vehicle can record the accelerator pedal angle, brake pedal angle, and steering wheel angle of the vehicle at each time step. An RGB three-channel camera is installed on the test vehicle, which can capture the road conditions at each time step. The installation position is the midpoint of the upper edge of the front windshield, and the resolution of the collected images is 800*600. The experimental road is a 3-km straight road. There will be pedestrians and various vehicles participating in the traffic on the experimental road, and the traffic density is set to 10 vehicles (persons) / km.
[0045] First, conduct Figure 1 the collection of the interactive trajectories of human experts. The test vehicle is manipulated by a human expert to drive on the set experimental road. The sampling frequency of the test vehicle is 30 hz. The image information collected each time is denoted as s. The angles of the accelerator pedal and the brake pedal are normalized to [-1, 1], and the steering wheel angle is normalized to [-1, 1]. The decision value information collected each time is a tuple composed of the pedal angle and the steering wheel angle, denoted as a. Sampling starts when the human expert controls the test vehicle and terminates when the test vehicle has traveled 2 km or when the test vehicle collides or drives out of the road. Each time sampling starts, the test vehicle is randomly deployed to a certain position within the first 1 km of the road. The information set collected by the test vehicle during the time interval from the start of sampling to the end of sampling is called a trajectory, denoted as τ. The human expert needs to collect 100 trajectories, and these 100 trajectories are denoted as χ.
[0046] χ = {τ (1) , τ (2) , ··· τ (100)}; τ = [s 1 , a 1 , s 2 , a 2 , ···, s m , a m
[0047] Among them, the superscript of τ indicates which trajectory this is, the subscript of s indicates which frame of image is collected, the subscript of a indicates the corresponding action under this frame of image, and the value of m depends on the speed of the test vehicle, traffic conditions, and sampling termination conditions (the test vehicle travels 2 km or the test vehicle collides or drives out of the road), etc. If it is stipulated that the speed of the test vehicle is above 30 km / h, then the value range of m is [0, 7200].
[0048] After collecting the trajectories of human experts, continue to collect Figure 1 The interaction trajectory of the generator network. The method of collecting the interaction trajectory of the generator network is the same as that of collecting the interaction trajectory of human experts, except that it is no longer a human expert who manipulates the test vehicle, but the generator network. Different from the interaction trajectory of human experts, each time the interaction trajectory of the generator network is collected, one trajectory is collected. Next, build the generator network. The input of the generator network is image information, which needs to be processed by a convolutional network, and the final output is the pedal angle and the steering wheel angle, that is, a one-dimensional vector containing two elements. Since autonomous driving is a partially observable Markov process, a recurrent neural network is introduced into the generator network. The image information is first processed by 5 convolutional layers, 3 pooling layers and 2 fully connected layers to obtain a one-dimensional vector containing 128 elements. The specific structure is as Figure 2 shown. The recurrent neural network contains 3 recurrent units and 3 fully connected layers. The parameters are shared among the recurrent units, and the parameters are also shared directly among the fully connected layers. The memory unit h is used to store historical information and consists of 64 elements. The specific structure is as Figure 3 shown, Figure 3 where x in t is the data obtained after processing the image information. When training the generator network, the generator will obtain the current scene s t-1 and the previous two scenes s t-2 , s 1 , x 2 and x 3 are calculated through the convolutional network and the fully connected network, and then y 3 is the action to be taken in the current scene after passing through the recurrent neural network. Figure 4 are the parameters of the generator network. In the figure, CNN represents the convolutional layer, MAXPOOL represents the max pooling layer, FC represents the fully connected layer, RNN represents the recurrent neural network layer, k represents the size of the convolutional kernel, s is the stride, and p is the extended boundary. In the RNN, there is also an input h, which represents the memory unit. The RNN will also output the updated memory unit each time it outputs. Figure 5 shows the overall architecture of the generator network.
[0049] Before starting the training, it is necessary to build Figure 1 the discriminator network in Figure 6 shown. The input of the discriminator network has two parts. One part is the image information collected by the test vehicle, and the other part is the action (pedal angle and steering wheel angle) made by the test vehicle in this scene. The processing network of the image information is similar to that of the generator network, and the generator network can be truncated at the recurrent neural network; the processing of the action is completed through two fully connected layers. The vectors obtained after processing the two parts of the information are spliced together, and then a vector containing two elements is output through two fully connected layers, respectively representing the probability that the discriminator determines that the action belongs to a human expert or the generator network. The specific structure of the network is as
[0050] After the generator network and the discriminator network are built, the network can be trained according to the collected trajectory information. For each trajectory collected by the generator network, a trajectory is randomly selected from the human expert interaction trajectories. Let the selected human expert interaction trajectory be:
[0051]
[0052] The trajectory collected by the generator network is:
[0053]
[0054] According to
[0055]
[0056] Perform gradient descent to update the generator network.
[0057] According to
[0058]
[0059] Perform gradient descent to update the discriminator network.
[0060] When the target recognition accuracy is reached, terminate the network update. At this time, the discriminator is no longer used, and the autonomous driving vehicle makes a decision based on the operation result of the generator.
[0061] The present invention is not limited to this embodiment, and any equivalent conceptions or changes within the technical scope disclosed in the present invention are included in the protection scope of the present invention.
Claims
1. An autonomous driving decision-making method based on an adaptive discriminator, characterized in that: It includes the following steps: A. Collect interaction trajectories Human experts drive a test vehicle on a real traffic road, make judgments and decisions based on their own experience through the information obtained, interact with the road environment, and obtain human expert interaction trajectories; Let the generator network be G(x), that is, the decision-making network to be learned; its input is the information required for making decisions, that is, the pictures of the road traffic scene taken by the camera; the output is the decision-making quantity, that is, the steering wheel angle and the pedal angle; Use the initialized generator network to interact with the environment to obtain the generator network interaction trajectory; B. Judge the data source Let the discriminator network be D(x). The network structure of the discriminator network needs to be designed according to the generator network. Its input is the set of human expert interaction trajectories and generator network interaction trajectories, that is, a set of information required for making decisions and the decision-making quantities obtained based on this information; its output is a probability value, that is, the probability that the discriminator judges that the input data comes from the human expert interaction trajectory; C. Dilute the information entropy volume Dilute the information entropy volume flowing into the discriminator by adding an information entropy regularization term to the objective function; define the first half of the discriminator network structure as D 1 (x), the input x passes through D 1 (x) and becomes z, and the second half is defined as D 2 (z); the volume of the information entropy flowing into the discriminator is defined as M, which is characterized by the upper bound of the mutual information, that is: M = ∫p(x)KL(p(z|x)||r(z))dx where p(x) is the distribution of the data input to the discriminator, and r(z) is the normal distribution; the information threshold is defined as I, and KL is the KL divergence, also called the information divergence, which is a measure to describe the difference between two probability distributions; p(z|x) is the probability distribution of z under the condition that x is determined; || is the symbol in the KL divergence formula: KL(P||Q) = ∫P(x)log(P(x) / Q(x))dx The obtained objective function is as follows: where p G (x) is the distribution of the generator. Writing the above equation in the form of an expectation gives the following equation: λmax((M - I), 0) in the above formula is a regularization term for dynamically adjusting the discriminant ability of the discriminator; Among them denotes maximizing the data items related to G(x), denotes minimizing the data items related to D 1 (x), D 2 (z); p * (x) represents the probability distribution of the interaction trajectories of human experts, and p G (x) represents the probability distribution of the interaction trajectories generated by the generator network; λ is the penalty coefficient, characterizing the proportion of the loss caused by the inflow of information volume exceeding the threshold D. Update the network Calculate the loss through the above objective function and update the generator network and the discriminator network. Use the updated generator network and discriminator network to go back to step A for repeated training until the recognition accuracy of the discriminator stabilizes at 49% - 51%, and then end.
2. The autonomous driving decision-making method based on an adaptive discriminator according to claim 1, characterized in that: The specific discriminant formula for the accuracy rate in step D is: where TRUE_TRACK is the number of scenarios in 5 human expert interaction trajectories, each scenario corresponds to an action, and FAKE_ACTION represents the number of actions that are judged by the discriminator network to be made by the generator network among these actions; FAKE_TRACK is the number of scenarios in 1 interaction trajectory of the generator network, each scenario corresponds to an action, and TRUE_ACTION is the number of actions that are judged by the discriminator to be made by the human expert among these actions.
Citation Information
Patent Citations
Segmentation loss-based generative adversarial network method
CN108665058A
Image segmentation network training method and device, image segmentation method and device and storage medium
CN111199550A