End-to-end automatic driving decision-making method and system based on brain-like reinforcement learning
By adopting brain-like reinforcement learning methods in end-to-end autonomous driving systems, modeling the return distribution and adding large field-angle visual information, the problems of high energy consumption and poor robustness in traditional technologies are solved, and low energy consumption and high robustness autonomous driving decisions are achieved.
Patent Information
- Application Number
- CN202510084169.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional end-to-end autonomous driving technology based on reinforcement learning has problems of high energy consumption and poor robustness, making it difficult to make stable decisions in complex traffic environments.
The brain-like reinforcement learning method is adopted to directly model the return distribution rather than the expectation, increase the fine-grained visual information of the large field of view angle, and divide the training process into two stages through course learning. First train with small field of view angle information to convergence, and then continue training with large field of view angle information.
It reduces the energy consumption of reinforcement learning in end-to-end autonomous driving, improves the robustness of the system in complex environments, and ensures that the vehicle can make stable and safe decisions.
Smart Images

Figure CN119928898A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of autonomous driving, and specifically relates to an end-to-end autonomous driving decision-making method and system based on brain-like reinforcement learning. Background Art
[0002] In recent years, as autonomous driving technology has become increasingly mature, autonomous vehicles have had a profound impact on urban transportation systems and urban planning, and are expected to greatly improve travel safety and efficiency.
[0003] Although traditional rule-based autonomous driving systems are fast and highly interpretable, they require a lot of manual parameter adjustments, making it difficult to adapt to complex traffic environments and seriously lacking in generalization capabilities. End-to-end autonomous driving technology based on reinforcement learning can continuously interact with the environment and directly map perception information to control signals, which not only simplifies the process of the autonomous driving system, but also improves its generalization ability. However, this method has high energy consumption due to the use of deep neural networks. Applying brain-like reinforcement learning to end-to-end autonomous driving can reduce the energy consumption of the autonomous driving decision-making system and reduce the annual carbon emissions of automobiles to the atmosphere. However, since the current brain-like reinforcement learning only models the scalar return expectation, and the low-precision encoding of the pulse neural network will lose some perception information, brain-like reinforcement learning is insufficient in information in end-to-end autonomous driving applications, so that the trained control strategy cannot make stable decisions. Therefore, there is an urgent need for a convergent method to help end-to-end autonomous driving based on brain-like reinforcement achieve low energy consumption and high robustness. Summary of the invention
[0004] The purpose of the present invention is to overcome the problems of high energy consumption and poor robustness in traditional end-to-end autonomous driving technology based on reinforcement learning, apply brain-like reinforcement learning to autonomous driving, and propose an end-to-end autonomous driving decision method and system based on brain-like reinforcement learning. The brain-like reinforcement learning involved in the present invention directly models the reward distribution instead of the traditional modeling of reward expectation to construct the loss function. The reward distribution contains more information related to the decision, which is conducive to improving the robustness of the autonomous driving system in a complex environment; adding fine-grained visual information with a large field of view angle to the input of brain-like reinforcement learning alleviates the loss of perceptual information caused by the low-precision encoding of the pulse neural network; fine-grained visual information with a large field of view angle is conducive to reducing the visual blind spot, but it is difficult to converge directly with fine-grained visual information with a large field of view angle. The present invention adopts the idea of curriculum learning and divides the training into two stages. In the first stage, fine-grained visual information with a small field of view angle is first trained until convergence, and in the second stage, fine-grained visual information with a large field of view angle is used to continue training on the basis of the model trained in the first stage until convergence.
[0005] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0006] In a first aspect, the present invention proposes an end-to-end autonomous driving decision-making method based on brain-like reinforcement learning, comprising:
[0007] Given the vehicle visual information, navigation information and control information, generate its fine-grained information based on the visual information;
[0008] Fusion of the visual information of the last few time steps and its fine-grained information, navigation information, and control information as the observation of the current time step t , extracting low-dimensional representations of observations using a group of encoders;
[0009] Reinforcement learning is used to train a policy network based on a spiking neural network. The policy network takes the observed low-dimensional representation as input during the training phase and uses μ, σ RL , σ IL is the output, where μ, σ RL Represents the mean and variance learned by the policy network during reinforcement learning, and generates actions based on the mean and variance; σ IL represents the variance learned by the policy network during imitation learning, where imitation learning refers to imitating the actions of the expert;
[0010] During the training process, we continuously RL The relationship between the threshold and the selected execution generation action or Expert Action Calculate the reward r based on the action performed t And get the observation o of the next time step t+1 , the quintuple Stored in the data container; when the parameter update time requirements of reinforcement learning are met, the reinforcement learning loss function is constructed based on the return distribution to complete the training of the policy network and encoder group;
[0011] Deploy the trained policy network and encoder group on an autonomous vehicle to achieve autonomous driving tasks.
[0012] As a preferred embodiment of the present invention, the training process is divided into two stages of training using a course learning method. In the first stage of training, the vehicle visual information is a small field of view RGB image not higher than 90 degrees. In the second stage of training, the vehicle visual information is a large field of view RGB image higher than 90 degrees. The second stage of training continues the training based on the first stage of training.
[0013] In a second aspect, the present invention provides an end-to-end autonomous driving decision-making system based on brain-like reinforcement learning, which is used to implement the above-mentioned end-to-end autonomous driving decision-making method.
[0014] The beneficial effects of the present invention are:
[0015] (1) This invention applies brain-inspired reinforcement learning to end-to-end autonomous driving for the first time, reducing the energy consumption of reinforcement learning in end-to-end autonomous driving.
[0016] (2) The present invention introduces reward distribution into reinforcement learning and integrates fine-grained visual information with a large field of view into the input. Compared with the expected reward, the reward distribution contains more information about the consequences of the agent's decision. The fine-grained visual information with a large field of view enhances the agent's perception of the environment. This rich information helps the vehicle make stable decisions, making it more suitable for complex scenarios such as end-to-end autonomous driving.
[0017] (3) In order to use visual information with a large field of view for training, the present invention also uses curriculum learning to divide the training process into two stages. The first stage uses visual information with a small field of view for training, and the second stage continues to use visual information with a large field of view for training based on the training in the first stage. The policy network trained by the method of the present invention is deployed on a real vehicle to verify its effectiveness in the real world. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a framework diagram of an end-to-end autonomous driving decision-making method based on brain-like reinforcement learning;
[0019] Figure 2 This is a schematic diagram of deploying the policy network on a real vehicle;
[0020] Figure 3 These are the small field of view RGB images and large field of view RGB images used in the course learning;
[0021] Figure 4 It is a performance comparison chart between the present invention and the traditional method. DETAILED DESCRIPTION
[0022] The present invention is further described and illustrated below in conjunction with specific embodiments. The embodiments are merely exemplary of the present disclosure and do not define the scope of limitation. The technical features of each embodiment of the present invention may be combined accordingly without conflicting with each other.
[0023] The accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0024] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the steps. For example, some steps may be decomposed, while some steps may be combined or partially combined, so the actual execution order may change according to the actual situation.
[0025] The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning proposed in this invention is as follows: Figure 1 As shown, it mainly includes the following steps:
[0026] Step 1: In a simulated environment, obtain the current information of the vehicle, including environmental visual information (RGB image with a small field of view), navigation information, speed and steering, which are represented as I, N, and V, respectively, where I is the RGB image, N is the navigation information, and V is the control information including speed and steering.
[0027] Step 2: Get the fine-grained information of the RGB image, specifically:
[0028] The RGB image is segmented using image segmentation technology and merged with the lane lines in the image detected using lane line detection technology to obtain an image, which is recorded as In this embodiment, image segmentation technology and lane line detection technology can be implemented using existing technologies. For example, image segmentation uses a pulsed DeepLab method, which is suitable for segmenting large field of view images; lane detection uses a pulsed CondNet method, which can detect lane instances and dynamically predict the line shape of each instance.
[0029] Step 3: Fuse the information of the last two time steps to get the observation of the current time step
[0030] Step 4: Since o t The dimension is very high, and its low-dimensional representation needs to be extracted as input. In this embodiment, the encoder network is used to get The corresponding low-dimensional representation is denoted as Concatenate them to get the concatenated low-dimensional representation For convenience, the following Co-encoder Group Right now
[0031] In this embodiment, Use image encoders, such as CNN convolutional neural networks or ResNet networks; f n ,f v Adopt an MLP-based network structure.
[0032] Step 5: Let the policy network be π φ ,Will Modeled as a parameterized hyperbolic tangent Gaussian function, the given low-dimensional representation h t As a policy network π φ Input, action It can be obtained by sampling the hyperbolic tangent Gaussian function, that is, The action Through π φ The parameterized mean μ and variance σ of the output RL It is calculated that the hyperbolic tangent Gaussian function converts π φ The output of is limited to the range [-1,1]; ∈~N 0,1 is the standard normal distribution; f φ (∈,h t ) represents the hyperbolic tangent Gaussian function. The action space is set as the executable actions of the vehicle, which can be combined into left turn, right turn, forward, etc. according to the speed and steering output by the strategy network.
[0033] The policy network can generate three values: μ, σ RL , σ IL ; σ RL represents the variance learned by the policy network during reinforcement learning; σ IL It represents the variance learned by the policy network in imitation learning. Imitation learning updates the policy network by learning expert actions, which is used to accelerate the training of reinforcement learning. According to the usual practice, expert actions are only provided by the simulator during training in a simulation environment, and its principle will not be described in detail here.
[0034] In the present invention, the strategy network adopts a spiking neural network composed of leaky integral-and-fire (LIF) neurons and non-spiking neurons, wherein the non-spiking neurons serve as the output layer of the spiking neural network. The non-spiking neurons here are obtained by setting the threshold to infinity on the basis of LIF neurons. The dynamic equation of the non-spiking neurons is:
[0035]
[0036] Here, V t represents the presynaptic membrane potential, V reset represents the resting membrane potential (usually 0), X t Indicates the input pulse signal, H t represents the postsynaptic membrane potential, and τ represents the membrane potential time constant. Non-spiking neurons do not trigger spikes, and use membrane potential to directly represent continuous values. In the present invention, non-spiking neurons are used as the output layer to directly output continuous values.
[0037] Step 6: To speed up training, obtain expert actions through the simulator According to σ RL, the relationship between the threshold and the selection execution still Right now:
[0038]
[0039] According to action a t Get the current reward r t And the next observation o t+1 .
[0040] Step 7: 5-tuple Store in data container middle.
[0041] Step 8: Repeat steps 1-7 multiple times to collect a certain amount of five-tuple data. When the parameter update time requirements of reinforcement learning are met, execute the following steps 9-15, and continue to repeat steps 1-7 to continuously update the data in the data container.
[0042] Step 9: From the container Randomly sample a batch of data from the t , o t+1 Encoding gets h t and h t+1 .
[0043] Step 10: Calculate the value network loss based on the reinforcement learning results to update the value network θ and the encoder group
[0044] The process of updating the value network θ is expressed as: Update encoder group The process uses conventional optimization techniques and will not be described in detail.
[0045] Where: J z (θ) represents the value network loss under reinforcement learning, θ represents the value network parameter, β θ represents the value network learning rate, represents the value network gradient, Express expectations, represents the output action of the policy network at the current time step t, represents the output action of the policy network at time step t+1, r t represents the reward at the current time step t, π φ′ represents the target policy network, h t+1 represents the encoding result of the t+1 time step, Show obedience A random variable with a distribution, Represents a given state-action pair The return distribution obtained by the target value network under the condition, p(.) represents the probability density function, Represents a given target policy network π φ′ right Perform a Bellman update, Represents a given state-action pair The distribution of rewards obtained by the conditional value network;
[0046] To J z (θ) to find the gradient, the formula is as follows:
[0047]
[0048] Where: express The distribution of γ represents the discount factor, α represents the temperature parameter, is a random variable representing the target value, It means that it can be mathematically defined as, Represents the policy network π φ Given a state h t+1 Get Action The probability of Represents state-action pairs The reward distribution of φ represents the policy network, It represents the gradient operation on θ.
[0049] Step 11: Calculate the policy network loss based on the reinforcement learning results to update the policy network parameters π φ :
[0050]
[0051] The process of updating the policy network parameter φ is expressed as:
[0052] Where: J π (φ) represents the policy network loss under reinforcement learning, φ represents the policy network parameter, β φ represents the policy network learning rate, represents the policy network gradient; Express expectations, represents the output action of the policy network at the current time step t, π φ (·|h t ) indicates that given h t The probability density of the action under the condition of express expectations; express The value of , α represents the temperature parameter, Indicates that at a given h t Action under the conditions probability.
[0053] To J π (φ) to find the gradient, the formula is as follows:
[0054]
[0055] Where: Express Perform gradient calculations, represents the gradient operation on φ, f φ (∈,h t ) represents the tanh-Gaussian sampling function.
[0056] Step 12: Calculate the policy network loss based on the imitation learning results to update the policy network φ and encoder group
[0057]
[0058] The process of updating the value network parameter θ is expressed as: Update encoder group The process uses conventional optimization techniques and will not be described in detail.
[0059] Where: J IL (φ) represents the policy network loss under imitation learning, The mean is μ and the standard deviation is σ IL Gaussian distribution.
[0060] Step 13: Calculate the temperature parameter loss based on the reinforcement learning results to update the temperature parameter α:
[0061]
[0062] Where: J(α) represents the temperature parameter loss, Indicates that at a given h t Action under the conditions The probability of φ (·|h t ) indicates that given h t The probability density of the action under the condition of represents the expected entropy.
[0063] Step 14: Soft update process of target policy network and target value network:
[0064] After obtaining the above parameters θ and φ, the target strategy network parameter φ′ and the target value network parameter θ′ are updated according to θ′←τθ+(1τ)θ′, φ′←τφ+(1τ)φ′, where τ represents the update rate.
[0065] Step 15: Repeat steps 9-14 until the loss function converges, and then proceed to step 16.
[0066] Step 16: Obtain the current information of the vehicle, including environmental visual information (large field of view RGB image), navigation information, speed and steering, which are represented as I, N, and V respectively. These parameters are consistent with the parameter definitions above.
[0067] Step 17: Repeat steps 2-15 and use the same method to complete the policy network π based on the large field of view RGB image φ and value network training.
[0068] Step 18: Deploy the trained policy network φ and encoder group on the actual vehicle To achieve autonomous driving tasks;
[0069] Specifically, Figure 2 As shown in the figure, in the real environment, the current information of the vehicle is obtained, including environmental visual information (large field of view RGB image), navigation information, speed and steering, which are represented as I, N, and V respectively; the fine-grained information of the large field of view RGB image is obtained, which is recorded as These parameters are consistent with the above parameter definitions.
[0070] Using encoder group Get the current time step observation The encoding result of It is used as the input of the policy network φ, based on the parameterized mean μ and variance σ of the output RL The action is obtained by sampling the hyperbolic tangent Gaussian function This action is used to interact with the real environment.
[0071] In order to verify the effect of the present invention, this embodiment performs a simulation experiment, which is carried out on the Town1 map of CARLA 0.9.13. The Town 1 map has river bridges, T-junctions, various buildings, pedestrians and real vehicles, simulating various scenarios of urban traffic. The experiment was first trained in a collision environment of a daytime scene, and then evaluated on the NoCrash benchmark under other weather conditions, including rain at night, rain during the day, and fog at night. At the same time, medium-density and high-density traffic flows are considered according to the number of vehicles and pedestrians, where medium-density traffic flows have 200 vehicles and 100 pedestrians, and high-density traffic flows have 300 vehicles and 150 pedestrians; in the first and second stages of training, visual information with a field of view of 90° and 130° are used respectively, and the sample images are for example Figure 3 Finally, the trained control strategy is deployed on the real vehicle AutoBots-W1.
[0072] The method of the present invention is denoted as BiRLAD. In order to demonstrate the performance of BiRLAD, it is compared with RLfOLD. RLfOLD is currently a reinforcement-based end-to-end autonomous driving SOTA method; Figure 4 The average and maximum scores of BiRLAD and RLfOLD were compared. The average and maximum scores of BiRLAD were 6930 and 11112, respectively, while the average and maximum scores of RLfOLD were 4502 and 6316, respectively. It can be seen that BiRLAD is significantly better than RLfOLD.
[0073] By comparing the training results of the first and second stages in BiRLAD, the impact of a large field of view on end-to-end autonomous driving was verified. In the first stage, BiRLAD was trained using images with a 90° field of view, and in the second stage, images with a 130° field of view were used for further training. The average rewards, number of collisions, success rates, and trajectory completion rates for medium-density and high-density traffic flows under different weather conditions in the two stages were statistically analyzed, as shown in Table 1. It can be seen that the number of collisions per kilometer in the second stage was significantly lower than that in the first stage, while the average rewards, success rates, and trajectory completion rates were higher. Experiments show that visual information with a large field of view can effectively avoid dangerous behaviors and improve robustness.
[0074] Table 1
[0075]
[0076] The real vehicle AutoBots-W1 can drive stably in a campus environment, demonstrating the effectiveness of BiRLAD in the real world.
[0077] Based on the same inventive concept, an embodiment of the present invention further provides an end-to-end autonomous driving decision system based on brain-like reinforcement learning, including:
[0078] A data acquisition module, which is used to acquire given vehicle visual information, navigation information and control information, and generate fine-grained information based on the visual information;
[0079] The observation-encoding module is used to fuse the visual information of the last few time steps and its fine-grained information, navigation information, and control information as the observation of the current time step. t , extracting low-dimensional representations of observations using a group of encoders;
[0080] A reinforcement learning module based on reward distribution is used to train a policy network based on a pulse neural network using reinforcement learning. The policy network takes the observed low-dimensional representation as input during the training phase and uses μ, σ RL , σ IL is the output, where μ, σ RL Represents the mean and variance learned by the policy network during reinforcement learning, and generates actions based on the mean and variance; σ IL represents the variance learned by the policy network during imitation learning, where imitation learning refers to imitating the actions of the expert;
[0081] During the training process, we continuously RL The relationship between the threshold and the selected execution generation action or Expert Action Calculate the reward r based on the action performed t And get the observation o of the next time step t+1 , the quintuple Stored in the data container; when the parameter update time requirements of reinforcement learning are met, the reinforcement learning loss function is constructed based on the return distribution to complete the training of the policy network and encoder group;
[0082] A deployment module, which is used to deploy the trained policy network and encoder group on the autonomous driving vehicle to achieve the autonomous driving task.
[0083] In this embodiment, the following may also be included:
[0084] A curriculum learning module is used to control the training process of a reinforcement learning module based on reward distribution. The training process is divided into two stages of training. In the first stage of training, the vehicle visual information is an RGB image with a small field of view not higher than 90 degrees. In the second stage of training, the vehicle visual information is an RGB image with a large field of view higher than 90 degrees. The second stage of training continues the training based on the first stage of training.
[0085] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment, and the implementation methods of the remaining modules will not be repeated here. The system embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.
[0086] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiments can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, the corresponding computer program instructions in the non-volatile memory are read into the memory by the processor of any device with data processing capabilities and run.
[0087] The above-mentioned embodiments only express several implementation modes of the present invention, and the description is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. For those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. An end-to-end autonomous driving decision-making method based on brain-like reinforcement learning, characterized in that: include: Given the vehicle visual information, navigation information and control information, generate its fine-grained information based on the visual information; Fusion of the visual information of the last few time steps and its fine-grained information, navigation information, and control information as the observation of the current time step t , extracting low-dimensional representations of observations using a group of encoders; Reinforcement learning is used to train a policy network based on a spiking neural network. The policy network takes the observed low-dimensional representation as input during the training phase and uses μ, σ RL , σ IL is the output, where μ, σ RL Represents the mean and variance learned by the policy network during reinforcement learning, and generates actions based on the mean and variance; σ IL represents the variance learned by the policy network during imitation learning, where imitation learning refers to imitating the actions of the expert; During the training process, we continuously RL The relationship between the threshold and the selected execution generation action or Expert Action Calculate the reward r based on the action performed t And get the observation o of the next time step t+1 , the quintuple Stored in the data container; when the parameter update time requirements of reinforcement learning are met, the reinforcement learning loss function is constructed based on the return distribution to complete the training of the policy network and encoder group; Deploy the trained policy network and encoder group on an autonomous vehicle to achieve autonomous driving tasks.
2. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1, characterized in that: The training process is divided into two stages of training using a course learning method. In the first stage of training, the vehicle visual information is a small field of view RGB image not higher than 90 degrees. In the second stage of training, the vehicle visual information is a large field of view RGB image higher than 90 degrees. The second stage of training continues the training based on the first stage of training.
3. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1, characterized in that: The control information includes vehicle speed and vehicle steering.
4. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1, characterized in that: The generating of fine-grained information based on visual information includes: Image segmentation technology is used to segment vehicle visual information, namely, RGB image, and merge it with the lane lines in the RGB image detected by lane line detection technology to obtain an image as fine-grained information of visual information.
5. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1 or 4, characterized in that: The encoder group used to extract low-dimensional representations of observations contains four independent encoders corresponding to vehicle visual information and its fine-grained information, navigation information, and control information, respectively.
6. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1, characterized in that: The five-tuple The generation methods include: S1, get the observation at the current time t It incorporates the information from the last two time steps; S2, using the encoder group to extract observation o t The low-dimensional representation h t ; S3, the given low-dimensional representation h t As a policy network π φ The input generates an action It is obtained by sampling the hyperbolic tangent Gaussian function, that is, where f φ (∈,h t ) represents the hyperbolic tangent Gaussian function, ∈~N(0,1) is the standard normal distribution; S4, given expert action and generate actions If σ RL If the value is less than the threshold, the generated action is executed, otherwise the expert action is executed, and the reward r at the current time t is obtained according to the action executed. t and the observation o at time t+1 t+1 ; S5, the above observation o at the current time t t , Generate Action Expert Actions Reward t , observation o at time t+1 t+1 As a quintuple Store in data container middle.
7. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 6, characterized in that: When the parameter update time requirement of reinforcement learning is met, a reinforcement learning loss function is constructed based on reward distribution to complete the training of the policy network and the encoder group, including: S6, from the data container Randomly sample a batch of data from the t , o t+1 Using the encoder group to process h t and h t+1 ; S7, calculate the value network loss based on the reinforcement learning results to update the value network and encoder group: The process of updating the value network parameters is expressed as: in: represents the value network loss under reinforcement learning, Express Find the gradient, θ represents the value network parameter, β θ represents the value network learning rate, represents the value network gradient, Express expectations, represents the output action of the policy network at the current time step t, represents the output action of the policy network at time step t+1, r t represents the reward at the current t-th time step, π φ′ represents the target policy network, h t+1 represents the encoding result of the t+1 time step, Show obedience A random variable with a distribution, Represents a given state-action pair The return distribution obtained by the target value network under the condition, p(.) represents the probability density function, Represents a given target policy network π φ′ right Perform a Bellman update, Represents a given state-action pair The distribution of rewards obtained by the conditional value network; S8, calculate the policy network loss based on the reinforcement learning results to update the policy network: The process of updating the policy network parameter φ is expressed as: Among them: J π (φ) represents the policy network loss under reinforcement learning, Express J π (φ) finds the gradient, φ represents the policy network parameter, β φ represents the policy network learning rate, represents the policy network gradient; Express expectations, represents the output action of the policy network at the current time step t, π φ (·|h t ) indicates that given h t The probability density of the action under the condition of express expectations; express The value of , α represents the temperature parameter, Indicates that at a given h t Action under the conditions probability. S9, calculate the policy network loss based on the imitation learning results to update the policy network and encoder group: The process of updating the value network parameter θ is expressed as: Among them: J IL (φ) represents the policy network loss under imitation learning, Express J IL (φ) finds the gradient, The mean is μ and the standard deviation is σ IL Gaussian distribution of S10, calculate the temperature parameter loss according to the reinforcement learning results to update the temperature parameter: Where: J(α) represents the temperature parameter loss, Indicates that at a given h t Action under the conditions The probability of φ (·|h t ) indicates that given h t The probability density of the action under the condition of represents the expected entropy. S11, soft update process of target strategy network and target value network: After obtaining the above parameters θ and φ, update the target strategy network parameter φ′ and the target value network parameter θ′ according to θ′←τθ+(1-τ)θ′, φ′←τφ+(1-τ)φ′, where τ represents the update rate; S12, repeat S6-S11 until the loss function converges.
8. The end-to-end autonomous driving decision-making method based on brain-like reinforcement learning according to claim 1, characterized in that: The strategy network adopts a pulse neural network, whose output layer is non-pulse neurons and the remaining layers are pulse neurons.
9. An end-to-end autonomous driving decision system based on brain-like reinforcement learning, used to implement the end-to-end autonomous driving decision method according to claim 1, characterized in that: include: A data acquisition module, which is used to acquire given vehicle visual information, navigation information and control information, and generate fine-grained information based on the visual information; The observation-encoding module is used to fuse the visual information of the last few time steps and its fine-grained information, navigation information, and control information as the observation of the current time step. t , extracting low-dimensional representations of observations using a group of encoders; A reinforcement learning module based on reward distribution is used to train a policy network based on a pulse neural network using reinforcement learning. The policy network takes the observed low-dimensional representation as input during the training phase and uses μ, σ RL , σ IL is the output, where μ, σ RL Represents the mean and variance learned by the policy network during reinforcement learning, and generates actions based on the mean and variance; σ IL represents the variance learned by the policy network during imitation learning, where imitation learning refers to imitating the actions of the expert; During the training process, we continuously RL The relationship between the threshold and the selected execution generation action or Expert Action Calculate the reward r based on the action performed t And get the observation o of the next time step t+1 , the quintuple Stored in the data container; when the parameter update time requirements of reinforcement learning are met, the reinforcement learning loss function is constructed based on the return distribution to complete the training of the policy network and encoder group; A deployment module, which is used to deploy the trained policy network and encoder group on the autonomous driving vehicle to achieve the autonomous driving task.
10. The end-to-end autonomous driving decision system based on brain-like reinforcement learning according to claim 9, characterized in that: Also includes: A curriculum learning module is used to control the training process of a reinforcement learning module based on reward distribution. The training process is divided into two stages of training. In the first stage of training, the vehicle visual information is an RGB image with a small field of view not higher than 90 degrees. In the second stage of training, the vehicle visual information is an RGB image with a large field of view higher than 90 degrees. The second stage of training continues the training based on the first stage of training.
Citation Information
Patent Citations
Memristor-based on-chip reinforcement learning pulse GAN model and design method
CN114943329A
End-to-end automatic driving method and system based on reinforcement learning driving world model
CN117218618A
Unmanned aerial vehicle path planning method based on pulse distributed reinforcement learning
CN118913295A
Cited By
Adaptive control method and system based on visual reinforcement learning
CN122156894A