End-to-End Autonomous Driving Method and System Based on a Reinforcement Learning Driving World Model
By adopting a driving world model based on reinforcement learning and a brain-like neural network structure in end-to-end autonomous driving technology, the problem of difficulty in converging and generalization of existing technologies in complex traffic environments is solved, and more efficient autonomous driving control is achieved.
Patent Information
- Application Number
- CN202311186440.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-13
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2043-09-13
AI Technical Summary
The existing end-to-end autonomous driving technology is difficult to converge and generalize in complex traffic environments, resulting in poor control effects.
Using a driving world model based on reinforcement learning, the car's forward-view camera inputs images, extracts environmental dynamics information through a generative world model, and simulates the nematode nervous system to establish a brain-like neural network structure, replacing the traditional perceptron network or convolutional network, and trains reinforcement learning vehicle intelligence through interaction with the environment.
Through continuous interactive training with the environment, reinforcement learning vehicle agents can handle complex traffic environments more effectively and improve the control accuracy and stability of autonomous driving.
Smart Images

Figure CN117218618B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of autonomous driving technology, and particularly to an end-to-end autonomous driving method and system based on a reinforcement learning driving world model. Background Art
[0002] With the rapid development of artificial intelligence, autonomous driving is in a period of rapid transformation from assisted driving to fully autonomous driving. Fully autonomous driving requires that autonomous vehicles can reasonably handle various complex traffic environments and execute driving tasks safely and efficiently. The current mainstream technical solutions for autonomous driving can be divided into two types: modular and end-to-end. The modular solution has strong interpretability, and each module can work efficiently in coordination. Therefore, this method is widely adopted in the industrial community. However, the modular method also has many disadvantages. There may be errors and redundancies in the communication between multiple modules, which ultimately leads to poor control effects limited by the architecture.
[0003] As a typical paradigm of general artificial intelligence, end-to-end autonomous driving technology has received extensive attention and has advantages such as small cumulative error and fast response. However, the end-to-end paradigm mostly uses algorithms such as imitation learning or reinforcement learning for training. The complex neural network architecture and high-dimensional information make it difficult to converge and have limited generalization. Summary of the Invention
[0004] The embodiments of the present application provide an end-to-end autonomous driving method and system based on a reinforcement learning driving world model. This method uses the front view camera of the vehicle as the input, applies a generative driving world model to extract and store environmental dynamics information, and simulates the nematode nervous system to establish a brain-like neural network structure to replace the perceptron network or convolutional network in traditional reinforcement learning. By continuously interacting with the environment, the reinforcement learning vehicle agent is trained to complete the end-to-end autonomous driving task.
[0005] To solve the above technical problems, the embodiments of the present application provide an end-to-end autonomous driving method based on a reinforcement learning driving world model, including the following steps: First, using the front view camera of the vehicle as the input image, based on the generative world model, obtain effective traffic information of the environment; the front view camera of the vehicle is a monocular camera; Next, simulate the nematode neural network to establish a brain-like neural circuit network; Finally, input the effective traffic information of the environment into the brain-like neural circuit network to obtain the control output for autonomous driving.
[0006] In some exemplary embodiments, based on a generative world model, effective traffic information of the environment is obtained, including: performing image encoding on the input image and extracting two-dimensional features from the image features; projecting the two-dimensional features into a three-dimensional space to obtain three-dimensional features, and predicting the depth probability distribution of each three-dimensional feature; based on the depth probability distribution, mapping the three-dimensional features into the bird's-eye view space by means of sum pooling to obtain image features from the bird's-eye view perspective; and obtaining the environmental dynamics information at the current moment based on the image features and the hidden features.
[0007] In some exemplary embodiments, obtaining the environmental dynamics information at the current moment based on the image features and the hidden features includes: fitting the posterior distribution of the environmental dynamics through the image features; fitting the prior distribution of the environmental dynamics through the hidden features; training with the smallest difference between the posterior distribution of the environmental dynamics and the prior distribution of the environmental dynamics, and based on the posterior distribution of the environmental dynamics and the hidden features, generating posterior features and prior features respectively to obtain the environmental dynamics information at the current moment; wherein, the posterior features are sampled and generated by the hidden features containing historical moment information, the action at the previous moment, and the image features; the prior features are sampled and generated by the hidden features containing historical moment information and the action at the previous moment; and the hidden features at the current moment are generated using the environmental dynamics information at the current moment as the hidden features at the next moment.
[0008] In some exemplary embodiments, assuming that both the posterior features and the prior features conform to a normal distribution, the generation processes of the posterior features and the prior features are expressed as:
[0009]
[0010] wherein, x k represents the feature vector; o k represents the input image; s k represents the posterior features; z k represents the prior features; h k represents the hidden features; x k = f e (o k ) represents the process of performing image encoding on the image at time k to obtain image features; q(s k ) ~ N(μ θ (h k , a k-1 , x k ), σ θ (h k , a k-1 , x k )) represents the generation process of the posterior features; represents the generation process of the prior features; a k-1 represents the action at the previous moment; hk+1 = f φ (h k , s k ) represents that the hidden variable at the next moment is encoded by a recurrent neural network.
[0011] In some exemplary embodiments, at future moments, the generation processes of the prior feature and the hidden feature at the next moment are represented as:
[0012]
[0013] Among them, represents the generation process of the prior feature; a k-1 represents the action at the previous moment; h k+T , z k+T respectively represent the hidden feature and the prior feature at future moment k + T; h k+T+1 = f φ (h k+T , z k+T ) represents that at future moment k + T, using the hidden feature h k+T and the prior feature z k+T , the process of generating the hidden feature at the next moment.
[0014] In some exemplary embodiments, the brain-like neural circuit network includes four layers of neurons; among them, the four layers of neurons are respectively: N s sensory neurons, Ni internal neurons, N c instruction neurons, N m motor neurons; between any two consecutive layers, for any source neuron, n so-t synapses are inserted; among them, n so-t satisfies: n so-t ≤ N t , the synaptic polarity satisfies the Bernoulli distribution, where N t represents the number of target neurons, and n so-t target neurons are randomly selected through the binomial distribution; between any two consecutive layers, for any target neuron j without synapses, m so-t synapses are inserted;
[0015] m so-t satisfies: Among them, is the number of synapses to target neuron i, the synaptic polarity satisfies the Bernoulli distribution, and m so-t source neurons are randomly selected through the binomial distribution; the instruction neurons are connected in a cycle, and for any instruction neuron, l so-t synapses are inserted, where l so-t satisfies: l so-t ≤ N c, the synaptic polarity satisfies the Bernoulli distribution, where N c represents the number of command neurons, and l so-t source neurons are randomly selected through the binomial distribution.
[0016] According to the characteristics of the current transmitted between neurons at synapses, each neuron is modeled as:
[0017]
[0018] where x(t) represents the synaptic current of the neuron, I(t) represents the external input of the synapse, A is the deviation matrix, and f I represents the neural network, and τ represents the time constant.
[0019] In some exemplary embodiments, the function g is used to represent the cerebral nerve circuit network, and the cerebral nerve circuit network is adopted to convert the environmental dynamics information into control action information, realizing the conversion process from perception to control. The conversion process is represented by the following formula:
[0020]
[0021] where a k represents the action at the historical moment; h k represents the hidden feature; s k represents the posterior feature; a k+T represents the action at the future moment; h k+T represents the hidden feature at the future moment; z k+T represents the posterior feature at the future moment.
[0022] In some exemplary embodiments, the generative world model is a driving world model after being trained by reinforcement learning; the process of reinforcement learning training includes: adopting a reinforcement learning algorithm introducing entropy regularization to establish the interaction process between the environment and the reinforcement learning agent and the reward function; among them, the value function and the policy are respectively fitted by neural networks controlled by parameters, and the value function and the policy are respectively defined as Q θ (s|a) and π φ (·|s); the Q function receives the input state-action pair (s, a) and outputs a real number as the Q value; the policy π φ (·|s) receives the input state s and outputs a Gaussian distribution with a mean of μ and a standard deviation of σ for actions as a stochastic policy; or performs random sampling of actions, and the sampling result is used as the decision action of the deterministic policy; the Q function and the policy are respectively optimized and solved through the following loss functions:
[0023]
[0024]
[0025] During the interaction process, the environmental information from t k to t k+T is input into the driving world model for training. After encoding the environmental information, the driving world model converts it into hidden variables, and then gives an action instruction through the class brain nerve circuit network. After receiving the action instruction, the environment gives a reward.
[0026] In a second aspect, an embodiment of the present application provides an end-to-end autonomous driving system based on a reinforcement learning driving world model, including an environmental information extraction module, a class brain nerve circuit network construction module, and an output module connected in sequence; the environmental information extraction module is used to use the front-view camera of the vehicle as the input image and obtain effective environmental traffic information based on the generative world model; the front-view camera of the vehicle is a monocular camera; the class brain nerve circuit network construction module is used to simulate the nematode neural network and establish a class brain nerve circuit network; the output module is used to input the effective environmental traffic information into the class brain nerve circuit network to obtain the control output of autonomous driving.
[0027] In some exemplary embodiments, the generative world model includes a connected perception module and the environmental memory module; the perception module is used to use the monocular camera image as the input image, encode the input image, and obtain the image features from the bird's-eye view perspective; the perception module includes a two-dimensional feature encoding unit, a three-dimensional feature encoding unit, and a sum pooling unit connected in sequence; the two-dimensional feature encoding unit is used to extract two-dimensional features from the image features; the three-dimensional feature encoding unit is used to project the two-dimensional features into the three-dimensional space to obtain three-dimensional features and predict the depth probability distribution of each three-dimensional feature; the sum pooling unit is used to map the three-dimensional features into the bird's-eye view space by using the sum pooling method according to the depth probability distribution to obtain the image features from the bird's-eye view perspective; the environmental memory module is used to obtain the environmental dynamics information at the current moment according to the image features and the hidden features.
[0028] The technical solutions provided by the embodiments of the present application have at least the following advantages:
[0029] An embodiment of the present application provides an end-to-end autonomous driving method and system based on a reinforcement learning driving world model. The method includes: First, use the front-view camera of the vehicle as the input image and obtain effective environmental traffic information based on the generative world model; Next, simulate the nematode neural network and establish a class brain nerve circuit network; Finally, input the effective environmental traffic information into the class brain nerve circuit network to obtain the control output of autonomous driving.
[0030] The present application provides an end-to-end autonomous driving method based on a reinforcement learning driving world model. This method uses the front-view camera of the vehicle as input, applies a generative world model to extract and store effective traffic information of the environment, and simulates the nematode nervous system to establish a brain-like neural network structure to replace the perceptron network or convolutional network in traditional reinforcement learning. By continuously interacting with the environment, the reinforcement learning vehicle agent is trained to complete the end-to-end autonomous driving task. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] One or more embodiments are illustrated by way of example in the accompanying drawings, which illustrations do not constitute a limitation of the embodiments, unless otherwise stated, the figures in the drawings do not constitute a scale limitation.
[0032] Figure 1 It is a schematic flow chart of the end-to-end autonomous driving method based on the reinforcement learning driving world model provided by an embodiment of the present application;
[0033] Figure 2 It is a schematic structural diagram of the end-to-end autonomous driving system based on the reinforcement learning driving world model provided by an embodiment of the present application;
[0034] Figure 3 It is a schematic architecture flow chart of the end-to-end autonomous driving method based on the reinforcement learning driving world model provided by an embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of the brain-like neural circuit network architecture provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] As can be seen from the background art, currently, the end-to-end paradigm mostly uses algorithms such as imitation learning or reinforcement learning for training. The complex neural network architecture and high-dimensional information make it difficult to converge and have limited generalization.
[0037] Currently, the development of generative artificial intelligence has not only brought breakthroughs in general artificial intelligence models to the field of natural language processing, but also an opportunity for further development of autonomous driving. The complex parameter structure of generative artificial intelligence can effectively deconstruct the effective traffic information of the environment. The brain-like neural network can simulate biological intelligence, accelerate training and improve the convergence effect. Reinforcement learning interacts with the environment to explore and has the ability of self-evolution, and can continuously explore edge scenarios to solve the long-tail problem. The combination of multiple technologies can better achieve a breakthrough in end-to-end autonomous driving technology.
[0038] To solve the above technical problems, an embodiment of the present application provides an end-to-end autonomous driving method based on a reinforcement learning driving world model, including: First, using the front-view camera of the vehicle as the input image, based on the generative world model, obtain the effective traffic information of the environment; the front-view camera of the vehicle is a monocular camera; Next, simulate the nematode neural network and establish a brain-like neural circuit network; Finally, input the effective traffic information of the environment into the brain-like neural circuit network to obtain the control output of autonomous driving. By providing an end-to-end autonomous driving method based on a reinforcement learning driving world model, the present application uses the front-view camera of the vehicle as the input, applies the generative driving world model to extract and store the environmental dynamics information, and simulates the nematode nervous system to establish a brain-like neural network structure to replace the perceptron network or convolutional network in traditional reinforcement learning. By continuously interacting with the environment, train the reinforcement learning vehicle agent to complete the end-to-end autonomous driving task.
[0039] The following will elaborate on each embodiment of the present application in conjunction with the accompanying drawings. However, those of ordinary skill in the art can understand that in each embodiment of the present application, many technical details are presented to help readers better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can still be implemented.
[0040] See Figure 1 , an embodiment of the present application provides an end-to-end autonomous driving method based on a reinforcement learning driving world model, including the following steps:
[0041] Step S1: Use the front-view camera of the vehicle as the input image, and based on the generative world model, obtain the effective traffic information of the environment; the front-view camera of the vehicle is a monocular camera.
[0042] Step S2: Simulate the nematode neural network and establish a brain-like neural circuit network.
[0043] Step S3: Input the effective traffic information of the environment into the brain-like neural circuit network to obtain the control output of autonomous driving.
[0044] In some embodiments, step S1, based on the generative world model, obtaining the effective traffic information of the environment includes:
[0045] Step S101: Perform image encoding on the input image and extract two-dimensional features from the image features.
[0046] Step S102: Project the two-dimensional features into three-dimensional space to obtain three-dimensional features and predict the depth probability distribution of each three-dimensional feature.
[0047] Step S103: Based on the depth probability distribution, use sum pooling to map the three-dimensional features into the bird's-eye view space to obtain the image features from the bird's-eye view perspective.
[0048] Step S104: Obtain the environmental dynamics information at the current moment based on the image features and the hidden features.
[0049] In some embodiments, in step S104, obtaining the environmental dynamics information at the current moment based on the image features and the hidden features includes:
[0050] Step S1041: Fit the posterior distribution of the environmental dynamics through the image features.
[0051] Step S1042: Fit the prior distribution of the environmental dynamics through the hidden features.
[0052] Step S1043: Train with the smallest difference between the posterior distribution of the environmental dynamics and the prior distribution of the environmental dynamics, and based on the posterior distribution of the environmental dynamics and the hidden features, generate posterior features and prior features respectively to obtain the environmental dynamics information at the current moment. Among them, the posterior features are sampled and generated by the hidden features containing historical moment information, the action at the previous moment, and the image features; the prior features are sampled and generated by the hidden features containing historical moment information and the action at the previous moment; use the environmental dynamics information at the current moment to generate the hidden features at the current moment as the hidden features at the next moment.
[0053] See Figure 2 , the embodiment of the present application further provides an end-to-end autonomous driving system based on a reinforcement learning driving world model, including an environmental information extraction module 101, a neural circuit network construction module 102, and an output module 103 connected in sequence; the environmental information extraction module 101 is used to use the front-view camera of the vehicle as the input image and obtain the effective traffic information of the environment based on the generative world model; the front-view camera of the vehicle is a monocular camera; the neural circuit network construction module 102 is used to simulate the nematode neural network and establish a neural circuit network; the output module 103 is used to input the effective traffic information of the environment into the neural circuit network to obtain the control output of autonomous driving.
[0054] In some embodiments, the generative world model includes a connected perception module and the environmental memory module; the perception module is configured to use a monocular camera image as an input image, perform image encoding on the input image, and obtain image features from a bird's-eye view perspective; the perception module includes a two-dimensional feature encoding unit, a three-dimensional feature encoding unit, and a sum pooling unit connected in sequence; the two-dimensional feature encoding unit is configured to extract two-dimensional features from the image features; the three-dimensional feature encoding unit is configured to project the two-dimensional features into a three-dimensional space to obtain three-dimensional features and predict the depth probability distribution of each three-dimensional feature; the sum pooling unit is configured to map the three-dimensional features into the bird's-eye view space in a sum pooling manner according to the depth probability distribution to obtain image features from a bird's-eye view perspective; the environmental memory module is configured to obtain the environmental dynamics information at the current moment according to the image features and the hidden features.
[0055] Specifically, the model architecture of the generative world model is as Figure 3 shown.
[0056] This application uses a generative world model to process the perception of the environment by an autonomous driving system. As Figure 3 shown, the actual input of the world model at time k is the image input o k , and this input will be encoded into a feature vector x k in the model. This process is mathematically defined as: x k = f e (o k ). The memory of the model for this vector can be divided into two parts, namely the posterior feature s k and the prior feature z k , both of which follow a Gaussian distribution. The posterior feature s k is sampled and generated by the hidden feature h k containing historical moment information, the action a k-1 (longitudinal and lateral accelerations) at the previous moment, and the image feature x k ; the prior feature is sampled and generated by the hidden feature h k and the action a k-1 at the previous moment. The hidden variable at the next moment is encoded through a recurrent neural network and can be expressed as h k+1 = f φ (h k , s k ).
[0057] In some embodiments, the posterior feature s k is sampled and generated by the hidden feature containing historical moment information, the action at the previous moment, and the feature vector; the prior feature z k is sampled and generated by the hidden feature containing historical moment information and the action at the previous moment.
[0058] In some embodiments, assuming that both the posterior feature and the prior feature conform to a normal distribution, the generation processes of the posterior feature and the prior feature are expressed as:
[0059]
[0060] where x k represents the feature vector; o k represents the input image; s k represents the posterior feature; z k represents the prior feature; h k represents the hidden feature; x k = f e (o k ) represents the process of image encoding the image at time k to obtain the image feature; q(s k ) ~ N(μ θ (h k , a k-1 , x k ), σ θ (h k , a k-1 , x k )) represents the generation process of the posterior feature; represents the generation process of the prior feature; a k-1 represents the action at the previous time; h k+1 = f φ (h k , s k ) indicates that the hidden variable at the next time is encoded through a recurrent neural network.
[0061] Since the generative world model cannot obtain image input at the future time k+T, the generative world model obtains future actions through prediction. Specifically, the generative world model does not generate posterior features at the future time k+T, but directly uses the hidden feature h k+T and the prior feature z k+T to generate the hidden feature h k+T+1 at the next time.
[0062] In some embodiments, at the future time k+T, the generation processes of the prior feature z k+T and the hidden feature at the next time are expressed as:
[0063]
[0064] where represents the generation process of the prior feature; a k-1 represents the action at the previous time; h k+T , z k+T represent the hidden feature and the prior feature at the future time k+T, respectively.
[0065] h k+T+1 = f φ (h k+T , z k+T ) represents the process of generating the hidden feature at the next moment using the hidden feature h k+T and the prior feature z k+T at the future moment k + T.
[0066] This application mimics the nematode nervous system and establishes a brain-like neural circuit network to replace the traditional neural network to enhance the learning ability of the model. This application mimics the activation mode of Caenorhabditis elegans neurons and the way they communicate with each other through electrical impulses to establish a brain-like neural circuit architecture, as Figure 4 shown.
[0067] See Figure 4 , in some embodiments, the brain-like neural circuit network includes four layers of neurons; among them, the four layers of neurons are respectively: N s sensory neurons, N i interneurons, N c command neurons, N m motor neurons; between any two consecutive layers, for any source neuron, n so-t synapses are inserted; among them, n so-t satisfies: n so-t ≤ N t , the synaptic polarity satisfies the Bernoulli distribution, where N t represents the number of target neurons, and n so-t target neurons are randomly selected through the binomial distribution; between any two consecutive layers, for any target neuron j without a synapse, m so-t synapses are inserted.
[0068] m so-t satisfies: where is the number of synapses to target neuron i, the synaptic polarity satisfies the Bernoulli distribution, and m so-t source neurons are randomly selected through the binomial distribution; the command neurons are connected in a cycle, and for any command neuron, l so-t synapses are inserted, where l so-t satisfies: l so-t ≤ N c , the synaptic polarity satisfies the Bernoulli distribution, where N c represents the number of command neurons, and l so-t source neurons are randomly selected through the binomial distribution. According to the characteristics of the current transmission between neuron synapses, each neuron is modeled as:
[0069]
[0070] Among them, x(t) represents the neuron synaptic current, I(t) represents the synaptic external input, A is the deviation matrix, and f I represents the neural network, and τ represents the time constant.
[0071] As described above, imitate the activation mode of Caenorhabditis elegans neurons and the way of communicating with each other through electrical pulses to establish a brain-like neural circuit network. Use the function g to represent the brain-like neural circuit network. Adopt the brain-like neural circuit network to convert environmental dynamics information into control action information, and realize the conversion process from perception to control. The conversion process is represented by the following formula:
[0072]
[0073] Among them, a k represents the action at the historical moment; h k represents the hidden feature; s k represents the posterior feature; a k+T represents the action at the future moment; h k+T represents the hidden feature at the future moment; z k+T represents the posterior feature at the future moment.
[0074] In some embodiments, the generative world model is a driving world model after being trained by reinforcement learning; the process of reinforcement learning training includes: adopting the SAC (Soft Actor Critic) reinforcement learning algorithm introducing entropy regularization to establish the interaction process between the environment and the reinforcement learning agent and the reward function; among them, the value function and the policy are respectively fitted by neural networks controlled by parameters, and the value function and the policy are respectively defined as Q θ (s|a) and π φ (·|s); the Q function receives the input state-action pair (s, a) and outputs a real number as the Q value; the policy π φ (·|s) receives the input state s and outputs a Gaussian distribution with a mean of μ and a standard deviation of σ for actions as the stochastic policy; or performs action random sampling, and the sampling result is used as the decision action of the deterministic policy; the Q function and the policy are respectively optimized and solved through the following loss functions:
[0075]
[0076]
[0077] During the interaction process, from t k to t k+TThe environmental information is input into the driving world model for training. After encoding the environmental information, the driving world model converts it into hidden variables, and then gives action instructions through the neural circuit network of the brain-like. After receiving the action instructions, the environment gives rewards. The goal of reinforcement learning training is to make the rewards converge to a relatively high value with a small amount of fluctuation allowed, so as to achieve the overall control of vehicle autonomous driving.
[0078] With the above technical solutions, the embodiments of the present application provide an end-to-end autonomous driving method and system based on a reinforcement learning driving world model. The method includes: First, using the front view camera of the vehicle as the input image, based on the generative world model, obtain the effective traffic information of the environment; the front view camera of the vehicle is a monocular camera; Next, simulate the neural network of nematodes to establish a neural circuit network of the brain-like; Finally, input the effective traffic information of the environment into the neural circuit network of the brain-like to obtain the control output of autonomous driving.
[0079] The present application provides an end-to-end autonomous driving method based on a reinforcement learning driving world model. The method uses the front view camera of the vehicle as the input, applies the generative world model to extract and store the effective traffic information of the environment, and simulates the nematode nervous system to establish a brain-like neural network structure to replace the perceptron network or convolutional network in traditional reinforcement learning. By continuously interacting with the environment, train the reinforcement learning vehicle agent to complete the end-to-end autonomous driving task.
[0080] Those of ordinary skill in the art can understand that the above embodiments are specific embodiments for implementing the present application. In actual applications, various changes can be made in form and details without departing from the spirit and scope of the present application. Any person skilled in the art can make their own changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be subject to the scope defined by the claims.
Claims
1. An end-to-end autonomous driving method based on a reinforcement learning driving world model, characterized in that, it includes the following steps: Using the front-view camera of the vehicle as the input image, based on the generative world model, obtain the effective traffic information of the environment; the front-view camera of the vehicle is a monocular camera; Simulate the nematode neural network and establish a neural circuit network similar to the brain; Input the effective traffic information of the environment into the neural circuit network similar to the brain to obtain the control output of autonomous driving; Based on the generative world model, obtaining the effective traffic information of the environment includes: Perform image encoding on the input image and extract two-dimensional features from the image features; Project the two-dimensional features into three-dimensional space to obtain three-dimensional features and predict the depth probability distribution of each three-dimensional feature; Based on the depth probability distribution, use sum pooling to map the three-dimensional features into the bird's-eye view space to obtain the image features in the bird's-eye view; Based on the image features and hidden features, obtain the environmental dynamics information at the current moment; Based on the image features and hidden features, obtaining the environmental dynamics information at the current moment includes: Fit the posterior distribution of environmental dynamics through the image features; Fit the prior distribution of environmental dynamics through the hidden features; Train with the smallest difference between the posterior distribution of environmental dynamics and the prior distribution of environmental dynamics, and based on the posterior distribution of environmental dynamics and the hidden features, generate posterior features and prior features respectively to obtain the environmental dynamics information at the current moment; where The posterior features are sampled and generated by the hidden features containing historical moment information, the action at the previous moment, and the image features; The prior features are sampled and generated by the hidden features containing historical moment information and the action at the previous moment; Use the environmental dynamics information at the current moment to generate the hidden features at the current moment as the hidden features at the next moment; Assume that both the posterior features and the prior features conform to the normal distribution, and the generation processes of the posterior features and the prior features are expressed as: (1) Among them, represents the feature vector; represents the input image; represents the posterior feature; represents the prior feature; represents the hidden feature; represents the process of k encoding the image at time to obtain the image feature; represents the generation process of the posterior feature; represents the action at the previous time; represents that the hidden variable at the next time is encoded by a recurrent neural network.
2. The end-to-end autonomous driving method based on the reinforcement learning driving world model according to claim 1, characterized in that, At a future moment, the generation processes of the prior features and the hidden features at the next moment are expressed as: (2) Among them, represents the generation process of prior features; represents the action at the previous moment; , respectively represent the hidden feature and prior feature at the future moment ; represents the process of generating the hidden feature at the next moment using the hidden feature and prior feature at the future moment .
3. The end-to-end autonomous driving method based on the reinforcement learning driving world model according to claim 1, characterized in that, The neural circuit network similar to the brain includes four layers of neurons; where The four layers of neurons are respectively: sensory neurons, internal neurons, command neurons, motor neurons; Between any two consecutive layers, for any source neuron, insert synapses; among them, satisfies: , the synaptic polarity satisfies the Bernoulli distribution, where target neurons are randomly selected through the binomial distribution; Between any two consecutive layers, any target neuron without synapses Insert synapses, Satisfy: ; Among them, is the number of synapses to the target neuron , the synaptic polarity satisfies the Bernoulli distribution, source neurons are randomly selected through the binomial distribution; Recurrent connections between command neurons. For any command neuron, insert synapses, where satisfies: , the synaptic polarity satisfies the Bernoulli distribution, source neurons are randomly selected through the binomial distribution; According to the characteristics of the current transmitted between neuron synapses, each neuron is modeled as: (3) Among them, represents the neuron synaptic current, represents the synaptic external input, A is the deviation matrix, represents the neural network, represents the time constant.
4. The end-to-end autonomous driving method based on the reinforcement learning driving world model according to claim 1, characterized in that, Using a function represents a neural circuit network. The neural circuit network is used to convert environmental dynamics information into control action information, realizing the conversion process from perception to control. The conversion process is represented by the following formula: (4) Among them, represents the action at the historical moment; represents the hidden feature; represents the posterior feature; represents the action at the future moment; represents the hidden feature at the future moment; represents the posterior feature at the future moment.
5. The end-to-end autonomous driving method based on the reinforcement learning driving world model according to claim 1, characterized in that, The generative world model is a driving world model trained through reinforcement learning; The reinforcement learning training includes: An enhanced learning algorithm with entropy regularization is adopted to establish the interaction process between the environment and the enhanced learning agent, as well as the reward function. Among them, the value function and the policy are respectively fitted by neural networks controlled by parameters, and the value function is represented by indicating that the function and the policy are respectively defined as and ; The function receives an input state-action pair and outputs a real value as the value ; Policy Receive the input state , and output a Gaussian distribution with a mean of and a standard deviation of as a stochastic policy; or perform random sampling of actions, and the sampling result is used as the decision action of the deterministic policy; The function and the strategy are optimized and solved through the following loss functions respectively: (5) (6) During the interaction process, to the environmental information is input into the driving world model for training. After encoding the environmental information, the driving world model converts it into hidden variables, and then gives action instructions through the neural circuit network. After receiving the action instructions, the environment gives rewards.
6. An end-to-end autonomous driving system based on a reinforcement learning driving world model, characterized in that, It includes an environmental information extraction module, a neural circuit network construction module similar to the brain, and an output module connected in sequence; The environmental information extraction module is used to take the front view camera of the vehicle as the input image and obtain effective environmental traffic information based on the generative world model; The artificial neural circuit network construction module is used to simulate the nematode neural network and establish an artificial neural circuit network; The output module is used to input the effective environmental traffic information into the artificial neural circuit network to obtain the control output for autonomous driving; The generative world model includes a connected perception module and the environmental memory module; The perception module is used to take a monocular camera image as the input image, perform image encoding on the input image, and obtain image features from the perspective of a bird's-eye view; the perception module includes a two-dimensional feature encoding unit, a three-dimensional feature encoding unit, and a sum pooling unit connected in sequence; the two-dimensional feature encoding unit is used to extract two-dimensional features from the image features; The three-dimensional feature encoding unit is used to project the two-dimensional features into three-dimensional space to obtain three-dimensional features and predict the depth probability distribution of each three-dimensional feature; the sum pooling unit is used to map the three-dimensional features into the bird's-eye view space in a sum pooling manner according to the depth probability distribution to obtain image features from the perspective of a bird's-eye view; the environmental memory module is used to obtain the environmental dynamics information at the current moment according to the image features and hidden features; Based on the image features and hidden features, obtain the environmental dynamics information at the current moment; Based on the image features and hidden features, obtaining the environmental dynamics information at the current moment includes: Fitting the environmental dynamics posterior distribution through the image features; Fitting the environmental dynamics prior distribution through the hidden features; Training with the difference between the environmental dynamics posterior distribution and the environmental dynamics prior distribution being minimized, and based on the environmental dynamics posterior distribution and the hidden features, generating posterior features and prior features respectively to obtain the environmental dynamics information at the current moment; where The posterior features are sampled and generated by the hidden features containing historical moment information, the action at the previous moment, and the image features; The prior features are sampled and generated by the hidden features containing historical moment information and the action at the previous moment; Using the environmental dynamics information at the current moment to generate the hidden features at the current moment as the hidden features for the next moment; Assuming that both the posterior features and the prior features conform to the normal distribution, the generation processes of the posterior features and the prior features are expressed as: (1) Among them, represents the feature vector; represents the input image; represents the posterior feature; represents the prior feature; represents the hidden feature; represents the process of image encoding the image at time k to obtain image features; represents the generation process of the posterior feature; represents the generation process of the prior feature; represents the action at the previous time; represents that the hidden variable at the next time is encoded by a recurrent neural network.
Citation Information
Patent Citations
Performance testing of robotic systems
CN114270369A
Bionic motion control method based on caenorhabditis elegans neural network
CN114897125A