End-to-end automatic driving reinforcement learning method and device

Through Transformer codec of multi-view images and point cloud data and world model updates, road topology and environment state are generated, which solves the inadequate reinforcement learning of existing end-to-end autonomous driving algorithms and realizes optimal strategy learning for the agent in complex scenarios.

CN120428567APending Publication Date: 2025-08-05CHERY INTELLIGENT VEHICLE TECH (HEFEI) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510568797.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing end-to-end autonomous driving algorithm lacks an effective reinforcement learning framework, making it difficult to implement optimal strategies in complex scenarios, and the simulation environment is difficult to cover all autonomous driving scenarios.

Method used

By obtaining multi-view image data, point cloud data and historical frame video data, the Transformer architecture is used to code and generate multi-information fusion tensors, generate road topology and filtering for driving paths, and update the environment status using the world model to achieve offline reinforcement learning.

Benefits of technology

Agencies that transcend human experience are trained, solve the problem of interaction between the agent and the simulation environment, and realize offline reinforcement learning of end-to-end autonomous driving algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428567A_ABST
    Figure CN120428567A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic driving, in particular to an end-to-end automatic driving reinforcement learning method and device, and the method comprises the steps: obtaining the multi-view image data, point cloud data and historical frame video data of a target vehicle, inputting the multi-view image data and the point cloud data into a Transform architecture for coding and decoding, and obtaining the multi-view image data, the point cloud data and the historical frame video data of the target vehicle; a multi-information fusion tensor is obtained; generating a road topology by using the multi-information fusion tensor, and screening the drivable paths of the vehicle intelligent agent by using the road topology to obtain a final planned path; and inputting the final planned path and the historical frame video data into the world model to generate future video data, and updating the environment state of the next time step of the vehicle agent by using the future video data. Therefore, the problems that training of an existing end-to-end automatic driving algorithm lacks an effective support of a reinforcement learning framework and a strategy is difficult to optimize in practice are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving technology, and in particular to an end-to-end autonomous driving reinforcement learning method and device. Background Art

[0002] In the evolution of autonomous driving technology, breakthroughs in end-to-end algorithms have consistently faced a core conflict between data dependency and strategic limitations. The current mainstream imitation learning paradigm attempts to replicate human decision-making behind the wheel by training models with massive amounts of human driving data. However, this approach is inherently constrained by the limitations of human experience: human drivers may misjudge when responding to unexpected situations, rely on empirical intuition rather than optimal solutions in complex scenarios, and even make suboptimal decisions due to fatigue or distraction. When a model learns solely by imitating human behavior, its performance is capped at the average level of human drivers. It is unable to break through the cognitive boundaries of humans in extreme scenarios, nor does it achieve a qualitative leap in rules and strategies.

[0003] In existing technologies, reinforcement learning is often used to abstract autonomous vehicles into intelligent agents. Through trial and error interactions with virtual or real environments, the optimal strategy is achieved through a reward-and-penalty mechanism. However, the training of existing end-to-end autonomous driving algorithms lacks an effective reinforcement learning framework, making it difficult to achieve optimal strategies in practice. For example, an end-to-end autonomous driving control method and device for urban scenarios based on an attention mechanism and graphical reinforcement learning adds rule-based vehicles as other traffic participants in the Carla simulation environment and performs online reinforcement learning training. However, the rules for these added traffic participants in this method must be set before reinforcement learning, which is not only labor-intensive but also difficult to cover all autonomous driving scenarios. End-to-end autonomous driving decision-making methods based on monocular RGB-D features and reinforcement learning require pre-training an intelligent agent network during the reinforcement learning process to obtain a Q-value training policy model. However, the interaction between the agent and the environment is also based on the Carla simulation environment, which means that the simulation environment still cannot cover all autonomous driving scenarios. Summary of the Invention

[0004] The present invention provides an end-to-end autonomous driving reinforcement learning method and device to solve the problems that the training of existing end-to-end autonomous driving algorithms lacks the support of an effective reinforcement learning framework and is difficult to optimize the strategy in practice.

[0005] The first aspect of the present invention provides an end-to-end autonomous driving reinforcement learning method, comprising the following steps: obtaining multi-perspective image data, point cloud data, and historical frame video data of a target vehicle, and inputting the multi-perspective image data and the point cloud data into a Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor; using the multi-information fusion tensor to generate a road topology, and using the road topology to filter the drivable paths of a vehicle intelligent body to obtain a final planned path; inputting the final planned path and the historical frame video data into a world model to generate future video data, and using the future video data to update the environmental state of the vehicle intelligent body in the next time step.

[0006] Optionally, inputting the multi-view image data and the point cloud data into a Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor includes:

[0007] extracting a plurality of bird's-eye view features from the multi-view image data and the point cloud data, and fusing the plurality of bird's-eye view features to obtain a global spatial semantic representation;

[0008] The global spatial semantic representation is input into the Transformer architecture, and the global spatial semantic representation is encoded and decoded through the self-attention mechanism and the cross-attention mechanism to obtain the multi-information fusion tensor.

[0009] Optionally, the generating of a road topology using the multi-information fusion tensor, and screening a drivable path of the vehicle agent using the road topology to obtain a final planned path, includes:

[0010] Performing mapping according to the multi-information fusion tensor to generate the road topology;

[0011] The drivable path of the vehicle agent is encoded as a query tensor, and the query tensor is filtered using the road topology to obtain a final planned path of the vehicle agent.

[0012] Optionally, inputting the final planned path and the historical frame video data into a world model to generate future video data, and using the future video data to update the environmental state of the vehicle agent at the next time step includes:

[0013] Performing discrete action marking on the final planned path to obtain a plurality of action prompt words;

[0014] The multiple action prompt words and the historical frame video data are input into the world model to generate the future video data, and the future video data is used to update the environmental state of the vehicle agent in the next time step.

[0015] The second aspect of the present invention provides an end-to-end autonomous driving reinforcement learning device, including: a coding and decoding module for acquiring multi-perspective image data, point cloud data and historical frame video data of a target vehicle, and inputting the multi-perspective image data and the point cloud data into a Transformer architecture for coding and decoding to obtain a multi-information fusion tensor; a screening module for generating a road topology using the multi-information fusion tensor, and using the road topology to screen the drivable paths of the vehicle intelligent body to obtain a final planned path; a generation module for inputting the final planned path and the historical frame video data into a world model to generate future video data, and using the future video data to update the environmental state of the vehicle intelligent body in the next time step.

[0016] Optionally, the encoding and decoding module includes:

[0017] An extraction and fusion unit, configured to extract a plurality of bird's-eye view features from the multi-view image data and the point cloud data, and fuse the plurality of bird's-eye view features to obtain a global spatial semantic representation;

[0018] The encoding and decoding unit is used to input the global spatial semantic representation into the Transformer architecture, encode and decode the global spatial semantic representation through the self-attention mechanism and the cross-attention mechanism to obtain the multi-information fusion tensor.

[0019] Optionally, the screening module includes:

[0020] A generating unit, performing mapping according to the multi-information fusion tensor to generate the road topology;

[0021] A screening unit is used to encode the drivable path of the vehicle agent into a query tensor and filter the query tensor using the road topology to obtain a final planned path of the vehicle agent.

[0022] Optionally, the generating module includes:

[0023] a discrete marking unit, configured to perform discrete marking of actions on the final planned path to obtain a plurality of action prompt words;

[0024] A generation and update unit is used to input the multiple action prompt words and the historical frame video data into the world model to generate the future video data, and use the future video data to update the environmental state of the vehicle intelligent body in the next time step.

[0025] The third aspect of the present invention provides a vehicle, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the end-to-end autonomous driving reinforcement learning method as described in the above embodiment.

[0026] The fourth aspect of the present invention provides a computer-readable storage medium, which stores a computer program. When the program is executed by a processor, it implements the above end-to-end autonomous driving reinforcement learning method.

[0027] The end-to-end autonomous driving reinforcement learning device and method proposed in the embodiments of the present invention models the environment of the intelligent agent through a world model, learns good strategies through trial and error of the intelligent agent in the world model, and trains an intelligent agent that surpasses human experience, thereby realizing offline reinforcement learning of the end-to-end autonomous driving algorithm; in particular, the intelligent agent's actions (i.e., the planned path) are converted into tokens and input into the world model as prompts, so that the world model can generate video data for a period of time in the future based on historical video data and actions taken by the intelligent agent, and update the environment of the intelligent agent based on the generated video data, thereby solving the problem of difficult interaction between the intelligent agent and the simulation environment.

[0028] Additional aspects and advantages of the present invention will be given in part in the following description and in part will become obvious from the following description, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0030] Figure 1 A flowchart of an end-to-end autonomous driving reinforcement learning method provided by an embodiment of the present invention;

[0031] Figure 2 This is a specific execution flow chart of an end-to-end autonomous driving reinforcement learning method provided by an embodiment of the present invention;

[0032] Figure 3 A block diagram of an end-to-end autonomous driving reinforcement learning device provided by an embodiment of the present invention;

[0033] Figure 4 The present invention is a block diagram of a vehicle provided in accordance with an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention and are not to be construed as limiting the present invention.

[0035] The following describes an end-to-end autonomous driving reinforcement learning method and apparatus according to an embodiment of the present invention with reference to the accompanying drawings.

[0036] Figure 1 A flowchart of an end-to-end autonomous driving reinforcement learning method provided by an embodiment of the present invention.

[0037] like Figure 1 As shown, the end-to-end autonomous driving reinforcement learning method includes the following steps:

[0038] In step S101, multi-view image data, point cloud data and historical frame video data of the target vehicle are obtained, and the multi-view image data and point cloud data are input into the Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor.

[0039] In some embodiments, multi-view image data and point cloud data are input into a Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor, including:

[0040] Extract multiple bird's-eye view features from multi-view image data and point cloud data, and fuse multiple bird's-eye view features to obtain a global spatial semantic representation;

[0041] The global spatial semantic representation is input into the Transformer architecture, and the global spatial semantic representation is encoded and decoded through the self-attention mechanism and the cross-attention mechanism to obtain a multi-information fusion tensor.

[0042] It should be noted that in the process of reinforcement learning, the environment state of the agent is S t , when the agent takes action a, the environment changes to S t+1 , define the reward function r, the agent will receive rewards due to changes in the environment, in order to make the agent consider long-term reward benefits when choosing actions, define the Q value function as Represents the weighted sum of the rewards for the next n time steps, where k i represents the reward weight at the i-th time step, r i Represents the reward value at the i-th time step. The traditional deep reinforcement learning algorithm based on deep deterministic policy gradient (DDPG) trains a total of two neural networks, one is the Actor-network, and the input is the environment state S t, the output is the optimal strategy action a that the agent should take, and the other neural network is Critic-network, the input is S t The output is the estimated Q-value function. Through the agent's continuous exploration of the environment, the two neural networks are trained. The trained actor-network can be used to guide the agent to take the best action in various environments.

[0043] In the actual implementation process, Figure 2 As shown, multi-view image data, point cloud data, and historical frame video data of the target vehicle are obtained, and multiple bird's-eye view features in the multi-view image data and point cloud data are extracted respectively, and the multiple bird's-eye view features are fused to obtain a global spatial semantic representation.

[0044] Furthermore, the global spatial semantic representation is input into the Transformer architecture, and the global spatial semantic representation is effectively encoded and decoded through the self-attention mechanism and cross-attention mechanism to obtain a multi-information fusion tensor.

[0045] In step S102, a road topology is generated using a multi-information fusion tensor, so as to filter the drivable paths of the vehicle intelligent body using the road topology to obtain a final planned path.

[0046] In some embodiments, a multi-information fusion tensor is used to generate a road topology, and the road topology is used to screen the drivable paths of the vehicle agent to obtain a final planned path, including:

[0047] Mapping is performed based on multi-information fusion tensors to generate road topology;

[0048] The drivable path of the vehicle agent is encoded as a query tensor, and the query tensor is filtered using the road topology to obtain the final planned path of the vehicle agent.

[0049] In the actual implementation process, Figure 2 As shown, mapping is performed based on the multi-information fusion tensor to generate road topology. After the road topology is generated, the model parameters of the mapping process will be frozen and will no longer be updated with the subsequent reinforcement learning process.

[0050] Furthermore, all drivable paths of the vehicle agent are used as action spaces, and the agent's goal is to t Therefore, the embodiment of the present invention encodes the path into a query tensor Query (in the reinforcement learning process, the model parameters of this part are also frozen), and the result output by the Transformer module is defined as the key K and the value V. According to the attention calculation formula where d kThe dimension of K is used to filter the action space of the agent according to the road topology, and some planned paths that violate the rules are removed to obtain the final planned path of the vehicle agent.

[0051] In step S103, the final planned path and the historical frame video data are input into the world model to generate future video data, and the future video data is used to update the environmental state of the vehicle agent in the next time step.

[0052] In some embodiments, the final planned path and historical frame video data are input into the world model to generate future video data, and the future video data is used to update the environmental state of the vehicle agent at the next time step, including:

[0053] The final planned path is discretized into action labels to obtain multiple action prompt words;

[0054] Multiple action prompts and historical frame video data are input into the world model to generate future video data, and the future video data is used to update the environmental state of the vehicle agent in the next time step.

[0055] In the actual implementation process, Figure 2 As shown in the figure, the final planned path is discretely marked with actions through the tokenizer to obtain multiple token action words, and the multiple token action words and historical frame video data are input into the world model to generate future video data. The world model is a generative large model established based on the diffusion model technology. The historical frame video data is used to generate future video data in the world model. At the same time, the world model uses the action word token of the agent as the Prompt prompt word, which has reflected the changes in the environment under the action of the agent. After the multiple action prompt words and the historical frame video data are input into the world model together (the parameters of the world model are also frozen during the reinforcement learning process), the world model can generate video data for a period of time in the future based on the historical video data and the actions taken by the agent, and use the future video data to update the environmental state S of the vehicle agent in the next time step. t+1 .

[0056] It should be noted that the reward function of the vehicle agent can be defined as r = ω1Δv + ω2Δδ + ω3Δa x +ω4Δa y +ω5C, where ω1Δv represents the speed variance during the planned path selected by the vehicle agent, ω2Δδ represents the front wheel angle variance during the planned path selected by the vehicle agent, and ω3Δa x ω4Δa ywhere ω5C represents the longitudinal and lateral acceleration variances of the vehicle agent's chosen path, respectively. The collision penalty is C = 1 if a collision occurs and C = 0 otherwise. Note that the specific form of the reward function is not limited here; those skilled in the art may adjust the reward function based on their specific circumstances.

[0057] In summary, according to the end-to-end autonomous driving reinforcement learning method proposed in the embodiment of the present invention, the environment of the intelligent agent is modeled through a world model, and good strategies are learned through trial and error of the intelligent agent in the world model, and an intelligent agent that surpasses human experience is trained, thereby realizing offline reinforcement learning of the end-to-end algorithm for autonomous driving; in particular, the action of the intelligent agent (i.e., the planned path) is converted into a token, which is input into the world model as a prompt word, so that the world model can generate video data for a period of time in the future based on historical video data and the actions taken by the intelligent agent, and update the environment of the intelligent agent based on the generated video data, thereby solving the problem of difficult interaction between the intelligent agent and the simulation environment.

[0058] Next, the end-to-end autonomous driving reinforcement learning device proposed in accordance with an embodiment of the present invention will be described with reference to the accompanying drawings.

[0059] Figure 3 1 is a block diagram of an end-to-end autonomous driving reinforcement learning device according to an embodiment of the present invention.

[0060] like Figure 3 As shown, the end-to-end autonomous driving reinforcement learning device 30 includes: a coding and decoding module 301, a screening module 302 and a generation module 303.

[0061] The encoding and decoding module 301 is used to obtain multi-view image data, point cloud data, and historical frame video data of the target vehicle, and input the multi-view image data and point cloud data into the Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor. The screening module 302 is used to generate a road topology using the multi-information fusion tensor, and then use the road topology to filter the drivable paths of the vehicle agent to obtain the final planned path. The generation module 303 is used to input the final planned path and historical frame video data into the world model to generate future video data, and use the future video data to update the environmental state of the vehicle agent at the next time step.

[0062] In some embodiments, the codec module 301 includes:

[0063] An extraction and fusion unit is used to extract multiple bird's-eye view features from multi-view image data and point cloud data, and fuse the multiple bird's-eye view features to obtain a global spatial semantic representation;

[0064] The encoder-decoder unit is used to input the global spatial semantic representation into the Transformer architecture, encode and decode the global spatial semantic representation through the self-attention mechanism and the cross-attention mechanism to obtain a multi-information fusion tensor.

[0065] In some embodiments, the screening module 302 includes:

[0066] The generation unit builds a map based on the multi-information fusion tensor to generate the road topology;

[0067] The screening unit is used to encode the drivable path of the vehicle agent into a query tensor and filter the query tensor using the road topology to obtain the final planned path of the vehicle agent.

[0068] In some embodiments, the generation module 303 includes:

[0069] A discrete marking unit is used to discretely mark the actions of the final planned path to obtain multiple action prompt words;

[0070] The generation and update unit is used to input multiple action prompt words and historical frame video data into the world model to generate future video data, and use the future video data to update the environmental state of the vehicle agent in the next time step.

[0071] It should be noted that the above explanation of the embodiment of the end-to-end autonomous driving reinforcement learning method is also applicable to the end-to-end autonomous driving reinforcement learning device of this embodiment and will not be repeated here.

[0072] According to the end-to-end autonomous driving reinforcement learning device proposed in the embodiment of the present invention, the environment of the intelligent agent is modeled through a world model, and good strategies are learned through trial and error of the intelligent agent in the world model, and an intelligent agent that surpasses human experience is trained, thereby realizing offline reinforcement learning of the end-to-end algorithm for autonomous driving; in particular, the action of the intelligent agent (i.e., the planned path) is converted into a token, which is input into the world model as a prompt word, so that the world model can generate video data for a period of time in the future based on historical video data and actions taken by the intelligent agent, and update the environment of the intelligent agent based on the generated video data, thereby solving the problem of difficult interaction between the intelligent agent and the simulation environment.

[0073] Figure 4 A schematic diagram of the structure of a vehicle provided by an embodiment of the present invention. The vehicle may include:

[0074] Memory 401 , processor 402 , and computer programs stored in the memory 401 and executable on the processor 402 .

[0075] When the processor 402 executes the program, the end-to-end autonomous driving reinforcement learning method provided in the above embodiment is implemented.

[0076] Furthermore, the vehicle further comprises:

[0077] The communication interface 403 is used for communication between the memory 401 and the processor 402 .

[0078] The memory 401 is used to store computer programs that can be run on the processor 402 .

[0079] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0080] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0081] Optionally, in a specific implementation, if the memory 401 , the processor 402 and the communication interface 403 are integrated on a chip, the memory 401 , the processor 402 and the communication interface 403 can communicate with each other through an internal interface.

[0082] The processor 402 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0083] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned end-to-end autonomous driving reinforcement learning method.

[0084] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.

[0085] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0086] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0087] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or N wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0088] It should be understood that the various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0089] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0090] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0091] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. A person skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. An end-to-end autonomous driving reinforcement learning method, characterized in that: The following steps are involved: Acquire multi-view image data, point cloud data, and historical frame video data of the target vehicle, and input the multi-view image data and the point cloud data into the Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor; Generating a road topology using the multi-information fusion tensor, and screening a drivable path of the vehicle intelligent body using the road topology to obtain a final planned path; The final planned path and the historical frame video data are input into a world model to generate future video data, and the future video data is used to update the environmental state of the vehicle agent in the next time step.

2. The end-to-end autonomous driving reinforcement learning method according to claim 1, characterized in that: Inputting the multi-view image data and the point cloud data into the Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor includes: extracting a plurality of bird's-eye view features from the multi-view image data and the point cloud data, and fusing the plurality of bird's-eye view features to obtain a global spatial semantic representation; The global spatial semantic representation is input into the Transformer architecture, and the global spatial semantic representation is encoded and decoded through the self-attention mechanism and the cross-attention mechanism to obtain the multi-information fusion tensor.

3. The end-to-end autonomous driving reinforcement learning method according to claim 1, characterized in that: The method of generating a road topology by using the multi-information fusion tensor, and screening a drivable path of the vehicle intelligent body by using the road topology to obtain a final planned path, includes: Performing mapping according to the multi-information fusion tensor to generate the road topology; The drivable path of the vehicle agent is encoded as a query tensor, and the query tensor is filtered using the road topology to obtain a final planned path of the vehicle agent.

4. The end-to-end autonomous driving reinforcement learning method according to claim 1, characterized in that: Inputting the final planned path and the historical frame video data into a world model to generate future video data, and using the future video data to update the environmental state of the vehicle agent at the next time step, includes: Performing discrete action marking on the final planned path to obtain a plurality of action prompt words; The multiple action prompt words and the historical frame video data are input into the world model to generate the future video data, and the future video data is used to update the environmental state of the vehicle agent in the next time step.

5. An end-to-end autonomous driving reinforcement learning device, characterized in that: include: A codec module is used to obtain multi-view image data, point cloud data, and historical frame video data of the target vehicle, and input the multi-view image data and the point cloud data into the Transformer architecture for encoding and decoding to obtain a multi-information fusion tensor; a screening module, configured to generate a road topology using the multi-information fusion tensor, and to screen the drivable paths of the vehicle agent using the road topology to obtain a final planned path; A generation module is used to input the final planned path and the historical frame video data into a world model to generate future video data, and use the future video data to update the environmental state of the vehicle agent in the next time step.

6. The end-to-end autonomous driving reinforcement learning device according to claim 5, characterized in that: The encoding and decoding module includes: An extraction and fusion unit, configured to extract a plurality of bird's-eye view features from the multi-view image data and the point cloud data, and fuse the plurality of bird's-eye view features to obtain a global spatial semantic representation; The encoding and decoding unit is used to input the global spatial semantic representation into the Transformer architecture, encode and decode the global spatial semantic representation through the self-attention mechanism and the cross-attention mechanism to obtain the multi-information fusion tensor.

7. The end-to-end autonomous driving reinforcement learning device according to claim 5, characterized in that: The screening module includes: A generating unit, performing mapping according to the multi-information fusion tensor to generate the road topology; A screening unit is used to encode the drivable path of the vehicle agent into a query tensor and filter the query tensor using the road topology to obtain a final planned path of the vehicle agent.

8. The end-to-end autonomous driving reinforcement learning device according to claim 5, characterized in that: The generation module includes: a discrete marking unit, configured to perform discrete marking of actions on the final planned path to obtain a plurality of action prompt words; A generation and update unit is used to input the multiple action prompt words and the historical frame video data into the world model to generate the future video data, and use the future video data to update the environmental state of the vehicle intelligent body in the next time step.

9. A vehicle, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the end-to-end autonomous driving reinforcement learning method according to any one of claims 1 to 4.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the end-to-end autonomous driving reinforcement learning method as described in any one of claims 1 to 4.

Citation Information

Cited By

  • End-to-end control system and method fusing multi-modal perception and strategy collaborative optimization mechanism

    CN121454923A

  • Intelligent driving method and device and storage medium

    CN122009220A