Automatic driving reinforcement teaching learning method, device, equipment, medium and product

By constructing a hybrid experience replay pool and performing probabilistic sampling in autonomous driving reinforcement learning, and integrating imitation learning with traditional reinforcement learning, the problems of training instability and reward function adjustment are solved, achieving more stable and efficient autonomous driving training.

CN121744597APending Publication Date: 2026-03-27CHINA FAW CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Reinforcement learning suffers from unstable training, difficulty in model convergence, slow training speed, and requires extensive reward function adjustments in the field of autonomous driving, which limits its application.

Method used

An abstract environment for reinforcement learning is constructed in a simulation platform to acquire and cluster multi-source trajectory data, build a hybrid experience replay pool, and combine probabilistic sampling for reinforcement teaching and learning of autonomous driving, thus integrating the advantages of imitation learning and traditional reinforcement learning.

Benefits of technology

This improves the stability of the autonomous driving learning process, reduces the workload of reward function adjustment, and enhances training efficiency and model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744597A_ABST
    Figure CN121744597A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic driving reinforcement teaching learning method, device and equipment, a medium and a product. The method comprises the following steps: in a simulation platform, constructing a reinforcement learning abstract environment containing a state, an action and a reward function based on an automatic driving scene; obtaining multi-source trajectory data in the reinforcement learning abstract environment, wherein the multi-source trajectory data comprises a preset teaching state transition trajectory, a different-orbit state transition trajectory and a same-orbit state transition trajectory; clustering the multi-source trajectory data to obtain a clustering result, performing probabilistic sampling according to the clustering result, and constructing a mixed experience playback pool; and carrying out automatic driving reinforcement teaching learning based on the mixed experience playback pool. According to the scheme, teaching knowledge used in imitation learning is integrated into a traditional reinforcement learning method, and the sampling and utilization mode of a probabilistic sampling optimization sample is combined, so that the learning ability of the traditional reinforcement learning method is kept, and meanwhile, the automatic driving learning process is more stable.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of autonomous driving technology, and in particular to an autonomous driving reinforcement teaching and learning method, device, equipment, medium and product. Background Technology

[0002] In recent years, reinforcement learning, with its core mechanism of learning optimal policies by interacting with the environment to obtain reward signals, has provided a feasible path to solve complex problems in autonomous driving. However, reinforcement learning training often relies heavily on adjusting reward functions, and reinforcement learning models often face problems such as training instability, difficulty in model convergence, and slow training, which limits its application in the field of autonomous driving. Summary of the Invention

[0003] This invention provides a method, apparatus, device, medium, and product for enhanced teaching and learning of autonomous driving, which can make the learning process of autonomous driving more stable while retaining the learning capabilities of traditional reinforcement learning methods.

[0004] In a first aspect, embodiments of the present invention provide an enhanced teaching and learning method for autonomous driving, comprising:

[0005] In the simulation platform, a reinforcement learning abstract environment containing state, action, and reward function is constructed based on the autonomous driving scenario;

[0006] In a reinforcement learning abstract environment, multi-source trajectory data is acquired, including preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories.

[0007] Clustering is performed on the multi-source trajectory data to obtain clustering results, and probabilistic sampling is performed based on the clustering results to construct a hybrid experience replay pool;

[0008] The autonomous driving enhancement teaching and learning is carried out based on the hybrid experience replay pool.

[0009] Furthermore, clustering is performed on the multi-source trajectory data to obtain clustering results, including:

[0010] Based on the states and actions involved in the multi-source trajectory data, similarity clustering is performed on the multi-source trajectory data to obtain clustering results.

[0011] Furthermore, based on the clustering results, probabilistic sampling is performed to construct a hybrid experience replay pool, including:

[0012] Based on the rewards involved in each type of data in the clustering results, determine the average reward level for each type of data;

[0013] The sampling probability of each type of data is determined based on the average reward level of each type of data.

[0014] Based on the sampling probability of various types of data, a hybrid experience replay pool is constructed.

[0015] Furthermore, based on the average reward level of various data types, the sampling probability of each data type is determined, including:

[0016] The average reward level of each type of data is input into the probability normalization function to determine the sampling probability of each type of data.

[0017] Furthermore, the states involved in the reinforcement learning abstract environment are composed of forward-looking camera images, bird's-eye view semantic segmentation images, and perception and localization information in the autonomous driving scenario;

[0018] The actions involved in the reinforcement learning abstract environment consist of steering control actions and throttle and brake control actions.

[0019] The reward function involved in the reinforcement learning abstract environment consists of a progress reward, a destination arrival reward, a centerline deviation penalty, and a dangerous driving penalty.

[0020] Furthermore, the different-track state transition trajectory and the same-track state transition trajectory are obtained by calling the environment reset function and the single-step iteration function to interact with the environment.

[0021] Secondly, embodiments of the present invention provide an enhanced teaching and learning device for autonomous driving, comprising:

[0022] The building blocks are used to construct reinforcement learning abstract environments containing states, actions, and reward functions based on autonomous driving scenarios in a simulation platform.

[0023] The acquisition module is used to acquire multi-source trajectory data in a reinforcement learning abstract environment. The multi-source trajectory data includes preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories.

[0024] The sampling module is used to cluster the multi-source trajectory data to obtain clustering results, and to perform probabilistic sampling based on the clustering results to construct a hybrid experience replay pool.

[0025] The learning module is used for enhanced teaching and learning of autonomous driving based on the hybrid experience replay pool.

[0026] Thirdly, embodiments of the present invention provide an electronic device, comprising:

[0027] At least one processor; and

[0028] A memory communicatively connected to the at least one processor; wherein,

[0029] The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method as described in the first aspect.

[0030] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions that cause a processor to execute the method described in the first aspect.

[0031] Fifthly, embodiments of the present invention provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in the first aspect.

[0032] The technical solution of this invention involves constructing a reinforcement learning abstract environment, including states, actions, and reward functions, based on an autonomous driving scenario within a simulation platform. Multi-source trajectory data is acquired within this abstract environment, including preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories. The multi-source trajectory data is clustered to obtain clustering results, and probabilistic sampling is performed based on these results to construct a hybrid experience replay pool. Autonomous driving reinforcement teaching is then conducted based on this hybrid experience replay pool. This solution integrates the teaching knowledge used in imitation learning into traditional reinforcement learning methods, combines probabilistic sampling to optimize sample sampling and utilization, and integrates the advantages of both imitation learning and traditional reinforcement learning methods. This allows for a more stable autonomous driving learning process while retaining the learning capabilities of traditional reinforcement learning methods.

[0033] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a flowchart of an enhanced teaching and learning method for autonomous driving provided according to Embodiment 1 of the present invention;

[0036] Figure 2This is a schematic diagram of the structure of an autonomous driving enhanced teaching and learning device according to Embodiment 2 of the present invention;

[0037] Figure 3 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation

[0038] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0039] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0040] Example 1

[0041] Figure 1 This is a flowchart of an enhanced teaching and learning method for autonomous driving according to Embodiment 1 of the present invention. This embodiment is applicable to situations involving enhanced teaching and learning for autonomous driving. The method can be executed by an enhanced teaching and learning device for autonomous driving, which can be implemented in software and / or hardware and integrated into an electronic device. Further, the electronic device includes, but is not limited to, computers, laptops, smartphones, servers, etc.

[0042] like Figure 1 As shown, the method includes:

[0043] S110. In the simulation platform, a reinforcement learning abstract environment containing state, action, and reward function is constructed based on the autonomous driving scenario.

[0044] In this embodiment of the invention, the simulation platform can be any simulation platform that supports autonomous driving learning, such as a high-fidelity, high-precision simulation platform developed specifically for autonomous driving tasks. The simulation platform provides a simulation environment equipped with a physics engine, and on this basis, it realizes the construction and rendering of models of vehicles, pedestrians, buildings, roads, traffic lights, weather, etc., achieving a digital twin of the real world within the virtual platform.

[0045] In the simulation platform, an abstraction of the reinforcement learning environment for autonomous driving is performed. This involves determining the states, actions, and reward functions involved in autonomous driving reinforcement learning, and then constructing a runnable reinforcement learning abstract environment through the design logic of the environment abstraction. This environment allows the agent to learn. The state can be understood as a comprehensive description of the vehicle's own state and surrounding environment information using quantifiable parameters. The core is to let the agent know the vehicle's current state, and the state must cover key information strongly related to driving decisions, avoiding redundant data. Actions can be understood as the set of operations that the vehicle can perform to change its current state. These must match the control logic of a real vehicle, and each action directly affects the subsequent state. The reward function can be understood as a reward-related function. The reward is a quantified feedback signal given based on the state change after the action is executed. The core is to guide the agent to learn behaviors that conform to safety, efficiency, and rules. Positive rewards encourage correct operations, while negative rewards (i.e., penalties) prohibit dangerous or illegal operations.

[0046] In one embodiment, the state involved in the reinforcement learning abstract environment consists of forward-looking camera images, bird's-eye view semantic segmentation images, and perception and localization information in the autonomous driving scenario; the actions involved in the reinforcement learning abstract environment consist of steering control actions and throttle and brake control actions; and the reward function involved in the reinforcement learning abstract environment consists of travel reward, destination arrival reward, centerline deviation penalty, and dangerous driving penalty.

[0047] Regarding the state, the forward-looking camera image in the autonomous driving scenario is a real-time image captured by a camera installed in front of the vehicle, providing the most direct visual view of the front environment for autonomous driving; the bird's-eye view semantic segmentation image is structured data that converts images from multiple cameras such as forward-looking and side-looking cameras into a top-down view through algorithms, and labels each pixel in the image with categories such as drivable area, lane lines, vehicles, pedestrians, obstacles, etc.; perception and localization information is the structured result extracted after processing the raw data collected by the forward-looking camera image, bird's-eye view semantic segmentation image and other sensors through algorithms. It includes two parts: the vehicle's own state and the surrounding environment information, such as vehicle coordinates, speed, distance from the road centerline, distance from obstacles in front, etc.

[0048] For actions, a two-dimensional action output is implemented, corresponding to the two major modules in vehicle control: the steering control module and the throttle and brake control module. The core objective of steering control actions is to change the vehicle's direction of travel or correct its lane position, ensuring that the vehicle travels along the expected path (such as straight lines, curves, and lane changes), mainly achieved by controlling the steering wheel angle. The core objective of throttle and brake control actions is to adjust the vehicle's speed (acceleration, deceleration, and stopping), ensuring that the vehicle maintains a safe distance from the vehicle in front and meets speed limits, mainly achieved by controlling the throttle opening (accelerator) or the braking system (brake).

[0049] Regarding the reward function, the travel reward is used to reward the vehicle for moving forward and prevent the vehicle from standing still or deviating from the target direction; the destination arrival reward is used to reinforce the core objective of reaching the destination and prevent the vehicle from driving aimlessly in the middle; the centerline deviation penalty is used to punish the vehicle according to the degree of deviation when the distance between the vehicle's center and the lane centerline exceeds a preset safety threshold during the vehicle's movement; and the dangerous driving penalty is used to punish the vehicle when it collides or violates core safety rules.

[0050] S120. Acquire multi-source trajectory data in a reinforcement learning abstract environment. The multi-source trajectory data includes preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories.

[0051] After constructing an abstract environment for reinforcement learning in a simulation platform, the agent can be manipulated to interact with the environment within the specific implementation of the reinforcement learning network to obtain state transition trajectories. These state transition trajectories refer to the sequence of states, actions, and rewards formed by the agent during its interaction with the environment, starting from an initial state and executing actions.

[0052] In this embodiment of the invention, to combine the advantages of imitation learning training, a data acquisition method integrating multiple data sources is proposed. These multiple sources specifically include three types of data: instruction data, dissimilar data, and concordant data. Dissimilar data adopts the widely used definition of an experience replay pool, which includes state transition trajectories collected in the earliest stages of training (i.e., dissimilar state transition trajectories). Concordant data consists of state transition trajectories generated more closely to the current network parameters (i.e., concordant state transition trajectories). Instruction data can include preset instruction state transition trajectories, which are state-action-reward sequences planned or recorded in advance by experts or algorithms based on task objectives. Each step of the state-action mapping within the trajectory represents a correct decision that meets safety and efficiency goals, essentially providing decision-making references for the agent. The introduction of instruction data can combine the advantages of imitation learning training, reducing training difficulty and improving training stability. It can also import data from two sources: human driving data and driving data generated by the convergence network, thus satisfying various application scenarios.

[0053] It should be noted that the proportions and sizes of the data in the preset teaching state transition trajectory, the different track state transition trajectory, and the same track state transition trajectory can be freely and flexibly defined according to actual training needs to achieve the best training effect.

[0054] In one embodiment, the different-track state transition trajectory and the same-track state transition trajectory are obtained by calling an environment reset function and a single-step iteration function to interact with the environment.

[0055] Regardless of whether the trajectory is on a different track or the same track, the acquisition process relies on the interaction between the environment reset function and the single-step iteration function. The only difference lies in the source of the action. Specifically: the environment state is initialized, i.e., the initial state is obtained, and the environment reset function is called to reset the internal state of the environment, providing a starting point for trajectory generation; based on the single-step iteration function, the action to be executed in the current state is input. After the action is executed, the environment returns the next state, the immediate reward, and the termination flag, and the single-step iteration loop is performed to realize trajectory generation. The action source for different tracks is historical actions, and the action source for the same track is the action under the current policy.

[0056] S130. Cluster the multi-source trajectory data to obtain clustering results, and perform probabilistic sampling based on the clustering results to construct a hybrid experience replay pool.

[0057] In this step, the multi-source trajectory data, including preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories, can be comprehensively considered. A clustering algorithm is used to cluster the state transition trajectories based on similarity, resulting in multiple data categories. The average reward level for each category of data is determined. Based on the average reward level, the sampling probability for each category is determined, i.e., the reward level is converted into sampling weights through quantization rules. According to the sampling probability of each category, probabilistic sampling is performed on each category of data until a hybrid experience replay pool of the desired size is obtained, thus constructing the hybrid experience replay pool. The size of the hybrid experience replay pool, i.e., the desired size, can be customized according to actual application needs and is not limited here.

[0058] S140. Perform autonomous driving reinforcement teaching and learning based on the hybrid experience replay pool.

[0059] After the hybrid experience replay pool is constructed, autonomous driving reinforcement teaching and learning can be carried out based on the hybrid experience replay pool. Specifically, batch training data can be obtained from the hybrid experience replay pool with certain weights (not limited here), that is, multiple state transition trajectories for training reinforcement learning networks can be obtained; reinforcement learning networks can be trained using the obtained batch training data, such as calculating the loss based on the batch training data, and updating the network parameters by minimizing the loss through gradient descent; by iteratively interacting, sampling, and training, the reinforcement learning network can be continuously optimized to obtain an autonomous driving decision-making algorithm that meets expectations.

[0060] It should be noted that in actual training, the mixed experience replay pool is not updated every time a new state transition trajectory is collected. Instead, a delayed update method is adopted, that is, a new round of probability calculation and mixed experience replay pool collection is performed only after a certain number of collection steps, in order to reduce the consumption of computing resources.

[0061] It should be noted that the present invention does not limit the specific network used for reinforcement learning, as long as it can be trained on the simulation platform of the present invention to obtain an autonomous driving decision-making algorithm that meets the expectations.

[0062] The technical solution of this invention involves constructing a reinforcement learning abstract environment, including states, actions, and reward functions, based on an autonomous driving scenario within a simulation platform. Multi-source trajectory data is acquired within this abstract environment, including preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories. The multi-source trajectory data is clustered to obtain clustering results, and probabilistic sampling is performed based on these results to construct a hybrid experience replay pool. Autonomous driving reinforcement teaching is then conducted based on this hybrid experience replay pool. This solution integrates the teaching knowledge used in imitation learning into traditional reinforcement learning methods, combines probabilistic sampling to optimize sample sampling and utilization, and integrates the advantages of both imitation learning and traditional reinforcement learning methods. This allows for a more stable autonomous driving learning process while retaining the learning capabilities of traditional reinforcement learning methods.

[0063] It's worth noting that the increased stability of the autonomous driving learning process is due to the integration of taught knowledge with traditional reinforcement learning methods. Taught knowledge can be used to extract a reasonable reward function during imitation learning, eliminating the need for manually setting and adjusting the reward function.

[0064] In one embodiment, clustering the multi-source trajectory data to obtain clustering results includes: performing similarity clustering on the multi-source trajectory data based on the states and actions involved in the multi-source trajectory data to obtain clustering results.

[0065] The multi-source trajectory data includes preset teaching state transition trajectories, different track state transition trajectories, and same track state transition trajectories. Based on the similarity of the states and actions included in each state transition trajectory, clustering results are obtained. In the clustering results, the similarity of the states and actions of different state transition trajectories under each category exceeds the similarity threshold.

[0066] In one embodiment, probabilistic sampling based on the clustering results to construct a hybrid experience replay pool includes: determining the average reward level of each type of data based on the rewards involved in each type of data in the clustering results; determining the sampling probability of each type of data based on the average reward level of each type of data; and sampling each type of data based on the sampling probability of each type of data to construct a hybrid experience replay pool.

[0067] For each data category in the clustering results, the average reward level of each state transition trajectory in that category is taken as the average level. Based on the average reward level of each data category, the sampling probability of each data category is determined. For example, the sampling probability of a certain data category with a high average reward level is high. Based on the sampling probability of each data category, probabilistic sampling is performed on each data category until a mixed experience replay pool of the expected size is obtained, thus realizing the construction of the mixed experience replay pool.

[0068] In one embodiment, determining the sampling probability of each type of data based on the average reward level of each type of data includes: inputting the average reward level of each type of data into a probability normalization function to determine the sampling probability of each type of data.

[0069] There are no restrictions on the probability normalization function; it can be any function that can achieve probability normalization, such as the Softmax function. By inputting the average reward level of any clustered category into the probability normalization function, the sampling probability of that category of data will be output.

[0070] Optionally, during network training in this invention, the following hyperparameter settings can be used: 2048 random exploration steps; 512 saved state transition trajectories on different tracks; 512 saved state transition trajectories on the same track; 1024 saved state transition trajectories for teaching; 4 total cluster categories; and 256 mixed experience replay pool size.

[0071] This invention proposes an end-to-end reinforcement learning training method for autonomous driving that is compatible with current mainstream simulation platforms. This method employs a unique, probabilistically optimized training sample sampling and utilization approach, enabling the incorporation of taught knowledge to increase the efficiency of network model training and improve model performance. Simultaneously, a general method is proposed to implement an abstract reinforcement learning environment that can be used for training autonomous driving algorithms, defining environmental states, actions, reward functions, etc. By training a reinforcement learning network using the framework of this invention, it can control a vehicle in an autonomous driving simulation platform to perform end-to-end control tasks. This allows for convenient training of reinforcement learning networks on mainstream simulation platforms, solving practical autonomous driving problems.

[0072] Example 2

[0073] Figure 2 This is a schematic diagram of an enhanced teaching and learning device for autonomous driving according to Embodiment 2 of the present invention. This embodiment is applicable to situations where enhanced teaching and learning for autonomous driving is implemented, such as... Figure 2 As shown, the specific structure of the device includes:

[0074] Module 21 is used to build a reinforcement learning abstract environment containing state, action and reward functions based on autonomous driving scenarios in the simulation platform;

[0075] The acquisition module 22 is used to acquire multi-source trajectory data in a reinforcement learning abstract environment. The multi-source trajectory data includes preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories.

[0076] The sampling module 23 is used to cluster the multi-source trajectory data to obtain clustering results, and to perform probabilistic sampling based on the clustering results to construct a hybrid experience replay pool.

[0077] Learning module 24 is used for autonomous driving reinforcement teaching learning based on the hybrid experience replay pool.

[0078] The autonomous driving reinforcement teaching and learning device provided in this embodiment constructs an abstract reinforcement learning environment, including states, actions, and reward functions, based on an autonomous driving scenario within a simulation platform using a construction module. An acquisition module acquires multi-source trajectory data within this abstract environment, including preset teaching state transition trajectories, cross-track state transition trajectories, and same-track state transition trajectories. A sampling module clusters the multi-source trajectory data to obtain clustering results and performs probabilistic sampling based on these results to construct a hybrid experience replay pool. A learning module then performs autonomous driving reinforcement teaching and learning based on this hybrid experience replay pool. This solution integrates the teaching knowledge used in imitation learning into traditional reinforcement learning methods, combines probabilistic sampling to optimize sample sampling and utilization, and integrates the advantages of both imitation learning and traditional reinforcement learning methods. This allows for a more stable autonomous driving learning process while retaining the learning capabilities of traditional reinforcement learning methods.

[0079] Furthermore, the sampling module 23 is specifically used for:

[0080] Based on the states and actions involved in the multi-source trajectory data, similarity clustering is performed on the multi-source trajectory data to obtain clustering results.

[0081] Furthermore, the sampling module 23 is specifically used for:

[0082] Based on the rewards involved in each type of data in the clustering results, determine the average reward level for each type of data;

[0083] The sampling probability of each type of data is determined based on the average reward level of each type of data.

[0084] Based on the sampling probability of various types of data, a hybrid experience replay pool is constructed.

[0085] Furthermore, the sampling module 23 is specifically used for:

[0086] The average reward level of each type of data is input into the probability normalization function to determine the sampling probability of each type of data.

[0087] Furthermore, the states involved in the reinforcement learning abstract environment are composed of forward-looking camera images, bird's-eye view semantic segmentation images, and perception and localization information in the autonomous driving scenario;

[0088] The actions involved in the reinforcement learning abstract environment consist of steering control actions and throttle and brake control actions.

[0089] The reward function involved in the reinforcement learning abstract environment consists of a progress reward, a destination arrival reward, a centerline deviation penalty, and a dangerous driving penalty.

[0090] Furthermore, the different-track state transition trajectory and the same-track state transition trajectory are obtained by calling the environment reset function and the single-step iteration function to interact with the environment.

[0091] The autonomous driving enhanced teaching and learning device provided in the embodiments of the present invention can execute the autonomous driving enhanced teaching and learning method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0092] Example 3

[0093] Figure 3 This is a schematic diagram of the structure of an electronic device implementing embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0094] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 performs various appropriate actions and processes based on the computer programs stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0095] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0096] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as autonomous driving reinforcement teaching learning methods.

[0097] In some embodiments, the autonomous driving reinforcement teaching and learning method can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the autonomous driving reinforcement teaching and learning method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the autonomous driving reinforcement teaching and learning method by any other suitable means (e.g., by means of firmware).

[0098] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0099] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0100] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0101] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0102] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0103] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0104] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0105] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for enhanced teaching and learning in autonomous driving, characterized in that, include: In the simulation platform, a reinforcement learning abstract environment containing state, action, and reward function is constructed based on the autonomous driving scenario; In a reinforcement learning abstract environment, multi-source trajectory data is acquired, including preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories. Clustering is performed on the multi-source trajectory data to obtain clustering results, and probabilistic sampling is performed based on the clustering results to construct a hybrid experience replay pool; The autonomous driving enhancement teaching and learning is carried out based on the hybrid experience replay pool.

2. The method according to claim 1, characterized in that, Clustering the multi-source trajectory data yields clustering results, including: Based on the states and actions involved in the multi-source trajectory data, similarity clustering is performed on the multi-source trajectory data to obtain clustering results.

3. The method according to claim 1, characterized in that, Based on the clustering results, probabilistic sampling is performed to construct a hybrid experience replay pool, including: Based on the rewards involved in each type of data in the clustering results, determine the average reward level for each type of data; The sampling probability of each type of data is determined based on the average reward level of each type of data. Based on the sampling probability of various types of data, a hybrid experience replay pool is constructed.

4. The method according to claim 3, characterized in that, Based on the average reward level of each type of data, determine the sampling probability of each type of data, including: The average reward level of each type of data is input into the probability normalization function to determine the sampling probability of each type of data.

5. The method according to claim 1, characterized in that, The states involved in the reinforcement learning abstract environment consist of forward-looking camera images, bird's-eye view semantic segmentation images, and perception and localization information in the autonomous driving scenario. The actions involved in the reinforcement learning abstract environment consist of steering control actions and throttle and brake control actions. The reward function involved in the reinforcement learning abstract environment consists of a progress reward, a destination arrival reward, a centerline deviation penalty, and a dangerous driving penalty.

6. The method according to claim 1, characterized in that, The different-track state transition trajectory and the same-track state transition trajectory are obtained by calling the environment reset function and the single-step iteration function to interact with the environment.

7. An automated driving enhanced teaching and learning device, characterized in that, include: The building blocks are used to construct reinforcement learning abstract environments containing states, actions, and reward functions based on autonomous driving scenarios in a simulation platform. The acquisition module is used to acquire multi-source trajectory data in a reinforcement learning abstract environment. The multi-source trajectory data includes preset teaching state transition trajectories, different-track state transition trajectories, and same-track state transition trajectories. The sampling module is used to cluster the multi-source trajectory data to obtain clustering results, and to perform probabilistic sampling based on the clustering results to construct a hybrid experience replay pool. The learning module is used for enhanced teaching and learning of autonomous driving based on the hybrid experience replay pool.

8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-6.