End-to-End Autonomous Driving Model Training Method and Device

By using the bird's-eye view semantic segmentation and features output from the student model in the end-to-end autonomous driving model, and combined with the compensation features calculated by the adapter, input the teacher model to output the control signal, the problem of causal inversion in behavioral cloning training is solved, and the output accuracy of the model is improved.

CN116842386BActive Publication Date: 2025-06-10SHANGHAI ARTIFICIAL INTELLIGENCE INNOVATION CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310811453.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-04
Publication Date
2025-06-10
Estimated Expiration
2043-07-04

AI Technical Summary

Technical Problem

Existing end-to-end autonomous driving models are trained using behavioral cloning, which can easily lead to causal inversion problems and affect the accuracy of the output results.

Method used

By inputting pre-acquisitioned samples into the student model, outputting aerial view semantic segmentation and aerial view features, and inputting them into the teacher model, outputting control signals. The preset adapter is used to calculate the difference between the intermediate features and the bird's-eye view features of each layer in the teacher model, obtain the compensation features, and use them as the input to the next layer of the teacher model, and finally output the control signal.

Benefits of technology

It effectively avoids the problem of causal inversion and improves the accuracy and feasibility of the output results of the end-to-end autonomous driving model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842386B_ABST
    Figure CN116842386B_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer vision technology, and discloses an end-to-end autonomous driving model training method and device. The method inputs pre-collected samples into a student model, outputs bird's-eye view semantic segmentation and bird's-eye view features; the samples include original images and corresponding point cloud data; then the bird's-eye view semantic segmentation and the bird's-eye view features are input into a teacher model, and a control signal is output to avoid the causal inversion problem caused by behavior cloning as much as possible, thereby improving the accuracy of the output results of the end-to-end autonomous driving model and improving the feasibility of end-to-end autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer vision technology, and particularly to an end-to-end autonomous driving model training method and apparatus. Background Art

[0002] Existing end-to-end autonomous driving models all adopt the "teacher model - student model" framework: the input of the teacher model is the perception ground truth (the real positions and states of surrounding vehicles, pedestrians, traffic lights, lane lines, etc.). The teacher model is generally rule-based or reinforcement learning-based. Because there is perception ground truth and the driving performance is good, it is used to collect data in the simulator. The input of the student model is sensor data (images captured by cameras, point clouds scanned by lidar, etc.). Generally, the end-to-end training is carried out by means of behavior cloning, that is, a neural network is directly used to map the sensor input to the control signal of the vehicle, and the training data used is the data collected by the teacher model. However, this behavior cloning technology has defects and there will be a problem of causal inversion. For example, there is data in the dataset of stopping before a red light and waiting for the green light to pass. In most of this data, there are surrounding vehicles coming; therefore, based on the behavior cloning method, the tendency is to learn the logic that when the surrounding vehicles start, the target vehicle starts following; but the real logic that should be learned is that when the traffic light turns from red to green, the target vehicle starts. The reason for the learning error is that the vehicles in the pictures and lidar are very large, while the traffic lights are very small, and the neural network tends to learn the easily observable phenomena. The result is that in actual tests, if there are no other vehicles at the intersection, the target vehicle will be at a loss and then get stuck.

[0003] Therefore, using the behavior cloning method for training in the existing end-to-end model easily leads to learning errors in the end-to-end autonomous driving method and affects the accuracy of the output results. Summary of the Invention

[0004] An embodiment of this application provides an end-to-end autonomous driving model training method to solve the problem in the prior art that using the behavior cloning method for training in the end-to-end model easily leads to learning errors in the end-to-end autonomous driving method and affects the accuracy of the output results.

[0005] Correspondingly, an embodiment of this application also provides an end-to-end autonomous driving model training apparatus, an electronic device, and a computer-readable storage medium to ensure the implementation and application of the above method.

[0006] To solve the above technical problems, an embodiment of this application discloses an end-to-end autonomous driving model training method, and the method includes:

[0007] Input the pre-collected samples into the student model, and output the bird's-eye view semantic segmentation and bird's-eye view features; the samples include the original images and the corresponding point cloud data;

[0008] Input the semantic segmentation of the bird's-eye view and the features of the bird's-eye view into the teacher model, and output to obtain a control signal.

[0009] Preferably, the teacher model includes multiple layers; the step of inputting the semantic segmentation of the bird's-eye view and the features of the bird's-eye view into the teacher model and outputting to obtain a control signal includes:

[0010] Input the semantic segmentation of the bird's-eye view into the teacher model;

[0011] Use a preset adapter to calculate the difference between the intermediate features output by each layer in the teacher model and the features of the bird's-eye view to obtain a compensation feature;

[0012] Take the compensation feature as the input of the next layer of the teacher model;

[0013] Take the output of the last layer of the teacher model as the control signal.

[0014] Preferably, the step of using a preset adapter to calculate the difference between the intermediate features output by each layer in the teacher model and the features of the bird's-eye view to obtain a compensation feature includes:

[0015] Obtain the intermediate features output by the corresponding layer in the teacher model;

[0016] Perform downsampling on the features of the bird's-eye view to obtain a downsampled feature with the same size as the intermediate feature;

[0017] Use the adapter to process the downsampled feature and the intermediate feature to obtain the compensation feature.

[0018] Preferably, the step of using the adapter to process the downsampled feature and the intermediate feature to obtain the compensation feature includes:

[0019] Concatenate the downsampled feature and the intermediate feature and input them into the adapter, and output to obtain the compensation feature after being processed by a convolutional neural network.

[0020] Preferably, the sample also includes a standard signal. For an incorrect sample, the method further includes:

[0021] Apply a mask to the compensation feature, and calculate the gap between the control signal and the standard control signal as a loss;

[0022] Use the backpropagation of the loss to train the adapter;

[0023] Wherein, the incorrect sample is a sample in which the predicted control signal is inconsistent with the output of the standard signal;

[0024] The prediction control signal is output after the input bird's-eye view semantic segmentation and the bird's-eye view features are processed in sequence by multiple layers of the teacher model.

[0025] Preferably, the student model is constructed by sampling the BEVFusion method or the BEVFormer method.

[0026] Preferably, the teacher model is constructed by sampling Roach method or LBC method.

[0027] The present application also discloses an end-to-end autonomous driving model training device, the device comprising:

[0028] A student model training module is used to input pre-collected samples into the student model and output bird's-eye view semantic segmentation and bird's-eye view features; the samples include original images and corresponding point cloud data;

[0029] The teacher model training module is used to input the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and output a control signal.

[0030] An embodiment of the present application also discloses an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, one or more of the methods described in the embodiments of the present application are implemented.

[0031] The embodiments of the present application further disclose a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, one or more methods described in the embodiments of the present application are implemented.

[0032] In an embodiment of the present application, pre-collected samples are input into a student model, and bird's-eye view semantic segmentation and bird's-eye view features are output; the samples include an original image and corresponding point cloud data; then the bird's-eye view semantic segmentation and the bird's-eye view features are input into a teacher model, and a control signal is output to avoid the causal inversion problem caused by behavioral cloning as much as possible, thereby improving the accuracy of the output results of the end-to-end autonomous driving model and improving the feasibility of end-to-end autonomous driving.

[0033] Additional aspects and advantages of the embodiments of the present application will be given in the following description, which will become apparent from the following description or be understood through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0035] Figure 1Flowchart of the end-to-end autonomous driving model training method provided by the embodiments of the present application;

[0036] Figure 2 Schematic diagram of the original image provided by the embodiments of the present application;

[0037] Figure 3 Schematic diagram of the point cloud data provided by the embodiments of the present application;

[0038] Figure 4 Schematic diagram of the bird's-eye view semantic segmentation provided by the embodiments of the present application;

[0039] Figure 5 Flowchart of the training method of the teacher model for alignment using an adapter provided by the embodiments of the present application;

[0040] Figure 6 Schematic diagram of the structure of the end-to-end autonomous driving model training device provided by the embodiments of the present application;

[0041] Figure 7 Schematic diagram of the structure of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0042] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0043] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present application means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their combinations. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.

[0044] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0045] The solution provided by the embodiments of the present application can be executed by any electronic device, such as a terminal device or a server. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this. Regarding the technical problems existing in the prior art, the end-to-end autonomous driving model training method, device and electronic device provided by the present application are intended to solve at least one of the technical problems in the prior art.

[0046] The following uses specific embodiments to detail the technical solution of the present application and how the technical solution of the present application solves the above technical problems. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the embodiments of the present application with reference to the accompanying drawings.

[0047] The embodiments of the present application provide a possible implementation manner, such as Figure 1 shown, a flowchart of an end-to-end autonomous driving model training method is provided. This solution can be executed by any electronic device, and optionally, can be executed on the server side or the terminal device.

[0048] As Figure 1 shown, the method may include the following steps:

[0049] Step 101, input the pre-collected samples into the student model, and output the bird's-eye view semantic segmentation and bird's-eye view features; the samples include the original image and the corresponding point cloud data;

[0050] The original image can be a surround-view camera picture, such as Figure 2 shown; the point cloud data can be a lidar point cloud image, such as Figure 3 shown. Inputting the surround-view camera picture and the lidar point cloud image into the pre-constructed student model can output the bird's-eye view semantic segmentation and bird's-eye view features, and the bird's-eye view semantic segmentation is as Figure 4 shown.

[0051] Step 102, input the bird's-eye view semantic segmentation and the bird's-eye view features into a teacher model, and output a control signal.

[0052] In the embodiment of the present application, the semantic segmentation of the bird's-eye view under the bird's-eye view can be used as the input of the teacher model to output the control signal of the vehicle. The specific operation of the vehicle can be controlled by the control signal output by the teacher model.

[0053] In an embodiment of the present application, pre-collected samples are input into a student model, and bird's-eye view semantic segmentation and bird's-eye view features are output; the samples include an original image and corresponding point cloud data; then the bird's-eye view semantic segmentation and the bird's-eye view features are input into a teacher model, and a control signal is output to avoid the causal inversion problem caused by behavioral cloning as much as possible, thereby improving the accuracy of the output results of the end-to-end autonomous driving model and improving the feasibility of end-to-end autonomous driving.

[0054] In an optional embodiment, the student model is constructed by sampling the BEVFusion method or the BEVFormer method.

[0055] In an optional embodiment, the teacher model is constructed by sampling Roach method or LBC method.

[0056] The Roach method or the LBC method both take the bird's-eye view semantic segmentation under the bird's-eye view as input, so the bird's-eye view semantic segmentation can be the bird's-eye view semantic segmentation under the bird's-eye view.

[0057] In an optional embodiment, the teacher model includes a plurality of layers; the step of inputting the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and outputting a control signal includes:

[0058] Inputting the bird's-eye view semantic segmentation into the teacher model;

[0059] Using a preset adapter, a difference between the intermediate features output by each layer in the teacher model and the bird's-eye view features is calculated to obtain compensation features;

[0060] Using the compensated features as input to the next layer of the teacher model;

[0061] The output of the last layer of the teacher model is used as the control signal.

[0062] Directly connect the output of the student model to the input of the teacher model. Since there are errors in the bird's-eye view semantic segmentation obtained by training the student model, the errors will further accumulate during the transmission between multiple layers (convolutional layers) inside the teacher model, affecting the accuracy of the output control signal. Therefore, the driving effect is poor. And there will also be certain errors in the teacher model itself. Therefore, in the embodiments of the present application, an adapter is used to calculate the error of each layer in the teacher model (that is, the difference between the intermediate feature output by this layer and the bird's-eye view feature). After obtaining the compensation feature, it is then output to the next layer. Using the adapter to make up for the error layer by layer can improve the accuracy of the output control signal and improve the driving effect.

[0063] In an optional embodiment, the using a preset adapter to calculate the difference between the intermediate feature output by each layer of the teacher model and the bird's-eye view feature to obtain a compensation feature includes:

[0064] Obtain the intermediate feature output by the corresponding layer in the teacher model;

[0065] Perform downsampling processing on the bird's-eye view feature to obtain a downsampled feature with the same size as the intermediate feature;

[0066] Use the adapter to process the downsampled feature and the intermediate feature to obtain the compensation feature.

[0067] In an optional embodiment, the using the adapter to process the downsampled feature and the intermediate feature to obtain the compensation feature includes:

[0068] Concatenate the downsampled feature and the intermediate feature and input them into the adapter, and after being processed by a convolutional neural network, output to obtain the compensation feature.

[0069] Exemplarily, Figure 5 shows a flowchart of a training method for a teacher model using an adapter for alignment. As Figure 5 shown, the bird's-eye view semantic segmentation can be the bird's-eye view semantic segmentation, and the bird's-eye view feature can be the bird's-eye view feature. The teacher model includes multiple layers (convolutional layers), such as teacher model module 1, teacher model module 2,..., teacher model module N. An adapter is configured between each layer, such as adapter 1, adapter 2, and each adapter consists of multiple convolutional layers.

[0070] When training the teacher model, input the bird's-eye view semantic segmentation into teacher model module 1, and after convolutional processing, output the intermediate feature H 1 ; downsample the bird's-eye view feature through multiple convolutional layers to the same size as the intermediate feature H 1 to obtain the downsampled feature F 1 ; the intermediate feature H 1and the downsampled feature F 1 After concatenation, it is input into Adapter 1. After being processed by the convolutional neural network in Adapter 1, the compensated feature H corresponding to Teacher Model Module 1 is obtained as the output. Adpt ; Then, the compensated feature H Adpt is input into Teacher Model Module 2.

[0071] After convolutional processing, Teacher Model Module 2 outputs the intermediate feature H 2 ; The bird's-eye view feature is downsampled through multiple layers of convolution to the same size as the intermediate feature H 2 to obtain the downsampled feature F 2 ; Then, the intermediate feature H 2 and the downsampled feature F 2 are concatenated and input into Adapter 2. After being processed by the convolutional neural network in Adapter 2, the compensated feature corresponding to Teacher Model Module 2 is obtained as the output.

[0072] By analogy, for Teacher Model Module N - 1, the corresponding compensated feature is obtained The compensated feature is input into Teacher Model Module N. After convolutional processing by Teacher Model Module N, the predicted control signal is output.

[0073] In an optional embodiment, the sample further includes a standard signal. For an error sample, the method further includes:

[0074] Applying a mask to the compensated feature and calculating the gap between the control signal and the standard control signal as the loss;

[0075] Training the adapter using the backpropagation of the loss;

[0076] wherein the error sample is a sample in which the predicted control signal is inconsistent with the output of the standard signal;

[0077] The predicted control signal is output after the input bird's-eye view semantic segmentation and the bird's-eye view feature are processed layer by layer through multiple layers of the teacher model.

[0078] The predicted control signal in the embodiment of the present application directly connects the output of the student model to the input of the teacher model, and is the signal output by the teacher model without adapter compensation. Specifically, the original image and the corresponding point cloud data are input into the first layer of the teacher model, and the output of the first layer is used as the input of the second layer, and so on. The predicted control signal is output through the last layer of the teacher model.

[0079] In the embodiment of the present application, a mask is applied to the compensated feature so that the compensated feature does not perform backpropagation.

[0080] Exemplarily, if there is a pedestrian in front, the standard control signal is to execute braking; while the control signal output by the teacher model is to step on the accelerator. The accelerator can be mapped to the interval of 0-1 according to the degree of depression, and the brake can also be mapped to the interval of -1-0 according to the degree of depression. The standard control signal and the control signal output by the teacher model are subtracted, and a loss is formed after subtraction. The neural network can be updated by the backpropagation of the loss to train the teacher model.

[0081] During the training process, the generation of compensation features can be guided by the backpropagation of the above loss.

[0082] For the teacher model and the student model obtained by the end-to-end autonomous driving model training method in the embodiments of the present application, tests were carried out in the Carla simulator, and two groups of publicly available end-to-end driving evaluation specifications: town05long and longest6 were used, and both were significantly better than the existing methods; at the same time, compared with the baseline model without using the adapter, the increase was relatively large.

[0083] Based on the same principle as the method provided in the embodiments of the present application, the embodiments of the present application also provide an end-to-end autonomous driving model training device, as Figure 6 shown, the device includes:

[0084] A student model training module 601, configured to input pre-collected samples into the student model and output a bird's-eye view semantic segmentation and bird's-eye view features; the samples include original images and corresponding point cloud data;

[0085] The original image can be a panoramic camera picture, as Figure 2 shown; the point cloud data can be a lidar point cloud image, as Figure 3 shown. Inputting the panoramic camera picture and the lidar point cloud image into the pre-constructed student model can output a bird's-eye view semantic segmentation and bird's-eye view features, and the bird's-eye view semantic segmentation is as Figure 4 shown.

[0086] A teacher model training module 602, configured to input the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and output a obtained control signal.

[0087] In the embodiments of the present application, the bird's-eye view semantic segmentation under the bird's-eye view can be used as the input of the teacher model, and the control signal of the vehicle is output. The specific operation of the vehicle can be controlled by the control signal output by the teacher model.

[0088] In the embodiments of the present application, the student model training module inputs the pre-collected samples into the student model and outputs the bird's-eye view semantic segmentation and the bird's-eye view features; the samples include the original images and the corresponding point cloud data; the teacher model training module inputs the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and outputs a control signal, which can avoid the problem of causal inversion caused by behavior cloning as much as possible, thereby improving the accuracy of the output result of the end-to-end autonomous driving model and the feasibility of end-to-end autonomous driving.

[0089] In an optional embodiment, the teacher model includes multiple layers; the teacher model training module 602 includes:

[0090] The first teacher model training sub-module is used to input the bird's-eye view semantic segmentation into the teacher model;

[0091] The second teacher model training sub-module is used to calculate the difference between the intermediate features output by each layer of the teacher model and the bird's-eye view features by using a preset adapter to obtain a compensation feature;

[0092] The third teacher model training sub-module is used to use the compensation feature as the input of the next layer of the teacher model;

[0093] The fourth teacher model training sub-module is used to use the output of the last layer of the teacher model as the control signal.

[0094] In an optional embodiment, the second teacher model training sub-module includes:

[0095] The first teacher model training unit is used to obtain the intermediate features output by the corresponding layer in the teacher model;

[0096] The second teacher model training unit is used to perform downsampling processing on the bird's-eye view features to obtain a downsampled feature with the same size as the intermediate feature;

[0097] The third teacher model training unit is used to process the downsampled feature and the intermediate feature by using the adapter to obtain the compensation feature.

[0098] In an optional embodiment, the third teacher model training unit includes:

[0099] The first teacher model training sub-unit is used to splice the downsampled feature and the intermediate feature and input them into the adapter, and output the compensation feature after being processed by a convolutional neural network.

[0100] In an optional embodiment, the sample further includes a standard signal. For an incorrect sample, the teacher model training module 602 further includes:

[0101] The fifth teacher model training sub-module is used to apply a mask to the compensation feature and calculate the gap between the control signal and the standard control signal as the loss;

[0102] The sixth teacher model training sub-module is used to train the adapter using the backpropagation of the loss;

[0103] Wherein, the error sample is a sample in which the predicted control signal is inconsistent with the standard signal output;

[0104] The predicted control signal is output after the input bird's-eye view semantic segmentation and the bird's-eye view feature are processed sequentially through multiple layers of the teacher model.

[0105] In an alternative embodiment, the student model is constructed by sampling the BEVFusion method or the BEVFormer method.

[0106] In an alternative embodiment, the teacher model is constructed by sampling the Roach method or the LBC method.

[0107] The end-to-end autonomous driving model training device provided by the embodiments of the present application can implement Figures 1 to 5 each process implemented in the method embodiments, and for the sake of avoiding repetition, it will not be elaborated here.

[0108] The end-to-end autonomous driving model training device of the embodiments of the present application can execute the end-to-end autonomous driving model training method provided by the embodiments of the present application, and its implementation principle is similar. The actions performed by each module and unit in the end-to-end autonomous driving model training device in the embodiments of the present application correspond to the steps in the end-to-end autonomous driving model training method in the embodiments of the present application. For the detailed function descriptions of the modules of the end-to-end autonomous driving model training device, reference can specifically be made to the descriptions in the corresponding end-to-end autonomous driving model training method shown above, and it will not be elaborated here.

[0109] Based on the same principle as the method shown in the embodiment of the present application, the embodiment of the present application also provides an electronic device, which may include but is not limited to: a processor and a memory; a memory for storing a computer program; a processor for executing the end-to-end autonomous driving model training method shown in any optional embodiment of the present application by calling a computer program. Compared with the prior art, the end-to-end autonomous driving model training method provided by the present application inputs pre-collected samples into the student model, and outputs bird's-eye view semantic segmentation and bird's-eye view features; the samples include the original image and the corresponding point cloud data; then the bird's-eye view semantic segmentation and the bird's-eye view features are input into the teacher model, and the control signal is output to avoid the causal inversion problem caused by behavior cloning as much as possible, thereby improving the accuracy of the output results of the end-to-end autonomous driving model and improving the feasibility of end-to-end autonomous driving.

[0110] In an optional embodiment, an electronic device is also provided, such as Figure 7 As shown, Figure 7 The electronic device 700 shown may be a server, including: a processor 701 and a memory 703. The processor 701 and the memory 703 are connected, such as through a bus 702. Optionally, the electronic device 700 may further include a transceiver 704. It should be noted that in actual applications, the transceiver 704 is not limited to one, and the structure of the electronic device 700 does not constitute a limitation on the embodiments of the present application.

[0111] Processor 701 may be a CPU (Central Processing Unit), a general purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 701 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0112] The bus 702 may include a path for transmitting information among the above components. The bus 702 can be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 702 can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 7 only a thick line is used in Figure 7 , but it does not mean that there is only one bus or one type of bus.

[0113] The memory 703 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it can also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0114] The memory 703 is used to store the application program code for implementing the solution of this application, and is controlled by the processor 701 to execute. The processor 701 is used to execute the application program code stored in the memory 703 to implement the content shown in the foregoing method embodiments.

[0115] Among them, the electronic device includes but is not limited to: mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The shown electronic device is only an example and should not impose any limitations on the functions and usage scope of the embodiments of this application.

[0116] The server provided in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected by wired or wireless communication, and this application does not limit this.

[0117] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content in the aforementioned method embodiment.

[0118] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0119] It should be noted that the above-mentioned computer-readable storage medium in the present application may also be a computer-readable signal medium or a combination of a computer-readable storage medium and a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. And in the present application, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0120] The above-mentioned computer-readable medium may be included in the above-mentioned electronic device; or it may exist separately without being assembled into the electronic device.

[0121] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.

[0122] According to one aspect of the present application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the end-to-end autonomous driving model training method and apparatus provided in the above various optional implementation manners.

[0123] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the module itself in some cases. For example, the student model training module can also be described as "the student model training module for inputting pre-collected samples into the student model and outputting the bird's-eye view semantic segmentation and bird's-eye view features".

[0126] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) disclosed in the present application that have similar functions.

Claims

1. An end-to-end autonomous driving model training method, characterized in that, the method includes: Inputting the pre-collected samples into the student model, and outputting bird's-eye view semantic segmentation and bird's-eye view features; the samples include original images and corresponding point cloud data; Inputting the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model, and outputting to obtain a control signal; The teacher model includes multiple layers; the inputting the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and outputting to obtain a control signal includes: Inputting the bird's-eye view semantic segmentation into the teacher model; Using a preset adapter to calculate the difference between the intermediate features output by each layer of the teacher model and the bird's-eye view features to obtain a compensation feature; Taking the compensation feature as the input of the next layer of the teacher model; Taking the output of the last layer of the teacher model as the control signal; The using a preset adapter to calculate the difference between the intermediate features output by each layer of the teacher model and the bird's-eye view features to obtain a compensation feature includes: Obtaining the intermediate features output by the corresponding layer in the teacher model; Performing downsampling processing on the bird's-eye view features to obtain downsampled features with the same size as the intermediate features; Using the adapter to process the downsampled features and the intermediate features to obtain the compensation feature.

2. The end-to-end autonomous driving model training method according to claim 1, characterized in that, the using the adapter to process the downsampled features and the intermediate features to obtain the compensation feature includes: Inputting the concatenated downsampled features and the intermediate features into the adapter, and outputting to obtain the compensation feature after being processed by a convolutional neural network.

3. The end-to-end autonomous driving model training method according to claim 1, characterized in that, the samples further include standard signals, and for incorrect samples, the method further includes: Applying a mask to the compensation feature, and calculating the gap between the control signal and the standard signal as a loss; Training the adapter using the backpropagation of the loss; wherein, the incorrect samples are samples where the predicted control signal is inconsistent with the output of the standard signal; The predicted control signal is output after the bird's-eye view semantic segmentation and the bird's-eye view features input are sequentially processed by multiple layers of the teacher model.

4. The end-to-end autonomous driving model training method according to claim 1, characterized in that, the student model is constructed using the BEVFusion method or the BEVFormer method.

5. An end-to-end autonomous driving model training device, characterized in that, the device includes: A student model training module, configured to input pre-collected samples into the student model and output bird's-eye view semantic segmentation and bird's-eye view features; the samples include original images and corresponding point cloud data; A teacher model training module, configured to input the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and output to obtain a control signal; The teacher model includes multiple layers; the inputting the bird's-eye view semantic segmentation and the bird's-eye view features into the teacher model and outputting to obtain a control signal includes: Input the bird's-eye view semantic segmentation into the teacher model; Use a preset adapter to calculate the difference between the intermediate features output by each layer in the teacher model and the bird's-eye view features, and obtain compensation features; Use the compensation features as the input for the next layer of the teacher model; Use the output of the last layer of the teacher model as the control signal; The step of using a preset adapter to calculate the difference between the intermediate features output by each layer in the teacher model and the bird's-eye view features, and obtaining compensation features includes: Obtain the intermediate features output by the corresponding layer in the teacher model; Perform downsampling on the bird's-eye view features to obtain downsampled features with the same size as the intermediate features; Use the adapter to process the downsampled features and the intermediate features to obtain the compensation features.

6. An electronic device, characterized in that, it includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the method according to any one of claims 1 to 4.

7. A computer-readable storage medium, characterized in that, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, it implements the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Model training method and device, object recognition method and device, vehicle and storage medium

    CN114973178A

  • Perception result acquisition method and device, computer equipment and storage medium

    CN115240168A