Scene coding model training method, device and medium for autonomous driving

By applying a masking strategy to the sample data of the autonomous driving scenario coding model, the model robustness is improved, the problem of inaccurate coding under abnormal noise data is solved, and more accurate path planning and decision-making are achieved.

CN116469069BActive Publication Date: 2025-09-09BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310316906.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-09-09
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

When existing autonomous driving technology processes abnormal noise data, the scene coding model is not robust enough, resulting in inaccurate coding and affecting the accuracy of path planning and decision-making.

Method used

By applying masking strategies to the sample dataset, especially randomly masking the vehicle driving status, obstacle status, and road network data, the robustness of the model is improved. This includes using MultiPath++ networks, Vectornet networks, Wayformer networks, etc. to build scene encoding models, and training the models through comparative learning.

Benefits of technology

The coding accuracy of the scene coding model under abnormal noise data is improved, the accuracy of path planning and decision-making is ensured, and the stability of the autonomous driving system is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469069B_ABST
    Figure CN116469069B_ABST
Patent Text Reader

Abstract

This disclosure provides a method, device, and medium for training a scene coding model for autonomous driving, relating to the fields of computer technology, particularly artificial intelligence and autonomous driving. The method comprises: obtaining a sample dataset; processing multiple first sample data based on at least one masking strategy to obtain multiple second sample data; inputting the multiple second sample data into a scene coding model to obtain, as output by the scene coding model, multiple driving scene codes corresponding to the multiple second sample data; and training the scene coding model based on the multiple driving scene codes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, in particular to the fields of artificial intelligence and autonomous driving, and specifically to a scene coding model training method for autonomous driving, a scene coding method, device, electronic device, computer-readable storage medium, and computer program product for autonomous driving. Background Art

[0002] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0003] AI-based data processing has been widely applied in various fields. In the field of autonomous driving, AI-based data processing can plan optimal driving trajectories for vehicles. Autonomous vehicles rely on the collaborative efforts of AI, visual computing, radar, monitoring devices, and global positioning systems, enabling them to be driven without the need for on-site human control. This represents a key development direction for future intelligent transportation.

[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides a scene coding model training method for autonomous driving, a scene coding method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for autonomous driving.

[0006] According to one aspect of the present disclosure, a method for training a scene coding model for autonomous driving is provided, including: obtaining a sample data set, the sample data set including a plurality of first sample data corresponding to a plurality of first moments, each first sample data including first data corresponding to the first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data including vehicle driving status data at the corresponding first moment, and each second data in the at least one second data including vehicle driving status data at the corresponding historical moment; processing the plurality of first sample data based on at least one mask strategy to obtain a plurality of second sample data, the at least one mask strategy including a first mask strategy, the first mask strategy being used to mask the first data of each third sample data in at least one third sample data in the plurality of first sample data; inputting the plurality of second sample data into a scene coding model to obtain a plurality of driving scene codes output by the scene coding model corresponding to the plurality of second sample data respectively; and training the scene coding model based on the plurality of driving scene codes.

[0007] According to another aspect of the present disclosure, a scene coding method for autonomous driving is provided, comprising: obtaining first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, the first data comprising at least one of vehicle driving status data, obstacle status data, and road network data at the current moment, and each of the at least one second data comprising at least one of vehicle driving status data, obstacle status data, and road network data at the corresponding historical moment; inputting the first data and the at least one second data into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the above-mentioned scene coding model training method for autonomous driving.

[0008] According to another aspect of the present disclosure, a scene coding model training device for autonomous driving is provided, comprising: a first acquisition unit, configured to acquire a sample data set, the sample data set including a plurality of first sample data corresponding to a plurality of first moments, each first sample data including first data corresponding to the first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data including vehicle driving status data at the corresponding first moment, and each second data in the at least one second data including vehicle driving status data at the corresponding historical moment; a processing unit, configured to process the plurality of first sample data based on at least one mask strategy to obtain a plurality of second sample data, the at least one mask strategy including a first mask strategy, the first mask strategy being used to mask the first data of each third sample data in at least one third sample data in the plurality of first sample data; an input unit, configured to input the plurality of second sample data into the scene coding model to obtain a plurality of driving scene codes output by the scene coding model corresponding to the plurality of second sample data respectively; and a training unit, configured to train the scene coding model based on the plurality of driving scene codes.

[0009] According to another aspect of the present disclosure, a scene coding device for autonomous driving is provided, including: a second acquisition unit, configured to acquire first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, the first data including at least one of vehicle driving status data, obstacle status data and road network data at the current moment, and each of the at least one second data including at least one of vehicle driving status data, obstacle status data and road network data at the corresponding historical moment; an encoding unit, configured to input the first data and at least one second data into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the above-mentioned scene coding model training method for autonomous driving.

[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned scene coding model training method for autonomous driving or the above-mentioned scene coding method for autonomous driving.

[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned scene coding model training method for autonomous driving or the above-mentioned scene coding method for autonomous driving.

[0012] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned scene coding model training method for autonomous driving or the above-mentioned scene coding method for autonomous driving.

[0013] According to one or more embodiments of the present disclosure, at least one masking strategy can be used to randomly mask part of the information in one or more sample data among multiple first sample data (for example, randomly mask the first data in the sample data), and perform model training based on the processed sample data, thereby improving the robustness of the model so that it can have more accurate scene coding expression when receiving abnormal noise data of the first data.

[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0016] Figure 1 A schematic diagram illustrating an exemplary system in which the various methods described herein may be implemented according to an embodiment of the present disclosure;

[0017] Figure 2 A flowchart of a scene coding model training method for autonomous driving according to an embodiment of the present disclosure is shown;

[0018] Figure 3 A flowchart of a scene coding method for autonomous driving according to an embodiment of the present disclosure is shown;

[0019] Figure 4 A structural block diagram of a scene coding model training device for autonomous driving according to an embodiment of the present disclosure is shown;

[0020] Figure 5 A structural block diagram of a scene coding apparatus for autonomous driving according to an embodiment of the present disclosure is shown;

[0021] Figure 6 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.

[0024] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.

[0025] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0026] Figure 1 FIG2 is a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1 , the system 100 includes a motor vehicle 110 , a server 120 , and one or more communication networks 130 coupling the motor vehicle 110 to the server 120 .

[0027] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to perform a method according to an embodiment of the present disclosure.

[0028] The server 120 may run one or more services or software applications that enable methods for training scene encoding models for autonomous driving. In some embodiments, the server 120 may also provide other services or software applications that may include non-virtual environments and virtual environments. Figure 1In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of the motor vehicle 110 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.

[0029] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0030] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0031] In some embodiments, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from motor vehicle 110. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.

[0032] The network 130 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a satellite communication network, a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including, for example, Bluetooth, WiFi), and / or any combination of these and other networks.

[0033] The system 100 may also include one or more databases 150. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repository 150 can reside in a variety of locations. For example, the data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repository 150 can be of different types. In some embodiments, the data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0034] In some embodiments, one or more of the databases 150 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0035] Motor vehicle 110 may include sensors 111 for sensing its surroundings. Sensors 111 may include one or more of the following: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radar, and laser radar (LiDAR). Different sensors offer different detection accuracy and range. Cameras may be mounted on the front, rear, or other locations of the vehicle. Visual cameras can capture real-time information about the vehicle's interior and exterior and present it to the driver and / or passengers. Furthermore, by analyzing the images captured by the visual cameras, information such as traffic light indications, intersection conditions, and the operating status of other vehicles can be obtained. Infrared cameras can detect objects in night vision conditions. Ultrasonic sensors can be mounted on all sides of the vehicle, utilizing the strong directionality of ultrasonic waves to measure the distance of external objects from the vehicle. Millimeter-wave radars can be mounted on the front, rear, or other locations of the vehicle, utilizing the properties of electromagnetic waves to measure the distance of external objects from the vehicle. LiDARs can be mounted on the front, rear, or other locations of the vehicle, detecting object edges and shapes for object recognition and tracking. Due to the Doppler effect, radar devices can also measure changes in the speed of the vehicle and moving objects.

[0036] The motor vehicle 110 may also include a communication device 112. The communication device 112 may include a satellite positioning module that can receive satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generate coordinates based on these signals. The communication device 112 may also include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as GSM / GPRS, CDMA, LTE, and other current or evolving wireless communication technologies (e.g., 5G technology). The communication device 112 may also have a vehicle-to-everything (V2X) module that is configured to implement vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may also include a module configured to communicate with a user terminal 145 (including but not limited to a smartphone, tablet computer, or wearable device such as a watch) via a wireless local area network or Bluetooth using the IEEE 802.11 standard, for example. Using the communication device 112, the motor vehicle 110 may also access the server 120 via the network 130.

[0037] The motor vehicle 110 may also include a control device 113. The control device 113 may include a processor that communicates with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, and braking system of the motor vehicle 110 (not shown) via multiple actuators in response to input from multiple sensors 111 or other input devices to control acceleration, steering, and braking, respectively, without human intervention or limited human intervention. Some processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an on-board processor, while other processing may be performed using computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.

[0038] Figure 1 The system 100 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with the present disclosure.

[0039] According to the embodiments of the present disclosure, Figure 2As shown, a scene coding model training method for autonomous driving is provided, including:

[0040] Step S201: Acquire a sample data set, where the sample data set includes a plurality of first sample data corresponding to a plurality of first moments, each first sample data including first data corresponding to the first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data including vehicle driving state data at the corresponding first moment, and each of the at least one second data including vehicle driving state data at the corresponding historical moment;

[0041] Step S202: Process the plurality of first sample data based on at least one masking strategy to obtain a plurality of second sample data, where the at least one masking strategy includes a first masking strategy, and the first masking strategy is used to perform masking processing on the first data of each third sample data in at least one third sample data in the plurality of first sample data;

[0042] Step S203: input the plurality of second sample data into the scene coding model to obtain a plurality of driving scene codes output by the scene coding model and corresponding to the plurality of second sample data respectively; and

[0043] Step S204: training a scene coding model based on multiple driving scene codes.

[0044] Therefore, through at least one masking strategy, part of the information in one or more sample data among multiple first sample data is randomly masked (for example, the first data in the sample data is randomly masked), and the model is trained based on the processed sample data, thereby improving the robustness of the model so that it can have more accurate scene coding expression when receiving abnormal noise data of the first data.

[0045] In some embodiments, the raw data of the sample data set may be derived from data collected by vehicles over a period of history, wherein the historical data may include data collected by one or more vehicles over a period of history (e.g., several months). For each vehicle's historical data, it may be vehicle driving status data at different times in the vehicle coordinate system collected at a certain collection frequency.

[0046] In some embodiments, the historical data for each vehicle at different times can also include obstacle status data and road network data within a certain area around the vehicle in the vehicle coordinate system. Therefore, by introducing obstacle status data and road network data, the accuracy of the model encoding representation can be further improved.

[0047] In some embodiments, the vehicle driving status data may include data such as the vehicle position, vehicle driving direction, vehicle length, vehicle width, vehicle type, driving speed, acceleration, turn signal, etc. at the corresponding moment.

[0048] In some embodiments, obstacle status data may include data such as the speed, acceleration, and distance between obstacles and the vehicle within a certain area around the vehicle at a given moment. Obstacles may include, for example, other motor vehicles, non-motor vehicles, and pedestrians. Road network data may include road network topology data within a certain area around the vehicle's location at a given moment, including, for example, lane markings, stop signs, and crosswalks.

[0049] In some embodiments, the raw data is collated to obtain a sample data set, wherein each first sample data includes first data corresponding to a first moment and at least one second data corresponding to at least one historical moment before the first moment. The first data may be vehicle driving status data of a vehicle at the first moment, and the second data may be vehicle driving status data of a corresponding vehicle at the corresponding historical moment.

[0050] In some exemplary embodiments, a first sample data may include vehicle driving status data of the corresponding vehicle at the first moment corresponding to the sample data, and 16 frames of vehicle driving status data corresponding to the most recent 16 historical moments before the first moment of the vehicle.

[0051] In some embodiments, the scene coding model of the present disclosure can be constructed based on at least one of a MultiPath++ network, a Vectornet network, and a Wayformer network.

[0052] In some embodiments, the above-mentioned scene coding model training method for autonomous driving may further include: dividing multiple first sample data into at least one sample pair, wherein the similarity of the vehicle driving scenes corresponding to the two sample data in each sample pair in at least one sample pair meets and exceeds a preset value; and, training the scene coding model based on multiple driving scene codings may include: training the scene coding model based on the distance between two scene codings in at least one scene coding pair in multiple driving scene codings, wherein at least one scene coding pair corresponds to at least one sample pair.

[0053] In some embodiments, multiple first sample data can be first divided into at least one sample pair according to the degree of scene similarity, wherein the similarity of the vehicle driving scenes corresponding to the two sample data in each sample pair meets a preset condition (for example, greater than a preset value).

[0054] In some embodiments, a certain number of negative sample pairs may also be prepared at the same time, that is, two sample data corresponding to dissimilar scenes are combined into a negative sample pair for model training.

[0055] In some embodiments, each sample data in each sample pair can be input into the scene coding model in turn to obtain the driving scene coding corresponding to each sample data, and the contrast learning training method can be applied to calculate the contrast loss, and the contrast loss can be applied to perform model training.

[0056] Therefore, by dividing the sample data into at least one sample pair according to scene similarity and performing comparative learning based on the sample pairs, the accuracy of the model encoding expression can be further improved.

[0057] In some embodiments, before inputting the plurality of first sample data into the model, they may first be masked based on at least one masking strategy to obtain a plurality of second sample data.

[0058] In some embodiments, the masking strategy may be a first masking strategy for performing masking processing on the first data of each of at least one third sample data in the plurality of first sample data.

[0059] In some embodiments, for each first sample data among the above-mentioned multiple first sample data, random mask processing or discard processing can be performed on the first data based on a first preset probability (for example, 50%), thereby achieving the above-mentioned mask processing.

[0060] In some embodiments, at least one mask strategy is obtained from a preset mask strategy set, and the preset mask strategy set also includes a second mask strategy, which is used to perform random masking processing on at least one second data of each fourth sample data in at least one fourth sample data among multiple first sample data.

[0061] Therefore, by randomly masking at least one second data in the sample data, the robustness of the model is improved, so that it can have more accurate scene coding expression when at least one second data has anomalies.

[0062] In some embodiments, a preset mask strategy set may be predetermined, including multiple preset mask strategies (eg, the first mask strategy and the second mask strategy described above), and at least one mask strategy for processing the first sample data may be obtained from the preset mask strategy set.

[0063] In some embodiments, the preset mask strategy set may include a second mask strategy, and the second mask strategy may be used to perform random masking processing on at least one second data of each of at least one fourth sample data in the plurality of first sample data.

[0064] In some embodiments, for each second data of at least one second data of each first sample data in the above-mentioned multiple first sample data, the second data can be randomly masked or discarded based on a second preset probability (for example, 50%), thereby realizing the above-mentioned masking processing.

[0065] In some embodiments, the number of at least one second data is multiple, and the preset mask strategy set also includes a third mask strategy, which is used to mask the third data of each fifth sample data in at least one fifth sample data in the multiple first sample data, wherein the third data is the second data corresponding to the earliest historical moment in at least one historical moment corresponding to the fifth sample data.

[0066] Therefore, by performing random masking on the second data of the earliest historical moment in the sample data, the model can further balance its dependence on the second data of the earliest historical moment and the first data, thereby further improving the robustness of the model.

[0067] In some embodiments, the preset mask strategy set may further include the third mask strategy described above.

[0068] In some embodiments, the masking strategy may be a third masking strategy for performing masking on the third data of each of at least one fifth sample data in the plurality of first sample data. The third data may be the second data corresponding to the earliest historical moment among the plurality of second data in the first sample data.

[0069] In some embodiments, for each first sample data among the above-mentioned multiple first sample data, its third data can be randomly masked or discarded based on a third preset probability (for example, 50%), thereby achieving the above-mentioned masking process.

[0070] In some embodiments, the first data includes obstacle status data and road network data corresponding to the corresponding first moment, and each second data in the at least one second data includes obstacle status data and road network data at the corresponding historical moment.

[0071] Therefore, by further introducing obstacle status data and road network data, the accuracy of the model encoding representation can be further improved; at the same time, by masking the obstacle status data and road network data accordingly, the robustness of the model can be further improved.

[0072] In some embodiments, each of the first data and the second data may further include obstacle status data and road network data at corresponding moments, respectively.

[0073] In some embodiments, when performing mask processing, the vehicle driving status data, obstacle status data, and road network data in the first data and / or the second data may be masked simultaneously based on corresponding strategies.

[0074] In some embodiments, when performing mask processing, based on corresponding strategies, only one or more of the vehicle driving status data, obstacle status data, and road network data in the first data and / or the second data may be masked.

[0075] In some embodiments, when performing mask processing, mask processing may be performed only on the vehicle driving status data in the first data and / or the second data based on corresponding strategies.

[0076] Therefore, we can further improve the robustness of the model through multi-dimensional masking strategies and further enhance the model's ability to handle abnormal data.

[0077] In some embodiments, the above-mentioned scene coding model training method for autonomous driving may further include: dividing multiple first sample data into multiple sample batches; and based on at least one mask strategy, processing the multiple first sample data to obtain multiple second sample data may include: for each sample batch in the multiple sample batches, randomly selecting at least one mask strategy from a preset mask strategy set based on a preset probability, processing the sample data in the sample batch to obtain sample data of the processed multiple sample batches as multiple second sample data.

[0078] In some embodiments, the plurality of first sample data may be divided into a plurality of sample batches, and for each sample batch, at least one masking strategy may be randomly selected from a preset masking strategy set based on a preset probability to process the sample data in the sample batch.

[0079] In some embodiments, when the preset mask strategy set includes the above three mask strategies, for each sample batch, one of the above three mask strategies can be randomly selected based on equal probability (for example, 1 / 3), and the sample data in the sample batch can be processed based on the selected mask strategy.

[0080] Therefore, the sample data is divided into multiple batches, and different masking strategies are randomly selected for each batch to be processed for model training, so that the model can comprehensively improve the robustness of the model through multi-dimensional masking strategies and further enhance the model's ability to handle abnormal data.

[0081] In some embodiments, as Figure 3 As shown, a scene coding method for autonomous driving is provided, including:

[0082] Step S301: Acquire first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, where the first data includes at least one of vehicle driving state data, obstacle state data, and road network data at the current moment, and each of the at least one second data includes at least one of vehicle driving state data, obstacle state data, and road network data at a corresponding historical moment;

[0083] Step S302: input the first data and at least one second data into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the above-mentioned scene coding model training method for autonomous driving.

[0084] In some embodiments, the vehicle driving status data, obstacle status data and road network data at the current moment and the vehicle driving status data, obstacle status data and road network data corresponding to each historical moment in at least one historical moment before the current moment (for example, 16 historical moments before the current moment) can be simultaneously input into the scene coding model to obtain the driving scene coding of the driving scene in which the vehicle is located at the current moment.

[0085] Therefore, by applying the above-mentioned model training method, based on at least one masking strategy, partial information in one or more sample data among multiple first sample data is randomly masked (for example, the first data in the sample data is randomly masked), and model training is performed based on the processed sample data, thereby improving the robustness of the model so that it can have a more accurate scene coding expression when receiving abnormal noise data of the first data; furthermore, applying the scene coding expression to subsequent path planning and decision-making operations can achieve more accurate planning and decision-making.

[0086] In some embodiments, as Figure 4 As shown, a scene coding model training device 400 for autonomous driving is provided, comprising:

[0087] A first acquisition unit 410 is configured to acquire a sample data set, the sample data set including a plurality of first sample data corresponding to a plurality of first moments, each first sample data including first data corresponding to the first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data including vehicle driving state data at the corresponding first moment, and each of the at least one second data including vehicle driving state data at the corresponding historical moment;

[0088] a processing unit 420 configured to process the plurality of first sample data based on at least one masking strategy to obtain a plurality of second sample data, the at least one masking strategy including a first masking strategy for performing masking processing on first data of each third sample data in at least one third sample data in the plurality of first sample data;

[0089] The input unit 430 is configured to input the plurality of second sample data into the scene coding model to obtain a plurality of driving scene codes output by the scene coding model and corresponding to the plurality of second sample data respectively; and

[0090] The training unit 440 is configured to train a scene coding model based on a plurality of driving scene codes.

[0091] Among them, the operations of units 410-440 in the scene coding model training device 400 for autonomous driving are similar to the operations of steps S201-S204 in the above-mentioned scene coding model training method for autonomous driving, and will not be repeated here.

[0092] In some embodiments, at least one mask strategy is obtained from a preset mask strategy set, and the preset mask strategy set also includes a second mask strategy, which is used to perform random masking processing on at least one second data of each fourth sample data in at least one fourth sample data among multiple first sample data.

[0093] In some embodiments, the number of at least one second data is multiple, and the preset mask strategy set also includes a third mask strategy, which is used to mask the third data of each fifth sample data in at least one fifth sample data in the multiple first sample data, wherein the third data is the second data corresponding to the earliest historical moment in at least one historical moment corresponding to the fifth sample data.

[0094] In some embodiments, the above-mentioned scene coding model training device for autonomous driving may further include: a first division unit, configured to divide multiple first sample data into multiple sample batches; and the first processing unit may be further configured to: for each sample batch in the multiple sample batches, randomly select at least one mask strategy from a preset mask strategy set based on a preset probability, and process the sample data in the sample batch to obtain sample data of the processed multiple sample batches as multiple second sample data.

[0095] In some embodiments, the above-mentioned scene coding model training device for autonomous driving may also include: a second division unit, configured to divide multiple first sample data into at least one sample pair, and the similarity of the vehicle driving scenes corresponding to the two sample data in each sample pair in at least one sample pair is greater than a preset value; and the training unit can be further configured to: train the scene coding model based on the distance between the two scene codes in at least one scene coding pair in multiple driving scene codes, wherein at least one scene coding pair corresponds to at least one sample pair.

[0096] In some embodiments, the first data includes obstacle status data and road network data corresponding to the corresponding first moment, and each second data in the at least one second data includes obstacle status data and road network data at the corresponding historical moment.

[0097] In some embodiments, as Figure 5 As shown, a scene coding device 500 for autonomous driving is provided, comprising:

[0098] A second acquiring unit 510 is configured to acquire first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, wherein the first data includes at least one of vehicle driving state data, obstacle state data, and road network data at the current moment, and each of the at least one second data includes at least one of vehicle driving state data, obstacle state data, and road network data at a corresponding historical moment;

[0099] The encoding unit 520 is configured to input the first data and at least one second data into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the above-mentioned scene coding model training method for autonomous driving.

[0100] Among them, the operations of unit 510 and unit 520 in the scene encoding device 500 for autonomous driving are similar to the operations of step S301 and step S302 in the above-mentioned scene method for autonomous driving, and will not be repeated here.

[0101] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0102] refer to Figure 6, a block diagram of an electronic device 600 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0103] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0104] Multiple components within electronic device 600 are connected to I / O interface 605, including an input unit 606, an output unit 607, a storage unit 608, and a communication unit 609. Input unit 606 can be any type of device capable of inputting information into electronic device 600. Input unit 606 can receive input numeric or character information and generate key signal input related to user settings and / or function control of the electronic device. It can include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 607 can be any type of device capable of presenting information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 608 can include, but is not limited to, a magnetic disk or an optical disk. Communication unit 609 allows electronic device 600 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. It can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0105] The computing unit 601 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the scene coding model training method for autonomous driving or the scene coding method for autonomous driving. For example, in some embodiments, the scene coding model training method for autonomous driving or the scene coding method for autonomous driving can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the scene coding model training method for autonomous driving or the scene coding method for autonomous driving can be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the above-mentioned scene coding model training method for autonomous driving or the above-mentioned scene coding method for autonomous driving in any other appropriate manner (for example, by means of firmware).

[0106] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0107] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0108] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0109] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0110] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0111] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0112] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0113] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.

Claims

1. A method for training a scene coding model for autonomous driving, the method comprising: Acquire a sample data set, the sample data set including a plurality of first sample data corresponding to a plurality of first moments, each first sample data including first data corresponding to a first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data including vehicle driving state data at the corresponding first moment, and each second data of the at least one second data including vehicle driving state data at the corresponding historical moment; Processing the plurality of first sample data based on at least one masking strategy to obtain a plurality of second sample data, the at least one masking strategy including a first masking strategy for performing masking processing on first data of each third sample data in at least one third sample data in the plurality of first sample data; Inputting the plurality of second sample data into the scene coding model to obtain a plurality of driving scene codes output by the scene coding model and corresponding to the plurality of second sample data respectively; as well as The scene coding model is trained based on the multiple driving scene codings.

2. The method according to claim 1, wherein The at least one masking strategy is obtained from a preset masking strategy set, and the preset masking strategy set also includes a second masking strategy, and the second masking strategy is used to perform random masking processing on at least one second data of each fourth sample data in at least one fourth sample data among the multiple first sample data.

3. The method according to claim 2, wherein: The number of the at least one second data is multiple, and the preset mask strategy set also includes a third mask strategy, and the third mask strategy is used to mask the third data of each fifth sample data in at least one fifth sample data among the multiple first sample data, wherein the third data is the second data corresponding to the earliest historical moment in at least one historical moment corresponding to the fifth sample data.

4. The method according to claim 3, further comprising: Dividing the plurality of first sample data into a plurality of sample batches; and The processing of the plurality of first sample data based on at least one mask strategy to obtain a plurality of second sample data includes: For each sample batch in the multiple sample batches, at least one masking strategy is randomly selected from the preset masking strategy set based on a preset probability, and sample data in the sample batch is processed to obtain sample data of multiple processed sample batches as the multiple second sample data.

5. The method according to any one of claims 1 to 4, further comprising: Divide the plurality of first sample data into at least one sample pair, wherein the similarity of the vehicle driving scenes corresponding to the two sample data in each sample pair in the at least one sample pair is greater than a preset value; and The training of the scene coding model based on the multiple driving scene codes includes: The scene coding model is trained based on a distance between two scene codes in at least one scene coding pair among the plurality of driving scene codes, wherein the at least one scene coding pair corresponds to the at least one sample pair.

6. The method according to any one of claims 1 to 4, wherein The first data includes obstacle status data and road network data corresponding to a corresponding first moment, and each second data in the at least one second data includes obstacle status data and road network data at a corresponding historical moment.

7. A scene coding method for autonomous driving, the method comprising: Acquire first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, wherein the first data includes at least one of vehicle driving state data, obstacle state data, and road network data at the current moment, and each second data of the at least one second data includes at least one of vehicle driving state data, obstacle state data, and road network data at a corresponding historical moment; The first data and the at least one second data are input into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the method described in any one of claims 1-6.

8. A scene coding model training device for autonomous driving, the device comprising: a first acquisition unit configured to acquire a sample data set, the sample data set comprising a plurality of first sample data corresponding to a plurality of first moments, each first sample data comprising first data corresponding to the first moment corresponding to the first sample data and at least one second data corresponding to at least one historical moment before the corresponding first moment, the first data comprising vehicle driving state data at the corresponding first moment, and each of the at least one second data comprising vehicle driving state data at the corresponding historical moment; a processing unit configured to process the plurality of first sample data based on at least one masking strategy to obtain a plurality of second sample data, the at least one masking strategy including a first masking strategy for performing masking processing on first data of each third sample data in at least one third sample data in the plurality of first sample data; an input unit configured to input the plurality of second sample data into the scene coding model to obtain a plurality of driving scene codes output by the scene coding model and corresponding to the plurality of second sample data respectively; as well as A training unit is configured to train the scene coding model based on the multiple driving scene codings.

9. The device according to claim 8, wherein The at least one masking strategy is obtained from a preset masking strategy set, and the preset masking strategy set also includes a second masking strategy, and the second masking strategy is used to perform random masking processing on at least one second data of each fourth sample data in at least one fourth sample data among the multiple first sample data.

10. The device according to claim 9, wherein The number of the at least one second data is multiple, and the preset mask strategy set also includes a third mask strategy, and the third mask strategy is used to mask the third data of each fifth sample data in at least one fifth sample data among the multiple first sample data, wherein the third data is the second data corresponding to the earliest historical moment in at least one historical moment corresponding to the fifth sample data.

11. The apparatus according to claim 10, further comprising: A first dividing unit is configured to divide the plurality of first sample data into a plurality of sample batches; and The first processing unit is further configured to: For each sample batch in the multiple sample batches, at least one masking strategy is randomly selected from the preset masking strategy set based on a preset probability, and sample data in the sample batch is processed to obtain sample data of multiple processed sample batches as the multiple second sample data.

12. The apparatus according to any one of claims 8 to 11, further comprising: The second dividing unit is configured to divide the plurality of first sample data into at least one sample pair, wherein the similarity of the vehicle driving scenes corresponding to the two sample data in each sample pair of the at least one sample pair is greater than a preset value; and The training unit is further configured to: The scene coding model is trained based on a distance between two scene codes in at least one scene coding pair among the plurality of driving scene codes, wherein the at least one scene coding pair corresponds to the at least one sample pair.

13. The device according to any one of claims 8 to 11, wherein The first data includes obstacle status data and road network data corresponding to a corresponding first moment, and each second data in the at least one second data includes obstacle status data and road network data at a corresponding historical moment.

14. A scene coding device for autonomous driving, the device comprising: a second acquiring unit configured to acquire first data at a current moment and at least one second data corresponding to at least one historical moment before the current moment, wherein the first data includes at least one of vehicle driving state data, obstacle state data, and road network data at the current moment, and each second data of the at least one second data includes at least one of vehicle driving state data, obstacle state data, and road network data at a corresponding historical moment; The encoding unit is configured to input the first data and the at least one second data into a scene coding model to obtain the driving scene coding at the current moment output by the scene coding model, wherein the scene coding model is trained according to the method described in any one of claims 1-6.

15. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, wherein When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Automatic driving simulation scene recognition method and device

    CN111666714A

  • Reinforcement learning on autonomous vehicles

    US20190332110A1