Generative diffusion model-based autonomous driving model, method, apparatus, and vehicle
By introducing a generative diffusion model into the autonomous driving model, the problems of low prediction accuracy and decision-making efficiency in the prior art are solved, and more efficient and accurate autonomous driving decisions are achieved.
Patent Information
- Application Number
- PCT/CN2024/099434
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-06-14
- Publication Date
- 2025-06-12
AI Technical Summary
Existing autonomous driving technologies have problems with prediction accuracy and inefficiency in the perception and decision-making process, especially in complex and dynamic environments.
An autonomous driving model based on a generative diffusion model is adopted, which includes an encoding layer, a prediction layer and a decoding layer. The encoding layer encodes the current perceived information, the prediction layer generates a prediction spatial representation of the future moment through discrete diffusion, and the decoding layer decodes the prediction spatial representation into autonomous driving decision information.
By generating multiple future possibilities, the accuracy of future predictions is improved, thereby improving the effectiveness and efficiency of autonomous driving decisions.
Smart Images

Figure CN2024099434_12062025_PF_FP_ABST
Abstract
Description
Autonomous driving model, method, device and vehicle based on generative diffusion model
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311676294.5 filed on December 7, 2023, the entire contents of which are incorporated by reference in their entirety into this application. Technical Field
[0003] The present disclosure relates to the field of computer technology, in particular to the field of autonomous driving and artificial intelligence technology, and specifically to an autonomous driving model, a training method for an autonomous driving model, an autonomous driving method implemented using an autonomous driving model, an autonomous driving device based on the autonomous driving model, a training device for the autonomous driving model, an electronic device, a computer-readable storage medium, a computer program product, and an autonomous driving vehicle. Background Art
[0004] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0005] Autonomous driving technology integrates technologies such as recognition, decision-making, positioning, communication security, and human-computer interaction. Artificial intelligence learning can assist in generating autonomous driving strategies.
[0006] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.
[0007] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.
[0008] Summary of the Invention
[0009] The present disclosure provides an autonomous driving model, method, device, and vehicle based on a generative diffusion model.
[0010] According to one aspect of the present disclosure, an autonomous driving model is provided, comprising a coding layer, a prediction layer, and a decoding layer, wherein the coding layer is configured to encode current perception information of an autonomous driving vehicle to obtain a discrete spatial representation of a current scene; the prediction layer is configured to perform discrete diffusion based on at least one scene discrete spatial representation including a discrete spatial representation of the current scene to determine a predicted spatial representation at a future moment; and the decoding layer is configured to decode the predicted spatial representation to obtain autonomous driving decision information at the future moment.
[0011] According to another aspect of the present disclosure, a method for training an autonomous driving model is provided, comprising: obtaining current sample perception information of an autonomous driving vehicle and real driving decision information corresponding to the current sample perception information; encoding the current sample perception information using the encoding layer of the autonomous driving model to obtain a sample discrete space representation of the current scene; performing discrete diffusion using the prediction layer of the autonomous driving model based on at least one sample scene discrete space representation including the sample discrete space representation of the current scene to determine a sample prediction space representation at a future moment; decoding the sample prediction space representation using the decoding layer of the autonomous driving model to obtain sample driving decision information at the future moment, and adjusting parameters of the autonomous driving model based on the difference between the sample driving decision information and the real driving decision information.
[0012] According to another aspect of the present disclosure, an autonomous driving method implemented using an autonomous driving model is provided, including: using the encoding layer of the autonomous driving model to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene; using the prediction layer of the autonomous driving model to perform discrete diffusion based on at least one scene discrete spatial representation including the discrete spatial representation of the current scene to determine a predicted spatial representation at a future moment; and using the decoding layer of the autonomous driving model to decode the predicted spatial representation to obtain autonomous driving decision information at the future moment.
[0013] According to another aspect of the present disclosure, an autonomous driving device based on an autonomous driving model is provided, comprising: an encoding unit configured to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene; a prediction unit configured to perform discrete diffusion based on at least one scene discrete spatial representation including the discrete spatial representation of the current scene to determine a predicted spatial representation at a future moment; and a decoding unit configured to decode the predicted spatial representation to obtain autonomous driving decision information at the future moment.
[0014] According to another aspect of the present disclosure, there is provided an apparatus for training an autonomous driving model, comprising: an acquisition unit configured to acquire current sample perception information of an autonomous driving vehicle and real driving decision information corresponding to the current sample perception information; an encoding unit configured to encode the current sample perception information to obtain a sample discrete space representation of a current scene; a prediction unit configured to perform discrete diffusion based on at least one sample scene discrete space representation including the sample discrete space representation of the current scene to determine a sample prediction space representation at a future moment; a decoding unit configured to decode the sample prediction space representation to obtain the sample driving decision information at the future moment; and a parameter adjustment unit configured to adjust parameters of the autonomous driving model according to a difference between the sample driving decision information and the real driving decision information.
[0015] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform the above method.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.
[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.
[0018] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising: an autonomous driving device according to an embodiment of the present disclosure, or one of an electronic device.
[0019] Using the embodiments of the present disclosure, an autonomous driving model can utilize the output of a generative diffusion model to determine its autonomous driving decisions. Diffusion models can derive multiple future possibilities by diffusing information from the past to the future, improving the accuracy of future predictions and further enhancing the effectiveness of autonomous driving decisions and predictions.
[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.
[0022] FIG1 shows a schematic diagram of an exemplary system in which various methods described herein may be implemented according to an embodiment of the present disclosure;
[0023] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure;
[0024] FIG3 shows an exemplary flowchart of a method for training an autonomous driving model according to an embodiment of the present disclosure;
[0025] FIG4 shows an exemplary flowchart of an autonomous driving method according to an embodiment of the present disclosure;
[0026] FIG5 shows an exemplary process of an autonomous driving method according to an embodiment of the present disclosure;
[0027] FIG6 shows an exemplary diagram of a prediction layer according to an embodiment of the present disclosure;
[0028] FIG7 shows an exemplary diagram of decoding layers according to an embodiment of the present disclosure;
[0029] FIG8 shows an exemplary process of training a discretized vocabulary according to an embodiment of the present disclosure;
[0030] FIG9 shows a structural block diagram of an automatic driving device according to an embodiment of the present disclosure;
[0031] FIG10 shows a structural block diagram of an apparatus for training an autonomous driving model according to an embodiment of the present disclosure; and
[0032] FIG11 shows a block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0033] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0034] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.
[0035] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.
[0036] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0037] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0038] FIG1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein may be implemented according to an embodiment of the present disclosure. Referring to FIG1 , the system 100 includes a motor vehicle 110, a server 120, and one or more communication networks 130 coupling the motor vehicle 110 to the server 120.
[0039] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to perform a method according to an embodiment of the present disclosure.
[0040] The server 120 can run one or more services or software applications that enable autonomous driving. In some embodiments, the server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In the configuration shown in Figure 1, the server 120 can include one or more components that implement the functions performed by the server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of the motor vehicle 110 can, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which can be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0041] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0042] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.
[0043] In some embodiments, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from motor vehicle 110. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.
[0044] The network 130 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 130 may be a satellite communication network, a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including, for example, Bluetooth, WiFi), and / or any combination of these and other networks.
[0045] The system 100 may also include one or more databases 150. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repository 150 can reside in a variety of locations. For example, the data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repository 150 can be of different types. In some embodiments, the data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.
[0046] In some embodiments, one or more of the databases 150 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.
[0047] Motor vehicle 110 may include sensors 111 for sensing its surroundings. Sensors 111 may include one or more of the following: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radar, and laser radar (LiDAR). Different sensors offer different detection accuracy and range. Cameras may be mounted on the front, rear, or other locations of the vehicle. Visual cameras can capture real-time information about the vehicle's interior and exterior and present it to the driver and / or passengers. Furthermore, by analyzing the images captured by the visual cameras, information such as traffic light indications, intersection conditions, and the operating status of other vehicles can be obtained. Infrared cameras can detect objects in night vision conditions. Ultrasonic sensors can be mounted on all sides of the vehicle, utilizing the strong directionality of ultrasonic waves to measure the distance of external objects from the vehicle. Millimeter-wave radars can be mounted on the front, rear, or other locations of the vehicle, utilizing the properties of electromagnetic waves to measure the distance of external objects from the vehicle. LiDARs can be mounted on the front, rear, or other locations of the vehicle, detecting object edges and shapes for object recognition and tracking. Due to the Doppler effect, radar devices can also measure changes in the speed of the vehicle and moving objects.
[0048] The motor vehicle 110 may also include a communication device 112. The communication device 112 may include a satellite positioning module that can receive satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generate coordinates based on these signals. The communication device 112 may also include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as GSM / GPRS, CDMA, LTE, and other current or evolving wireless communication technologies (e.g., 5G technology). The communication device 112 may also have a vehicle-to-everything (V2X) module that is configured to implement vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may also include a module configured to communicate with a user terminal 145 (including but not limited to a smartphone, tablet computer, or wearable device such as a watch) via a wireless local area network or Bluetooth using the IEEE 802.11 standard, for example. Using the communication device 112, the motor vehicle 110 may also access the server 120 via the network 130.
[0049] The motor vehicle 110 may also include a control device 113. The control device 113 may include a processor that communicates with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, and braking system of the motor vehicle 110 (not shown) via multiple actuators in response to input from multiple sensors 111 or other input devices to control acceleration, steering, and braking, respectively, without human intervention or with limited human intervention. Some processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an on-board processor, while other processing may be performed using computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.
[0050] The system 100 of FIG. 1 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with this disclosure.
[0051] The diffusion model is a generative model that draws inspiration from the diffusion phenomenon in physics. In physics, the process of diffusion of matter from an area of high concentration to an area of low concentration is similar to the information loss caused by noise interference. Therefore, the diffusion model introduces noise and attempts to generate new information by removing it.
[0052] The present disclosure provides a new autonomous driving model based on a diffusion model.
[0053] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure.
[0054] As shown in FIG2 , the autonomous driving model 200 includes a coding layer 210 , a prediction layer 220 , and a decoding layer 230 .
[0055] The encoding layer 210 is configured to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene.
[0056] The prediction layer 220 is configured to perform discrete diffusion based on at least one scene discrete spatial representation including a discrete spatial representation of a current scene to determine a predicted spatial representation at a future time.
[0057] The decoding layer 230 is configured to decode the predicted spatial representation to obtain autonomous driving decision information at a future moment.
[0058] The autonomous driving model provided by the embodiments of the present disclosure can be used to determine the model's autonomous driving decisions using the output of a generative diffusion model. The diffusion model can derive multiple future possibilities by diffusing from the past to the future, improving the accuracy of future predictions and further enhancing the effectiveness of autonomous driving decisions and predictions.
[0059] The principles of the present disclosure will be described in detail below.
[0060] The encoding layer 210 is configured to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene.
[0061] The perception information may include sensor input collected by at least one sensor installed on the autonomous driving vehicle. The perception information for the vehicle's surrounding environment may include at least one of the following items: perception information from one or more cameras, perception information from one or more lidars, and perception information from one or more millimeter-wave radars. The current perception information may include sensor input collected at the current time t. The historical perception information may include historical perception information from all moments since the start of the autonomous driving process (for example, from time 0 to time t-1), and may also include historical perception information within a predetermined time period before the current time t (for example, from time tk to time t, where k represents a predetermined time range).
[0062] In some embodiments, the encoding layer 210 can be used to encode the perceptual information to obtain a discrete spatial representation of the scene corresponding to the perceptual information. In some examples, the discrete spatial representation of the scene referred to here refers to the discrete representation of the perceptual information in the BEV space. For example, the current perceptual information can be mapped to the bird's-eye view BEV space to obtain a continuous BEV representation of the current perceptual information. The current perceptual information can be mapped to the BEV space using models such as BEVFormer and BEVFusion to obtain a continuous BEV representation of the current perceptual information. Then, the continuous BEV representation can be discretized according to the pre-trained vocabulary to obtain a discrete spatial representation of the current scene. For example, the continuous BEV representation is b t For example, if the dimensions of the continuous BEV representation are W×H×C, the nearest vocabulary vector can be found in the secondary table for each of the W×H C-dimensional vectors in the continuous BEV representation, and the nearest vocabulary vector can be used to replace the corresponding vector in the BEV representation. In this way, the perceptual information in the BEV space can be represented using a limited number of vocabulary vectors. The vectors in the vocabulary can also be C-dimensional.
[0063] While the principles of this disclosure have been described above using the representation of perceptual information in BEV space, it is understood that, without departing from the principles of this disclosure, perceptual information can also be mapped to other spaces other than BEV space and a discrete spatial representation of the corresponding scene can be determined. Furthermore, any suitable method other than a vocabulary can be used to determine a discrete spatial representation of a scene.
[0064] The prediction layer 220 is configured to perform discrete diffusion based on at least one scene discrete spatial representation including a discrete spatial representation of a current scene to determine a predicted spatial representation at a future time.
[0065] The at least one scene discrete spatial representation may include a discrete spatial representation of the current scene and a discrete spatial representation of at least one historical scene. The discrete spatial representation of the current scene may be a discrete representation encoded based on the perceptual information at the current moment, and the discrete spatial representation of the historical scene may be a discrete representation encoded in the same manner based on the perceptual information at at least one previous moment, such as a discrete spatial representation in a BEV space.
[0066] In some embodiments, the predicted spatial representation of the future time instant may be a spatial representation of predicted perceptual information, that is, predicted information for the perceptual information at the future time instant.
[0067] For the spatial positions corresponding to the W×H vectors in the discrete spatial representation of each scene, the prediction layer 220 can be configured to perform feature transformation based on the information at each position in the discrete spatial representation of each scene and the information at each position in the predetermined information to obtain the information represented by the vector at the corresponding position in the predicted spatial representation. The size of the predetermined information can also be W×H. The predetermined information can have the same size as the predicted spatial representation and have predefined default content. For example, the predetermined information can include predetermined noise, or any other suitable information.
[0068] When determining the vector representation at each location in the prediction spatial representation, each scene discrete spatial representation can be spatially transformed to enable the network to perceive the spatial information in the scene. In some examples, the prediction layer can be configured to perform a spatial transformation on each scene discrete spatial representation to obtain a corresponding transformed spatial representation.
[0069] Various suitable neural network models can be used to perform spatial transformations on the discrete spatial representations of each scene. For example, the Swin Transformer or any other suitable spatial transformation model can be used to process the discrete spatial representations of each scene to achieve spatial transformation of information. The Swin Transformer can be used to spatially aggregate information within the current perception information, thereby enhancing the network's perception and understanding of the information within the current perception information.
[0070] In some examples, the step of performing spatial transformation on the discrete spatial representations of each scene may further include performing coordinate system transformation on the spatial representation using the driving trajectory of the autonomous driving vehicle. For example, the coordinate system transformation may be performed after the discrete spatial representations of each scene are processed using the Swin Transformer. The driving trajectory from the current moment to the future moment may be determined based on the autonomous driving decision information for the current moment output by the decoding layer at the previous moment, and the transformed spatial representation in the coordinate system of the current moment may be mapped to the transformed spatial representation in the coordinate system of the future moment based on the driving trajectory. For another example, the coordinate system transformation may also be performed before the discrete spatial representations of each scene are processed using the Swin Transformer. The discrete spatial representation of the scene in the coordinate system of the current moment may be mapped to the discrete spatial representation of the scene in the coordinate system of the future moment based on the driving trajectory, and then the discrete spatial representation of the scene in the coordinate system of the future moment may be processed using the Swin Transformer to output the transformed spatial representation in the coordinate system of the future moment.
[0071] For each position in the predicted spatial representation, feature transformation is performed based on the vector representation of the corresponding position in each transformed spatial representation and the vector representation of the corresponding position in the predetermined information to determine the vector representation of the position in the predicted spatial representation.
[0072] The feature transformation described above can be implemented using various suitable neural network models. For example, a generative diffusion model can be used to implement the feature transformation. In a generative diffusion model, vector representations of corresponding positions in the transformed spatial representation and vector representations of corresponding positions in the predetermined information can be input into a Transformer network to implement the feature transformation.
[0073] In some implementations, the final prediction space representation can be obtained through multiple feature transformations. That is, the output of the neural network model used for feature transformation (such as the Transformer mentioned above) and the scene discrete space representation can be input into the feature transformation network again, and the above steps can be repeated multiple times to obtain the prediction space representation.
[0074] For example, performing feature transformation based on the vector representation of the corresponding position in each transformed discrete spatial representation and the vector representation of the corresponding position in the predetermined information may include: using a Transformer to process the vector representation of the corresponding position in each transformed spatial representation and the vector representation of the corresponding position in the predetermined information to obtain first future scene information. Then, using a Transformer to process the vector representation of the corresponding position in the transformed spatial representation and the vector representation of the corresponding position in the first future scene information to obtain second future scene information. Prediction information may be determined based on the second future scene information.
[0075] In some examples, the second future scene information can be directly used as the prediction space input into the decoding layer of the autonomous driving model. In other examples, the feature transformation step can be repeated using the second future scene information and the transformed spatial representation to obtain the final prediction space representation. The number of repetitions of the feature transformation step can be determined based on actual conditions.
[0076] The autonomous driving model can also combine current additional information to obtain prediction information. In some examples, the current additional information may include at least one of historical decision information, interaction information, and traffic information. It will be appreciated that, without departing from the principles of this disclosure, the additional information may include any information that can assist in autonomous driving decision-making.
[0077] Autonomous driving decision information may include information such as planned trajectories or vehicle control signals (e.g., signals controlling the throttle, brakes, steering, etc.). That is, historical decision information may include sequences of historical trajectories and / or historical control signals output by the autonomous driving model prior to the current moment. In some examples, historical decision information may include historical decision information from all moments since the start of the autonomous driving process (e.g., for moment t, historical decision information includes autonomous driving decision information from moments 0 to t-1). It may also include historical decision information within a predetermined time period prior to the current moment t (e.g., autonomous driving decision information from moment tk to moment t, where k represents a predetermined time range).
[0078] Interaction information may include at least one of traffic control information, passenger interaction information, and safety officer interaction information. Traffic control information may include actions and / or language from outside the vehicle for traffic control purposes. Passenger interaction information and safety officer interaction information may include actions and / or language collected inside the vehicle while the passenger or safety officer is inside, actions and / or language used to communicate with the vehicle while the passenger or safety officer is outside the vehicle, and instructions sent from the passenger or safety officer to the vehicle via a remote communication device. Interaction information may be information collected by sensors such as cameras and microphones, or information received remotely via a communication device. In some implementations, interaction information may include interaction information acquired at the current time t, or historical interaction information acquired prior to the current time t. Historical interaction information may include historical interaction information from all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), or historical interaction information within a predetermined time period prior to the current time t (e.g., from time tk to time t, where k represents a predetermined time range).
[0079] Traffic information may include at least one of speed limit information, map information, and navigation information for the current route. For example, map information may include lane information, stop line information, traffic light information, etc. In the example, traffic information may include lane-level or road-level traffic information. Traffic information may include traffic information at the current time t, and may also include historical traffic information at at least one previous time before the current time t. In some examples, historical traffic information may include historical traffic information at all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), and may also include historical traffic information within a predetermined time period before the current time t (e.g., from time tk to time t, where k represents a predetermined time range).
[0080] In this case, the prediction layer can perform discrete diffusion on the scene discrete spatial representation and the current additional information to obtain the predicted spatial representation.
[0081] The principle of the present disclosure will be described below by taking the case where the current additional information is interactive information as an example.
[0082] In some implementations, when the interaction information includes natural language information, the interaction information in natural language form can be vectorized using a pre-trained language model (PLM) or a large language model (LLM) to convert the natural language information into a tensor form. When the interaction information includes action information represented by an image, a suitable image recognition algorithm can be used to convert the information in the image into natural language information or vectorized information that can be processed by the machine. The tensor representation of the interaction information can be obtained using the above method. The tensor representation of the interaction information can be made to have the same size as the discrete space representation of the scene by configuring parameters, such as W×H.
[0083] In the example, the prediction layer can perform discrete diffusion based on the scene discrete space representation and the current interaction information to obtain the prediction space representation. For example, the discrete space representations of each scene and the tensor representation of the interaction information can be stacked together to obtain a modified discrete space representation of the scene. The vector representations at each position of the modified discrete space representation of the scene and the vector representation at the same position in the predetermined information can be input into the Transformer network for feature transformation to obtain information of the vector representation at the corresponding position in the prediction space representation. As mentioned above, the scene discrete space representation can be spatially transformed before the feature transformation. The final prediction space representation can be obtained by repeating the feature transformation step multiple times.
[0084] In the case where the current additional information includes historical decision information and / or traffic information, the historical decision information and / or traffic information can also be used in a similar manner to modify the scene discrete space representation. The historical decision information and / or traffic information can be encoded into a tensor representation of the same size as the scene discrete space representation in any suitable manner, and the scene discrete space representation and the tensor representation of the current additional information can be stacked together to obtain the modified scene discrete space representation. In the example, the historical trajectory of the vehicle can be determined based on the historical decision information, and the coordinate points of the historical trajectory can be encoded using a recurrent neural network or a Transformer to obtain a tensor representation of the historical decision information. In the example, the tensor representation of the traffic information can be obtained by mapping the traffic information to the BEV space.
[0085] Without departing from the principles of the present disclosure, the scene discrete space representation may be modified using the additional information in any other suitable manner, so that the modified scene discrete space representation incorporates both the scene discrete space representation information and the additional information. This enables the model to simultaneously learn and consider both the scene discrete space representation information and the additional information during reasoning.
[0086] The decoding layer 230 is configured to decode the predicted spatial representation to obtain autonomous driving decision information at a future moment.
[0087] In some embodiments, a Transformer network can be used to decode the predicted spatial representation to obtain autonomous driving decision information for future moments.
[0088] When the input of the autonomous driving model also includes current interaction information, the decoding layer can also output interaction information at a future moment as a response to the current interaction information.
[0089] When the prediction layer is configured to perform discrete diffusion based only on the discrete spatial representation of the scene, the predicted information for the future moment only includes the predicted information of the perception information. In this case, the decoding layer can be configured to decode the tensor representation of the predicted spatial representation and the current interaction information to obtain the interaction information and autonomous driving decision information for the future moment.
[0090] When the prediction layer is configured to perform discrete diffusion based on the scene discrete spatial representation and current interaction information to obtain the prediction spatial representation, the prediction information for the future moment includes both the prediction information of the perception information and the prediction information of the interaction information. In this case, the decoding layer can be configured to decode the prediction spatial representation to obtain the interaction information and autonomous driving decision information for the future moment.
[0091] By utilizing the embodiments of the present disclosure, the principle of the generative diffusion model can be used to obtain predictive information for generating autonomous driving decision information, so that the diffusion model's predictive ability for the future can be utilized in the autonomous driving decision process to improve the effectiveness of autonomous driving decisions.
[0092] FIG3 shows an exemplary flow chart of a method for training an autonomous driving model according to an embodiment of the present disclosure. The method shown in FIG3 can be used to train the autonomous driving model described in conjunction with FIG2.
[0093] In step S302, current sample perception information of the autonomous vehicle and actual driving decision information corresponding to the current sample perception information can be obtained. Training sample data for the autonomous driving model can be obtained by collecting data from human drivers driving the vehicle. The sample perception information can be obtained by collecting perception information collected by sensors installed on the vehicle at each moment, and the actual driving decision information can be collected accordingly by collecting the operating signals of the human driver at each moment.
[0094] In step S304, the current sample perception information may be encoded using the coding layer of the autonomous driving model to obtain a discrete spatial representation of the current scene.
[0095] In step S306, the prediction layer of the autonomous driving model may be used to perform discrete diffusion based on at least one sample scene discrete space representation including a sample discrete space representation of the current scene to determine a sample prediction space representation at a future moment.
[0096] In step S308, the sample prediction space representation can be decoded using the decoding layer of the autonomous driving model to obtain sample driving decision information at the future moment.
[0097] The above steps S304 to S308 can be performed using the autonomous driving model described in conjunction with FIG. 2 to generate sample driving decision information based on the current sample perception information.
[0098] In step S310 , the parameters of the autonomous driving model may be adjusted according to the difference between the sample driving decision information and the actual driving decision information.
[0099] For example, the decision loss function L of the autonomous driving model can be determined according to formula (1): cts :
[0100] Where t represents the current time, represents the sample driving decision information output by the autonomous driving model at time t, c t represents the actual driving decision information at time t.
[0101] In some embodiments, real future information corresponding to the current sample perception information can also be obtained, and the parameters of the autonomous driving model can be adjusted based on the difference between the sample prediction spatial representation and the spatial representation of the real future information.
[0102] Among them, the sample prediction space representation output by the autonomous driving model is the discretized representation of the prediction information at the future moment in the BEV space. The real future information collected during actual driving can be mapped into a discrete representation in the BEV space b t+1 The prediction loss function L of the autonomous driving model can be determined according to formula (2): fp :
[0103] Where t represents the current time, represents the predicted spatial representation at time t+1 generated based on the current perception information at time t, b t+1 It represents the spatial representation of the real future information at time t+1 in the BEV space.
[0104] In some embodiments, when the autonomous driving model is also configured to process interaction information, the decoding layer of the autonomous driving model may further output sample prediction interaction information. In some examples, step S308 may include using the decoding layer of the autonomous driving model to decode the sample prediction space representation and the tensor representation of the current sample interaction information to obtain the sample prediction interaction information and sample driving decision information at the future moment. In other examples, step S306 may include performing discrete diffusion based on the sample scene discrete space representation and the current sample interaction information to obtain the sample prediction space representation. In this case, step S308 may include using the decoding layer to decode the sample prediction space representation to obtain the sample prediction interaction information and sample driving decision information.
[0105] When the decoding layer of the autonomous driving model outputs sample predicted interaction information, the real interaction information corresponding to the current sample interaction information can be obtained, and the parameters of the autonomous driving model can be adjusted to maximize the probability that the sample predicted interaction information is the real interaction information. The interaction loss function L can be determined according to formula (3): nll : L nll =∑ t logp(u t ) (3)
[0106] Among them, t represents the current moment, p(u t ) represents the predicted probability of the true interaction information at time t.
[0107] Based on the above decision loss function L cts , prediction loss function L fp And the interaction loss function L nll At least one of the above may determine a loss function for adjusting parameters of the autonomous driving model. The parameters in the autonomous driving model may be adjusted by minimizing the loss function.
[0108] In some embodiments, step S304 may include: mapping the current sample perception information to the bird's-eye view BEV space to obtain a sample continuous BEV representation of the current sample perception information, and discretizing the sample continuous BEV representation according to a pre-trained vocabulary to obtain a sample discrete space representation of the current scene.
[0109] Among them, the vocabulary can be generated in the following manner: obtaining sample sensor input; mapping the sample sensor input to the BEV space to obtain a continuous BEV representation of the sample sensor input; for the vector representation at each position of the continuous BEV representation of the sample sensor input, replacing the vector representation with the nearest vocabulary vector in the vocabulary to obtain a discrete representation of the sample sensor input; decoding the discrete representation of the sample sensor input to obtain a recovered sensor input; adjusting the parameters of the vocabulary by minimizing the difference between the recovered sensor input and the sample sensor input.
[0110] FIG4 illustrates an exemplary flowchart of an autonomous driving method according to an embodiment of the present disclosure. The autonomous driving model described in conjunction with FIG2 can be utilized to implement the autonomous driving method illustrated in FIG4 . The advantages of the autonomous driving model described in conjunction with FIG2 also apply to autonomous driving method 400 and are not further elaborated herein.
[0111] In step S402, the current perception information of the autonomous driving vehicle can be encoded using the encoding layer of the autonomous driving model to obtain a discrete spatial representation of the current scene.
[0112] In step S404, the prediction layer of the autonomous driving model may be used to perform discrete diffusion based on at least one discrete spatial representation of a scene including a discrete spatial representation of a current scene to determine a predicted spatial representation at a future moment.
[0113] In step S406, the predicted spatial representation may be decoded using the decoding layer of the autonomous driving model to obtain autonomous driving decision information for the future moment.
[0114] FIG5 shows an exemplary process of an autonomous driving method according to an embodiment of the present disclosure.
[0115] As shown in Figure 5, sensor inputs at time t-2, time t-1, and time t, namely, sensor information (t-2) 501-1, sensor information (t-1) 501-2, and sensor information (t) 501-3, can be encoded to obtain continuous representations 502-1, 502-1, and 502-3 of the sensor information in the BEV space. Furthermore, the continuously represented sensor information can be discretized using a vocabulary to obtain discrete spatial representations of the corresponding scene, namely, BEV(t-2) 503-1, BEV(t-1) 503-2, and BEV(t) 503-3.
[0116] Prediction layer 510 can be used to perform discrete diffusion based on BEV(t-2), BEV(t-1), and BEV(t) to predict prediction information at time t+1. Historical decision information (t-2) 504-1, historical decision information (t-1) 504-2, and decision information (t) 504-3 can be simultaneously input into the prediction layer along with BEV(t-2), BEV(t-1), and BEV(t) to generate prediction result PRED_BEV(t+1) 506. As previously described, coordinate system transformations can be performed on BEV(t-2), BEV(t-1), and BEV(t) using decision information (t-2), decision information (t-1), and decision information (t), respectively.
[0117] After obtaining the prediction result PRED_BEV(t+1), PRED_BEV(t+1) can be compared with the actual sample information BEV(t+1) 503-4, and the parameters of the autonomous driving model can be adjusted based on the difference between PRED_BEV(t+1) and the actual BEV(t+1). The perception information (t+1) 501-4 can be encoded to obtain a continuous representation 502-4 of the BEV space, and the continuous representation 502-4 can be discretized to obtain a discrete representation BEV(t+1) 503-4. Furthermore, the prediction information PRED_BEV(t+2) 507 at time t+2 can be further predicted based on PRED_BEV(t+1), BEV(t), and BEV(t-1). PRED_BEV(t+2) can be compared with the true sample information BEV(t+2) 503-5, and the parameters of the autonomous driving model can be adjusted based on the difference between PRED_BEV(t+2) and the true BEV(t+2). The perception information (t+2) 501-5 can be encoded to obtain a continuous representation 502-5 of the BEV space, and the continuous representation 502-5 can be discretized to obtain a discrete representation BEV(t+2) 503-5.
[0118] Furthermore, the interaction information encoding unit 520 can be used to encode the current interaction information at time t, and the traffic information encoding unit 530 can be used to encode the current traffic information at time t to obtain a tensor representation 504 of the traffic and interaction prompt information at time t. The decoding layer can be used to decode the tensor representation of the traffic and interaction prompt information and the predicted information PRED_BEV(t+1) 505 to obtain the interaction information and driving decision at time t+1. The decoding layer 540 can also be used to decode the tensor representation of the traffic and interaction prompt information and the predicted information PRED_BEV(t+2) to obtain the interaction information and driving decision at time t+2.
[0119] FIG6 shows an exemplary diagram of a prediction layer according to an embodiment of the present disclosure.
[0120] As shown in FIG6 , before discrete diffusion begins, the prediction space representation is initialized to predetermined information PRED_BEV(t+1,0), such as information of a special word [MASK].
[0121] The discrete representations 601-1, 601-2, and 601-3 of the perception information can be spatially transformed using a Swin Transformer to obtain transformed discrete representations 602-1, 602-2, and 602-3. Accordingly, the predetermined information PRED_BEV(t+1,0) 603 can also be discretely represented to obtain transformed predetermined information 604.
[0122] For the W×H vectors in the transformed discrete representation, the vector at the same position in the discrete representations 602-1, 602-2, and 602-3 (the lower left corner is shown as an example in FIG6 ) is extracted and input into the Transformer together with the vector at the same position in the transformed predetermined information to obtain the information at the corresponding position in the first prediction information PRED_BEV(t+1,1). The first prediction information PRED_BEV(t+1,1) can be obtained by traversing the W×H vectors in the discrete representation in a similar manner.
[0123] Next, the above discrete diffusion process can be repeated with the first prediction information PRED_BEV(t+1,1) replacing the predetermined information PRED_BEV(t+1,0) to obtain the second prediction information PRED_BEV(t+1,2). Further, the above discrete diffusion process can be repeated with the first prediction information PRED_BEV(t+1,1) replacing the first prediction information PRED_BEV(t+1,1) to obtain the third prediction information PRED_BEV(t+1,3).
[0124] The third prediction information PRED_BEV(t+1,3) may be determined as the prediction information PRED_BEV(t+1) output for time t+1.
[0125] Figure 7 shows an exemplary diagram of the decoding layer according to an embodiment of the present disclosure. As shown in Figure 7, traffic and interaction prompt information and prediction information can be flattened into a one-dimensional vector and decoded by Transformer network 710. When traffic and interaction prompt information are input first and prediction information is input later, Transformer network 710 can decode the natural language interaction information first and then decode the driving decision information.
[0126] FIG8 illustrates an exemplary process of training a discretized vocabulary according to an embodiment of the present disclosure.
[0127] As shown in Figure 8 , the raw sensor input 801 can be processed using a sensor space fusion model 802 to obtain a continuous BEV representation 803. Model 802 can be implemented using BEVFormer, BEVFusion, or any other suitable model. Each vector in the continuous BEV representation can then be discretized by searching for its nearest neighbor in a vocabulary 804 to obtain a discrete representation 805 of the sensor input. Vocabulary 804 can include multiple vectors of the same dimension as the vectors in the BEV representation. Discrete representation 805 can then be decoded using a decoding unit 806 to recover the BEV representation into sensor data 807. Decoding unit 806 can be implemented using any method that corresponds to the encoding process of model 802. By minimizing the difference between the raw input sensor data 801 and the recovered sensor data 807, parameters of sensor space fusion model 802 and vocabulary 804 can be adjusted for use in the autonomous driving model provided by embodiments of the present disclosure.
[0128] FIG9 shows a block diagram of an automatic driving device 900 according to an embodiment of the present disclosure. As shown in FIG9 , the automatic driving device 900 includes an encoding unit 910 , a prediction unit 920 , and a decoding unit 930 .
[0129] The encoding unit 910 can be configured to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene.
[0130] The prediction unit 920 may be configured to perform discrete diffusion based on at least one scene discrete spatial representation including the discrete spatial representation of the current scene to determine a predicted spatial representation at a future time.
[0131] The decoding unit 930 can be configured to decode the predicted spatial representation to obtain autonomous driving decision information at a future moment.
[0132] It should be understood that the various modules or units of the apparatus 900 shown in FIG9 may correspond to the various steps in the method 400 described with reference to FIG4 . Thus, the operations, features, and advantages described above for the method 400 are also applicable to the apparatus 900 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.
[0133] FIG10 shows a structural block diagram of an apparatus 1000 for training an autonomous driving model according to an embodiment of the present disclosure.
[0134] As shown in FIG. 10 , the apparatus 1000 includes an acquiring unit 1010 , an encoding unit 1020 , a predicting unit 1030 , a decoding unit 1040 , and a parameter adjusting unit 1050 .
[0135] The acquisition unit 1010 can be configured to acquire current sample perception information of the autonomous driving vehicle and actual driving decision information corresponding to the current sample perception information.
[0136] The encoding unit 1020 may be configured to encode the current sample perception information to obtain a sample discrete space representation of the current scene.
[0137] The prediction unit 1030 may be configured to perform discrete diffusion according to at least one sample scene discrete space representation including the sample discrete space representation of the current scene, to determine a sample prediction space representation at a future moment.
[0138] The decoding unit 1040 may be configured to decode the sample prediction space representation to obtain sample driving decision information at a future moment.
[0139] The parameter adjustment unit 1050 may be configured to adjust the parameters of the autonomous driving model according to the difference between the sample driving decision information and the actual driving decision information.
[0140] It should be understood that the various modules or units of the apparatus 1000 shown in FIG10 may correspond to the various steps in the method 300 described with reference to FIG3 . Thus, the operations, features, and advantages described above for the method 300 are also applicable to the apparatus 1000 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.
[0141] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various units discussed herein may be separated into multiple units, and / or at least some functionality of multiple units may be combined into a single unit.
[0142] It should also be understood that various technologies can be described herein in the general context of software hardware elements or program modules. The various units described above with respect to Figures 9 and 10 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program codes / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of units 910 to 930, preferably 1010 to 1050, can be implemented together in a system on chip (SoC). SoC can include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and can optionally execute the received program code and / or include embedded firmware to perform functions.
[0143] According to another aspect of the present disclosure, an electronic device is also provided, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the autonomous driving method according to an embodiment of the present disclosure.
[0144] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, where the computer instructions are used to enable the computer to execute the automatic driving method according to an embodiment of the present disclosure.
[0145] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the autonomous driving method according to an embodiment of the present disclosure when executed by a processor.
[0146] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising an autonomous driving device according to an embodiment of the present disclosure and one of the above-mentioned electronic devices.
[0147] With reference to Figure 11, a block diagram of an electronic device 1100 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0148] As shown in FIG11 , the electronic device 1100 includes a computing unit 1101 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. Various programs and data required for the operation of the electronic device 1100 can also be stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0149] Multiple components in the electronic device 1100 are connected to the I / O interface 1105, including: an input unit 1106, an output unit 1107, a storage unit 1108, and a communication unit 1109. The input unit 1106 can be any type of device that can input information to the electronic device 1100. The input unit 1106 can receive input digital or character information and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1107 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1108 can include but is not limited to a magnetic disk and an optical disk. The communication unit 1109 allows the electronic device 1100 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device and / or the like.
[0150] The computing unit 1101 may be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1101 performs the various methods and processes described above, such as methods (or processes) 300 and 400. For example, in some embodiments, the method (or process) 300 may be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the methods (or processes) 300 and 400 described above may be performed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the methods (or processes) 300 and 400 in any other appropriate manner (eg, by means of firmware).
[0151] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0152] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0153] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0155] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0156] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0157] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0158] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.
Claims
1. An autonomous driving model, comprising a coding layer, a prediction layer and a decoding layer, wherein: The encoding layer is configured to encode current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of the current scene; The prediction layer is configured to perform discrete diffusion based on at least one discrete spatial representation of a scene including a discrete spatial representation of the current scene to determine a predicted spatial representation at a future time; as well as The decoding layer is configured to decode the prediction space representation to obtain the autonomous driving decision information at the future moment.
2. The automatic driving model according to claim 1, wherein: Encoding the current perception information of the autonomous vehicle includes: Mapping the current perception information to a bird's-eye view BEV space to obtain a continuous BEV representation of the current perception information; The continuous BEV representation is discretized according to a pre-trained vocabulary to obtain a discrete spatial representation of the current scene.
3. The automatic driving model according to claim 2, wherein: Performing discrete diffusion according to at least one discrete spatial representation of a scene including a discrete spatial representation of the current scene comprises: Performing spatial transformation on each discrete spatial representation of the scene to obtain a corresponding transformed spatial representation; For each position in the predicted spatial representation, feature transformation is performed based on the vector representation of the corresponding position in each transformed spatial representation and the vector representation of the corresponding position in the predetermined information to determine the vector representation of the position in the predicted spatial representation.
4. The automatic driving model as claimed in claim 3, wherein: Performing feature transformation according to the vector representation of the corresponding position in each of the transformed spatial representations and the vector representation of the corresponding position in the predetermined information comprises: Using a Transformer, processing the vector representation of the corresponding position in each of the transformed spatial representations and the vector representation of the corresponding position in the predetermined information to obtain first future scene information; Using a Transformer, processing the vector representation of the corresponding position in each of the transformed spatial representations and the vector representation of the corresponding position in the first future scene information to obtain second future scene information; The predicted spatial representation is determined based on the second future scene information.
5. The automatic driving model as claimed in claim 3, wherein: Performing spatial transformation on the discrete spatial representation of each scene to obtain a corresponding transformed spatial representation includes: The discrete space representation of the scene is processed using Swin Transformer.
6. The automatic driving model as claimed in claim 5, wherein: Performing spatial transformation on the discrete spatial representation of each scene further includes: Determine a driving trajectory from the current moment to the future moment according to the automatic driving decision information for the current moment output by the decoding layer at the previous moment; The transformed spatial representation in the coordinate system at the current moment is mapped to the transformed spatial representation in the coordinate system at the future moment based on the driving trajectory.
7. The automatic driving model as claimed in claim 3, wherein: The predetermined information includes predetermined noise.
8. The automatic driving model as claimed in claim 3, wherein: The decoding layer is configured to decode the predicted spatial representation and the tensor representation of the current interaction information to obtain the interaction information and the autonomous driving decision information at the future moment.
9. The automatic driving model according to any one of claims 1 to 8, wherein: Performing discrete diffusion according to at least one discrete spatial representation of a scene including a discrete spatial representation of the current scene comprises: Discrete diffusion is performed according to the discrete spatial representation of the scene and the current interactive information to obtain the predicted spatial representation.
10. The automatic driving model according to claim 9, wherein: The decoding layer is configured to decode the predicted spatial representation to obtain the interaction information and autonomous driving decision information at the future moment.
11. The automatic driving model according to any one of claims 1 to 10, wherein: The at least one scene discrete space representation includes a discrete space representation of the current scene and a discrete space representation of at least one historical scene.
12. A method for training an autonomous driving model, comprising: Acquire current sample perception information of the autonomous driving vehicle and actual driving decision information corresponding to the current sample perception information; Encoding the current sample perception information using the encoding layer of the autonomous driving model to obtain a sample discrete space representation of the current scene; Using the prediction layer of the autonomous driving model, discrete diffusion is performed based on at least one sample scene discrete spatial representation including the sample discrete spatial representation of the current scene to determine a sample prediction spatial representation at a future moment; Decoding the sample prediction space representation using a decoding layer of the autonomous driving model to obtain sample driving decision information at the future moment, and The parameters of the autonomous driving model are adjusted according to the difference between the sample driving decision information and the actual driving decision information.
13. The method of claim 12, further comprising: Acquire real future information corresponding to the current sample perception information; The parameters of the autonomous driving model are adjusted according to the difference between the spatial representation of the sample prediction and the spatial representation of the real future information.
14. The method according to claim 12 or 13, wherein: Decoding the sample prediction space representation using a decoding layer of the autonomous driving model includes: The decoding layer of the autonomous driving model is used to decode the tensor representation of the sample prediction space representation and the current sample interaction information to obtain the sample prediction interaction information and sample driving decision information at the future moment.
15. The method of any one of claims 12 to 14, wherein using the prediction layer of the autonomous driving model to perform discrete diffusion according to at least one sample scene discrete space representation including a sample discrete space representation of the current scene comprises: The prediction layer of the autonomous driving model is used to discretely diffuse the predetermined information according to the discrete spatial representation of the sample scene and the current sample interaction information. Utilizing the decoding layer of the autonomous driving model to decode the sample prediction space representation to obtain the sample driving decision information at the future moment includes: utilizing the decoding layer of the autonomous driving model to decode the sample prediction space representation to obtain the sample driving decision information and sample prediction interaction information at the future moment.
16. The method of claim 14 or 15, further comprising: Acquire real interaction information corresponding to the current sample interaction information; The parameters of the autonomous driving model are adjusted to maximize the probability that the sample predicted interaction information is the true interaction information.
17. The method according to any one of claims 12 to 16, wherein: Encode the current sample perception information: Mapping the current sample perception information to a bird's-eye view BEV space to obtain a sample-continuous BEV representation of the current sample perception information; The sample continuous BEV representation is discretized according to the pre-trained vocabulary to obtain a sample discrete space representation of the current scene.
18. The method of claim 17, wherein: The vocabulary is generated in the following way: Get sample sensor input; Mapping the sample sensor input to the BEV space to obtain a continuous BEV representation of the sample sensor input; For each position of the continuous BEV representation of the sample sensor input, replace the vector representation with the nearest vocabulary vector in the vocabulary to obtain a discrete representation of the sample sensor input; decoding the discrete representation of the sample sensor input to obtain a recovered sensor input; Parameters of the vocabulary are adjusted by minimizing a difference between the restored sensor input and the sample sensor input.
19. An automatic driving method implemented by using the automatic driving model according to any one of claims 1 to 11, comprising: Encoding the current perception information of the autonomous driving vehicle using the encoding layer of the autonomous driving model to obtain a discrete spatial representation of the current scene; performing discrete diffusion using a prediction layer of the autonomous driving model based on at least one discrete spatial representation of a scene including a discrete spatial representation of the current scene to determine a predicted spatial representation at a future time; as well as The prediction space representation is decoded using a decoding layer of the autonomous driving model to obtain autonomous driving decision information at the future moment.
20. An automatic driving device based on the automatic driving model according to any one of claims 1 to 11, comprising: an encoding unit configured to encode the current perception information of the autonomous driving vehicle to obtain a discrete spatial representation of a current scene; A prediction unit configured to perform discrete diffusion based on at least one discrete spatial representation of a scene including a discrete spatial representation of the current scene to determine a predicted spatial representation at a future moment; as well as A decoding unit is configured to decode the predicted spatial representation to obtain the autonomous driving decision information at the future moment.
21. A device for training an autonomous driving model, comprising: an acquisition unit configured to acquire current sample perception information of the autonomous driving vehicle and real driving decision information corresponding to the current sample perception information; an encoding unit configured to encode the current sample perception information to obtain a sample discrete space representation of the current scene; A prediction unit configured to perform discrete diffusion according to at least one sample scene discrete space representation including the sample discrete space representation of the current scene to determine a sample prediction space representation at a future moment; a decoding unit configured to decode the sample prediction space representation to obtain the sample driving decision information at the future moment, and A parameter adjustment unit is configured to adjust the parameters of the automatic driving model according to the difference between the sample driving decision information and the actual driving decision information.
22. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 12 to 19.
23. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 12-19.
24. A computer program product comprising a computer program, wherein: The computer program implements the method of any one of claims 12 to 19 when executed by a processor.
25. An autonomous driving vehicle comprising: One of the automatic driving device according to claim 20 and the electronic device according to claim 22.
Citation Information
Patent Citations
Discrete Transform-based point cloud 3D target detection method and model
CN116152579A
Automatic driving model capable of autonomously interacting with personnel outside vehicle and training method
CN116776151A
Automatic driving track prediction method and device, electronic equipment and storage medium
CN117104275A
Automatic driving model, method and device based on generative diffusion model and vehicle
CN117519206A
Vehicle trajectory prediction method and system, computer device and storage medium
WO2023221348A1
Cited By
Automatic driving behavior decision robust optimization method guided by long-period historical trajectory
CN120993734A
Automatic driving scene risk assessment method fusing diffusion model and random forest
CN121233964A
Diffusion model training method and device based on spatial knowledge graph guidance, spatio-temporal data generation method and device, equipment and medium
CN121436072A