Autonomous driving model, method, apparatus and vehicle capable of achieving multi-modal interaction
By designing a multimodal interactive autonomous driving model to process historical decision-making information, perception information, traffic information and interaction information, the shortcomings of the autonomous driving model in terms of interpretability and controllability are solved, and more efficient autonomous driving decisions are achieved.
Patent Information
- Application Number
- PCT/CN2024/099441
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-29
- Filing Date
- 2024-06-14
- Publication Date
- 2025-06-05
AI Technical Summary
The output results of the autonomous driving model are insufficient in terms of interpretability and controllability, and it is difficult to effectively explain and control autonomous driving decisions.
A multimodal interaction autonomous driving model is designed, including an input layer, a coding layer, an autoregressive reasoning layer and a decoding layer. It can receive and process historical decision information, perception information, traffic information and interaction information, and generate hidden states and autonomous driving decision information at the next moment through the autoregressive reasoning layer.
Through the multimodal interaction processing method, the interpretability and controllability of the autonomous driving model are improved, so that the model can better understand the impact of the driving environment and historical operations, thereby generating more reliable autonomous driving decisions.
Smart Images

Figure CN2024099441_05062025_PF_FP_ABST
Abstract
Description
Autonomous driving model, method, device and vehicle capable of realizing multimodal interaction
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application No. 202311616173.1 filed on November 29, 2023, the entire contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present disclosure relates to the fields of computer technology, in particular to the fields of autonomous driving and artificial intelligence technology, and specifically to an autonomous driving model, an autonomous driving method implemented using the autonomous driving model, an autonomous driving device based on the autonomous driving model, an electronic device, a computer-readable storage medium, a computer program product, and an autonomous driving vehicle. Background Art
[0004] An autonomous driving vehicle may be equipped with an autonomous driving model for generating control signals for autonomous driving. The autonomous driving model can be trained using massive amounts of data, thereby improving the effectiveness of the autonomous driving strategy generated by the autonomous driving model.
[0005] However, the interpretability and controllability of the results output by the autonomous driving model are bottlenecks in the application of unmanned driving models.
[0006] Summary of the Invention
[0007] The present disclosure provides an autonomous driving model, method, device, and vehicle capable of realizing multimodal interaction.
[0008] According to one aspect of the present disclosure, an autonomous driving model is provided, comprising an input layer, an encoding layer, an autoregressive inference layer, and a decoding layer, wherein the input layer is configured to receive historical decision information, perception information, traffic information, and interaction information at a current moment; the encoding layer is configured to encode the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor for representing the historical decision information, the perception information, the traffic information, and the interaction information, respectively; the autoregressive inference layer is configured to infer an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for the next moment; and the decoding layer is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0009] According to another aspect of the present disclosure, there is provided an autonomous driving method implemented using an autonomous driving model, including: obtaining input information, the input information including historical decision information, perception information, traffic information and interaction information at a current moment; encoding the historical decision information, the perception information, the traffic information and the interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing the historical decision information, the perception information, the traffic information and the interaction information, respectively; inferring an input tensor formed by the first tensor, the second tensor, the third tensor and the fourth tensor to obtain a hidden state for a next moment; and decoding based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0010] According to another aspect of the present disclosure, an automatic driving device based on an automatic driving model is provided, including: an acquisition unit, configured to acquire input information, the input information including historical decision information, perception information, traffic information and interaction information at a current moment; an encoding unit, configured to encode the historical decision information, the perception information, the traffic information and the interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing the historical decision information, the perception information, the traffic information and the interaction information, respectively; an inference unit, configured to infer the input tensor formed by the first tensor, the second tensor, the third tensor and the fourth tensor to obtain a hidden state for the next moment; and a decoding unit, configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and automatic driving decision information for the next moment.
[0011] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform the above method.
[0012] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.
[0013] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.
[0014] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising: an autonomous driving device according to an embodiment of the present disclosure, or one of an electronic device.
[0015] By utilizing the embodiments of the present disclosure, the autonomous driving model can understand the current driving environment by reasoning about perception information, traffic information, and interaction information. It can further combine reasoning about historical decision data to better understand the impact of historical operations on the autonomous driving process, thereby making the output results of the autonomous driving model more interpretable and controllable.
[0016] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.
[0018] FIG1 shows a schematic diagram of an exemplary system in which various methods described herein may be implemented according to an embodiment of the present disclosure;
[0019] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure;
[0020] FIG3 shows an exemplary flowchart of an autonomous driving method implemented by using an autonomous driving model according to an embodiment of the present disclosure;
[0021] FIG4 shows a structural block diagram of an automatic driving device according to an embodiment of the present disclosure; and
[0022] FIG5 shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0023] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0024] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.
[0025] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.
[0026] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0027] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0028] FIG1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein may be implemented according to an embodiment of the present disclosure. Referring to FIG1 , the system 100 includes a motor vehicle 110, a server 120, and one or more communication networks 130 coupling the motor vehicle 110 to the server 120.
[0029] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to perform a method according to an embodiment of the present disclosure.
[0030] The server 120 may run one or more services or software applications that enable autonomous driving. In some embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In the configuration shown in Figure 1, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of the motor vehicle 110 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.
[0031] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0032] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.
[0033] In some embodiments, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from motor vehicle 110. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.
[0034] The network 130 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 130 may be a satellite communication network, a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including, for example, Bluetooth, WiFi), and / or any combination of these and other networks.
[0035] The system 100 may also include one or more databases 150. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repository 150 can reside in a variety of locations. For example, the data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repository 150 can be of different types. In some embodiments, the data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.
[0036] In some embodiments, one or more of the databases 150 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.
[0037] Motor vehicle 110 may include sensors 111 for sensing its surroundings. Sensors 111 may include one or more of the following: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radar, and laser radar (LiDAR). Different sensors offer different detection accuracy and range. Cameras may be mounted on the front, rear, or other locations of the vehicle. Visual cameras can capture real-time information about the vehicle's interior and exterior and present it to the driver and / or passengers. Furthermore, by analyzing the images captured by the visual cameras, information such as traffic light indications, intersection conditions, and the operating status of other vehicles can be obtained. Infrared cameras can detect objects in night vision conditions. Ultrasonic sensors can be mounted on all sides of the vehicle, utilizing the strong directionality of ultrasonic waves to measure the distance of external objects from the vehicle. Millimeter-wave radars can be mounted on the front, rear, or other locations of the vehicle, utilizing the properties of electromagnetic waves to measure the distance of external objects from the vehicle. LiDARs can be mounted on the front, rear, or other locations of the vehicle, detecting object edges and shapes for object recognition and tracking. Due to the Doppler effect, radar devices can also measure changes in the speed of the vehicle and moving objects.
[0038] The motor vehicle 110 may also include a communication device 112. The communication device 112 may include a satellite positioning module that can receive satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generate coordinates based on these signals. The communication device 112 may also include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as GSM / GPRS, CDMA, LTE, and other current or evolving wireless communication technologies (e.g., 5G technology). The communication device 112 may also have a vehicle-to-everything (V2X) module that is configured to implement vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may also include a module configured to communicate with a user terminal 145 (including but not limited to a smartphone, tablet computer, or wearable device such as a watch) via a wireless local area network or Bluetooth using the IEEE 802.11 standard, for example. Using the communication device 112, the motor vehicle 110 may also access the server 120 via the network 130.
[0039] The motor vehicle 110 may also include a control device 113. The control device 113 may include a processor that communicates with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, and braking system of the motor vehicle 110 (not shown) via multiple actuators in response to input from multiple sensors 111 or other input devices to control acceleration, steering, and braking, respectively, without human intervention or limited human intervention. Some processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an on-board processor, while other processing may be performed using computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.
[0040] The system 100 of FIG. 1 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with this disclosure.
[0041] End-to-end autonomous driving models can continuously achieve better performance based on massive data, but explainability and controllability are bottlenecks in the application of end-to-end autonomous driving models.
[0042] In order to improve the effect of the autonomous driving model, the present disclosure provides a new autonomous driving model.
[0043] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure.
[0044] As shown in FIG2 , the autonomous driving model 200 includes an input layer 210 , an encoding layer 220 , an autoregressive inference layer 230 , and a decoding layer 240 .
[0045] The input layer 210 is configured to receive historical decision information 201 , perception information 202 , traffic information 203 , and interaction information 204 at a current moment.
[0046] The encoding layer 220 is configured to encode historical decision information, perception information, traffic information and interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing historical decision information, perception information, traffic information and interaction information respectively.
[0047] The autoregressive inference layer 230 is configured to perform inference on an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for the next moment;
[0048] The decoding layer 240 is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0049] Utilizing the autonomous driving model provided by the embodiments of the present disclosure, the autonomous driving model can understand the current driving environment by reasoning about perception information, traffic information, and interaction information. It can further combine reasoning about historical decision data to better understand the impact of historical operations on the autonomous driving process, thereby making the output results of the autonomous driving model more interpretable and controllable.
[0050] The principles of the present disclosure will be described in detail below.
[0051] The input layer 210 is configured to receive historical decision information 201 , perception information 202 , traffic information 203 , and interaction information 204 at a current moment.
[0052] The historical decision information 201 may include the autonomous driving decision information output by the decoding layer at at least one previous moment before the current moment t. In some implementations, the autonomous driving decision information may include information such as a planned trajectory or a control signal for the vehicle (e.g., a signal for controlling the throttle, brake, steering amplitude, etc.). That is, the historical decision information may include a sequence of historical trajectories and / or historical control signals output by the autonomous driving model before the current moment. In some examples, the historical decision information may include historical decision information at all moments since the start of the autonomous driving process (e.g., for moment t, the historical decision information includes autonomous driving decision information from moments 0 to t-1), or historical decision information within a predetermined time period before the current moment t (e.g., autonomous driving decision information from moment tk to moment t, where k represents a predetermined time range).
[0053] The perception information 202 may include sensor input collected by at least one sensor installed on the autonomous driving vehicle. The perception information for the vehicle's surrounding environment may include at least one of the following items: perception information of one or more cameras, perception information of one or more lidars, and perception information of one or more millimeter-wave radars. The perception information 202 may include sensor input collected by the sensor at the current time t, and may also include historical perception information of sensor input collected by the sensor at at least one previous time before the current time t. In some examples, the historical perception information may include historical perception information of all times since the start of the autonomous driving process (for example, from time 0 to time t-1), and may also include historical perception information within a predetermined time period before the current time t (for example, from time tk to time t, where k represents a predetermined time range).
[0054] Traffic information 203 may include at least one of speed limit information, map information, and navigation information for the current route. For example, map information may include lane information, stop line information, traffic light information, etc. In the example, traffic information may include lane-level or road-level traffic information. Traffic information 203 may include traffic information at the current time t, and may also include historical traffic information at at least one previous time before the current time t. In some examples, historical traffic information may include historical traffic information at all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), and may also include historical traffic information within a predetermined time period before the current time t (e.g., from time tk to time t, where k represents a predetermined time range).
[0055] Interaction information 204 may include at least one of traffic control information, passenger interaction information, and safety officer interaction information. Traffic control information may include gestures and / or speech from outside the vehicle used for traffic control purposes. Passenger interaction information and safety officer interaction information may include gestures and / or speech collected from inside the vehicle while the passenger or safety officer is inside, gestures and / or speech used to communicate with the vehicle while the passenger or safety officer is outside the vehicle, and instructions sent from the passenger or safety officer to the vehicle via a remote communication device. Interaction information 204 may be information collected by sensors such as cameras and microphones, or information received remotely via a communication device. In some implementations, interaction information 204 may include interaction information acquired at the current time t, or historical interaction information acquired prior to the current time t. Historical interaction information may include historical interaction information from all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), or historical interaction information within a predetermined time period prior to the current time t (e.g., from time tk to time t, where k represents a predetermined time range).
[0056] The encoding layer 220 can be configured to encode historical decision information, perception information, traffic information and interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing historical decision information, perception information, traffic information and interaction information, respectively, wherein the first tensor, the second tensor, the third tensor and the fourth tensor have the same spatial representation.
[0057] In some embodiments, the encoding layer may include a recurrent neural network or a Transformer network, and may be configured to encode historical decision information using the recurrent neural network or the Transformer network. The historical trajectory of the vehicle may be determined based on the historical decision information, and the coordinates of the trajectory points at each historical moment in the vehicle coordinate system at the current moment t may be determined. Each trajectory coordinate in the historical trajectory may be input into the recurrent neural network or the Transformer network to obtain a one-dimensional vector or a two-dimensional tensor for representing the coordinates of the trajectory point. Furthermore, the one-dimensional vectors or two-dimensional tensors corresponding to the coordinates of the trajectory points at each historical moment may be stacked in chronological order to obtain a two-dimensional or three-dimensional vector with an added time dimension. As the first tensor used to represent historical decision information.
[0058] In some embodiments, the coding layer may further include a coding network for mapping information into a bird's-eye view BEV space, such as a BEVFormer, and configured to map the perception information into the BEV space to obtain a BEV representation of the perception information. The perception information collected by the sensor at each moment can be input into the BEVFormer network, and a BEV representation of the perception information at that moment can be obtained. In the example, the BEV representation can be a three-dimensional vector. In the case where the input information includes perception information at multiple moments, the BEV representations of the perception information at each moment can be stacked in chronological order to obtain a four-dimensional tensor with an added time dimension. As the second tensor used to represent the perceptual information.
[0059] The encoding layer can also be configured to map traffic information into BEV space to obtain a bird's-eye view BEV representation of traffic information. For example, the traffic information at each moment can be vectorized and encoded using BEVFormer to obtain the BEV representation of the traffic information at that moment. In the case where the input information includes traffic information at multiple moments, the BEV representations of the traffic information at each moment can be stacked in chronological order to obtain a four-dimensional tensor with an added time dimension. As the third tensor used to represent traffic information.
[0060] It is understandable that the BEVFormer for processing perception information and the BEVFormer for processing traffic information can be configured separately according to actual conditions.
[0061] In some embodiments, the encoding layer may further include a pre-trained language model (PLM). The pre-trained language model may be used to vectorize natural language to convert natural language information into information that can be processed by a machine. The pre-trained language model may be any model that can process input natural language information. In some examples, a large language model (LLM) may also be used to vectorize natural language. When the input interaction information includes action information represented by an image, a suitable image recognition algorithm may be used to convert the information in the image into natural language information or vectorized information that can be processed by a machine. The interaction information may be encoded using a pre-trained language model to convert the natural language information into a fourth tensor containing multiple dimensions (such as 2 dimensions). The natural language information may also include timestamp information.
[0062] Although the first tensor, the second tensor, the third tensor, and the fourth tensor are the results of output through different encoding methods, the above tensors can have the same length and width dimensions. For example, the first tensor used to represent historical decision information and the fourth tensor used to represent interaction information, although not BEV representation, can have the same spatial representation as the second tensor and the third tensor represented in the BEV space. In this way, the first tensor, the second tensor, the third tensor, and the fourth tensor can be processed uniformly in subsequent model processing, so that the model can uniformly process and reason on different input information when making inference decisions, so as to make autonomous driving decisions while considering input information of different modalities.
[0063] The autoregressive inference layer 230 is configured to perform inference on an input tensor formed of the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for the next moment.
[0064] In some embodiments, the autoregressive reasoning layer 230 can be implemented by a world model. In some implementations, the autoregressive reasoning layer can be implemented by a recurrent neural network (such as a long short-term memory network LSTM, a gated recurrent unit GRU), a Transformer, or a recursive structure with a memory mechanism (such as a recursive memory Transformer network (Recurrent Memory Transformer)) or a diffusion model. Using the autoregressive method, the input signal at time t can be used to predict the output result at time t+1. The combination of the first tensor, the second tensor, the third tensor and the fourth tensor output by the encoding layer can be used as the input of the autoregressive reasoning layer 230. For example, the first tensor, the second tensor, the third tensor and the fourth tensor can be flattened into a two-dimensional tensor sequence, and the above two-dimensional tensor sequence can be used as the input of the autoregressive reasoning layer.
[0065] Taking the recursive memory Transformer network as an example, the memory tensor M0 at the initial moment can be pre-set. Among them, the parameters in M0 can be randomly initialized parameters. The memory tensor M at time t can be t It is input into the Transformer layer together with the two-dimensional tensor sequence, so that each vector in the memory tensor is used to process the two-dimensional tensor sequence based on the attention mechanism to obtain the memory tensor M at the next moment. t+1 As the hidden state of the next moment. Among them, the memory tensor M t It is equivalent to the query parameter (Q) of the input Transformer layer, and the two-dimensional tensor sequence is equivalent to the key (K) and value (V) parameters of the input Transformer layer.
[0066] The decoding layer 240 is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0067] In some embodiments, the decoding layer may include a Transformer network. The Transformer can be used to transform the hidden state M at time t+1. t+1 Decoding is performed to obtain the autonomous driving decision information at time t+1, such as the control signals of throttle, brake, steering amplitude, etc. at time t+1.
[0068] You can also use Transformer to calculate the hidden state M at time t+1 t+1 , to obtain the interactive information output at time t+1. The interactive information output at time t+1 may include a natural language description, which can be used to respond to the interactive information at time t in natural language form.
[0069] You can also use the hidden state M at time t+1 t+1 Generate future prediction information. The future prediction information may include a future prediction image at time t+1, indicating the obstacle position at time t+1 or perception information at future times. The hidden state M that can be represented by BEV t+1 Perform spatial mapping to convert it to the sensor (such as camera) coordinate system. Image diffusion and deconvolution can be further used to transform M in the sensor coordinate system. t+1 The image is processed to obtain the future prediction image at time t+1. The future prediction image can be used to train the autonomous driving model in a self-supervised manner, so that the autonomous driving model has accurate future prediction capabilities, thereby improving the accuracy of the decision information output by the autonomous driving model.
[0070] The autonomous driving model provided by the embodiments of the present disclosure can receive multimodal input information (including historical decision information, perception information, traffic information, and interaction information) and process the multimodal input information to generate output. In the output, the driving decision and interaction information for the current moment can be output accordingly. In this way, the autonomous driving model can learn and consider information of different modalities when performing reasoning, thereby improving the controllability of the output results of the autonomous driving model. In addition, by outputting multimodal output results including interaction information, it can help to explain the decisions of the autonomous driving model, thereby improving the interpretability of the autonomous driving model.
[0071] According to another aspect of the present disclosure, an autonomous driving method is provided.
[0072] FIG3 illustrates an exemplary flowchart of an autonomous driving method 300 implemented using an autonomous driving model according to an embodiment of the present disclosure. The autonomous driving model described in conjunction with FIG2 can be used to implement autonomous driving method 300. The advantages of the autonomous driving model described in conjunction with FIG2 are also applicable to autonomous driving method 300 and will not be further elaborated here.
[0073] In step S302, input information may be obtained, including historical decision information, perception information, traffic information, and interaction information at the current moment;
[0074] In step S304, the historical decision information, the perception information, the traffic information, and the interaction information may be encoded to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor for representing the historical decision information, the perception information, the traffic information, and the interaction information, respectively. The first tensor, the second tensor, the third tensor, and the fourth tensor have the same spatial representation.
[0075] In step S306, the input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor may be inferred to obtain a hidden state for the next moment;
[0076] In step S308, decoding can be performed based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0077] In some embodiments, the historical decision information includes autonomous driving decision information output by the decoding layer at at least one previous moment before the current moment.
[0078] In some embodiments, the perception information includes sensor input collected by at least one sensor mounted on the autonomous vehicle.
[0079] In some embodiments, the traffic information includes at least one of speed limit information, map information, and navigation information of the current travel route.
[0080] In some embodiments, the interactive information includes at least one of traffic control information, interactive information from passengers, and interactive information from security officers.
[0081] In some embodiments, a recurrent neural network or a Transformer network may be used to encode historical decision information.
[0082] In some embodiments, the perceptual information may be mapped to a bird's-eye view (BEV) space to obtain a bird's-eye view (BEV) representation of the perceptual information.
[0083] In some embodiments, traffic information may be mapped to a bird's-eye view BEV space to obtain a bird's-eye view BEV representation of the traffic information.
[0084] In some embodiments, a pre-trained language model may be used to encode the interaction information.
[0085] In some embodiments, the autoregressive inference layer can be implemented by a recurrent neural network (such as a long short-term memory network LSTM, a gated recurrent unit GRU), a Transformer, or a recursive structure with a memory mechanism (such as a recurrent memory Transformer network (Recurrent Memory Transformer)) or a diffusion model.
[0086] By utilizing the autonomous driving method provided by the embodiments of the present disclosure, the autonomous driving model can understand the current driving environment by reasoning about perception information, traffic information, and interaction information. It can further combine reasoning about historical decision data to better understand the impact of historical operations on the autonomous driving process, thereby making the output results of the autonomous driving model more interpretable and controllable.
[0087] According to another aspect of the present disclosure, an automatic driving device based on an automatic driving model is provided.
[0088] FIG4 shows a block diagram of an automatic driving device 400 according to an embodiment of the present disclosure. As shown in FIG4 , the automatic driving device 400 includes an acquisition unit 410 , an encoding unit 420 , an inference unit 430 , and a decoding unit 440 .
[0089] The acquisition unit 410 may be configured to acquire input information, where the input information includes historical decision information, perception information, traffic information, and interaction information at a current moment.
[0090] The encoding unit 420 may be configured to encode the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor for representing the historical decision information, the perception information, the traffic information, and the interaction information, respectively. The first tensor, the second tensor, the third tensor, and the fourth tensor have the same spatial representation;
[0091] The inference unit 430 may be configured to perform inference on an input tensor formed of the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a next moment.
[0092] The decoding unit 440 can be configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
[0093] It should be understood that the various modules or units of the apparatus 400 shown in FIG4 may correspond to the various steps in the method 300 described with reference to FIG3 . Thus, the operations, features, and advantages described above for the method 300 are also applicable to the apparatus 400 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.
[0094] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various units discussed herein may be separated into multiple units, and / or at least some functionality of multiple units may be combined into a single unit.
[0095] It should also be understood that various technologies can be described herein in the general context of software hardware elements or program modules. The various units described above with respect to Figure 4 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of units 410 to 440 can be implemented together in a system on chip (SoC). SoC can include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and can optionally execute the received program code and / or include embedded firmware to perform a function.
[0096] According to another aspect of the present disclosure, an electronic device is also provided, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the autonomous driving method according to an embodiment of the present disclosure.
[0097] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, where the computer instructions are used to enable the computer to execute the automatic driving method according to an embodiment of the present disclosure.
[0098] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the autonomous driving method according to an embodiment of the present disclosure when executed by a processor.
[0099] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising an autonomous driving device according to an embodiment of the present disclosure and one of the above-mentioned electronic devices.
[0100] With reference to Figure 5, a block diagram of an electronic device 500 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0101] As shown in Figure 5, electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In RAM 503, various programs and data required for the operation of electronic device 500 can also be stored. Computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.
[0102] Multiple components within electronic device 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into electronic device 500. Input unit 506 can receive input numeric or character information and generate key signal input related to user settings and / or function control of the electronic device. It may include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. It may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0103] The computing unit 501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method (or process) 300. For example, in some embodiments, the method (or process) 300 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method (or process) 300 described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the method (or process) 300 in any other appropriate manner (eg, by means of firmware).
[0104] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0108] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0109] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0110] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0111] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.
Claims
1. An autonomous driving model, comprising an input layer, an encoding layer, an autoregressive inference layer and a decoding layer, wherein: The input layer is configured to receive historical decision information, perception information, traffic information and interaction information at the current moment; The encoding layer is configured to encode the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor for representing the historical decision information, the perception information, the traffic information, and the interaction information, respectively; The autoregressive inference layer is configured to perform inference on an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a next moment; as well as The decoding layer is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
2. The automatic driving model according to claim 1, wherein: The encoding layer is configured to encode the historical decision information using a recurrent neural network or a Transformer network.
3. The automatic driving model according to claim 2, wherein: The historical decision information includes autonomous driving decision information output by the decoding layer at least one previous moment before the current moment.
4. The automatic driving model according to any one of claims 1 to 3, wherein: The coding layer is configured to map the perceptual information to a bird's eye view (BEV) space to obtain a bird's eye view (BEV) representation of the perceptual information.
5. The autonomous driving model of claim 4, wherein the perception information comprises sensor input collected by at least one sensor mounted on the autonomous driving vehicle.
6. The automatic driving model according to any one of claims 1 to 5, wherein: The encoding layer is configured to map the traffic information to a bird's eye view (BEV) space to obtain a bird's eye view (BEV) representation of the traffic information.
7. The automatic driving model according to claim 6, wherein: The traffic information includes at least one of speed limit information, map information and navigation information of a current travel route.
8. The automatic driving model according to any one of claims 1 to 7, wherein: The encoding layer is configured to encode the interaction information using a pre-trained language model.
9. The automatic driving model as claimed in claim 8, wherein: The interactive information includes at least one of traffic control information, interactive information from passengers and interactive information from security personnel.
10. The automatic driving model according to any one of claims 1 to 9, wherein: The autoregressive inference layer is implemented by one of a recurrent neural network, a Transformer, a recursive memory Transformer, and a diffusion model.
11. An automatic driving method implemented by using the automatic driving model according to any one of claims 1 to 10, comprising: Acquiring input information, wherein the input information includes historical decision information, perception information, traffic information, and interaction information at a current moment; Encoding the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor respectively used to represent the historical decision information, the perception information, the traffic information, and the interaction information; Performing inference on an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a next moment; as well as Decoding is performed based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.
12. An automatic driving device based on the automatic driving model according to any one of claims 1 to 10, comprising: an acquisition unit configured to acquire input information, wherein the input information includes historical decision information, perception information, traffic information, and interaction information at a current moment; an encoding unit configured to encode the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor respectively used to represent the historical decision information, the perception information, the traffic information, and the interaction information; an inference unit, configured to infer an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a next moment; as well as The decoding unit is configured to decode based on the hidden state at the next moment to obtain the interaction information for the next moment and the automatic driving decision information for the next moment.
13. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of claim 11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to claim 11.
15. A computer program product comprising a computer program, wherein: The computer program implements the method of claim 11 when executed by a processor.
16. An autonomous driving vehicle, comprising: One of the automatic driving device according to claim 12 and the electronic device according to claim 13.
Citation Information
Patent Citations
Decision-making method and device for autonomous vehicle, equipment and medium
CN115366920A
Automatic driving model capable of autonomously interacting with personnel outside vehicle and training method
CN116776151A
Automatic driving model outputting explanation information, training method and device and vehicle
CN116861230A
Automatic driving model capable of performing natural language interaction and training method thereof
CN117010265A
Automatic driving method and device capable of achieving autonomous escape by following instructions and vehicle
CN117539253A
Cited By
Vehicle trajectory prediction method and device based on space-time cooperation and game driving
CN121019599A
Intelligent driving method and device and storage medium
CN122009220A
Intelligent driving method and device, and storage medium
CN122009220B