Autonomous driving method, apparatus and vehicle capable of following instructions for self-recovery

By adopting autoregressive inference and decoding technology in the autonomous driving system, combined with remote assistance, the problem of insufficient interpretability and controllability in autonomous driving technology is solved, and more efficient autonomous escape and cloud guidance are achieved.

WO2025112451A1PCT designated stage expired Publication Date: 2025-06-05BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099367
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-06-14
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing autonomous driving technology has problems of insufficient explanatory and controllability when dealing with complex driving environments and autonomous escape.

Method used

An autonomous driving method and device are adopted, which includes obtaining historical decision information, perception information, traffic information and interaction information at the current moment, performing encoding and autoregressive reasoning to generate hidden states and decoding to obtain information for autonomous driving decisions. The method further includes proactively initiating an interactive request to request remote assistance, when appropriate.

Benefits of technology

It improves the interpretability and controllability of the autonomous driving model, realizes the generation of rapid vehicle escape solutions, and supports cloud guidance for multiple autonomous driving vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099367_05062025_PF_FP_ABST
    Figure CN2024099367_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of computer technology and particularly relates to the technical fields of autonomous driving and artificial intelligence. Provided are an autonomous driving method, apparatus and vehicle capable of following instructions for self-recovery. The autonomous driving method comprises: acquiring input information; encoding the input information to obtain an input tensor corresponding to the input information; performing autoregressive inference on the input tensor to obtain a hidden state for a first moment after a current moment; and performing decoding on the basis of the hidden state of the first moment to obtain interaction information for the first moment and autonomous driving decision-making information for the first moment, wherein the interaction information for the first moment comprises a signal used for indicating that assistance is required during autonomous driving. In this way, prompt information for natural language interaction can be used to instruct an autonomous driving model to control a vehicle, thereby achieving the rapid generation of a self-recovery scheme. In addition, the vehicle can perform autonomous self-recovery on the basis of self-recovery scheme instructions.
Need to check novelty before this filing date? Find Prior Art

Description

Automatic driving method, device and vehicle capable of following instructions to achieve autonomous escape

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311612484.0 filed on November 29, 2023, the entire contents of which are incorporated by reference in their entirety into this application. Technical Field

[0003] The present disclosure relates to the fields of computer technology, in particular to the fields of autonomous driving and artificial intelligence technology, and specifically to an autonomous driving method, device, and vehicle capable of assisting in escaping distress. Background Art

[0004] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0005] Autonomous driving technology integrates technologies such as recognition, decision-making, positioning, communication security, and human-computer interaction. Artificial intelligence learning can assist in generating autonomous driving strategies.

[0006] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.

[0007] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.

[0008] Summary of the Invention

[0009] The present disclosure provides an automatic driving method, device, and vehicle that can follow instructions to achieve assisted autonomous escape.

[0010] According to one aspect of the present disclosure, an autonomous driving method is provided, comprising: obtaining input information, the input information including historical decision information, perception information, traffic information and interaction information at a current moment; encoding the historical decision information, the perception information, the traffic information and the interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing the historical decision information, the perception information, the traffic information and the interaction information, respectively; performing autoregressive inference on an input tensor formed by the first tensor, the second tensor, the third tensor and the fourth tensor to obtain a hidden state for a first moment after the current moment; and decoding based on the hidden state at the first moment to obtain interaction information for the first moment and autonomous driving decision information for the first moment; wherein the interaction information at the first moment includes a signal for indicating that the autonomous driving process requires assistance.

[0011] According to another aspect of the present disclosure, an automatic driving device is provided, including: an acquisition unit configured to acquire input information, the input information including historical decision information, perception information, traffic information and interaction information at a current moment; an encoding unit configured to encode the historical decision information, the perception information, the traffic information and the interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing the historical decision information, the perception information, the traffic information and the interaction information, respectively; an inference unit configured to infer the input tensor formed by the first tensor, the second tensor, the third tensor and the fourth tensor to obtain a hidden state for a first moment after the current moment; a decoding unit configured to decode based on the hidden state at the first moment to obtain interaction information for the next moment and automatic driving decision information for the first moment, wherein the interaction information at the first moment includes a signal for indicating that the automatic driving process requires assistance.

[0012] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform the above method.

[0013] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.

[0014] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0015] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising: an autonomous driving device according to an embodiment of the present disclosure, or one of an electronic device.

[0016] Using the embodiments of the present disclosure, natural language interactive prompts can be used to guide the autonomous driving model to control the vehicle, enabling the rapid generation of a vehicle escape plan. The vehicle can then autonomously escape according to the escape plan instructions. The autonomous driving model can proactively initiate interactive requests for remote assistance when appropriate, and cloud-based guidance can be implemented for multiple autonomous vehicles.

[0017] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0019] FIG1 shows a schematic diagram of an exemplary system in which various methods described herein may be implemented according to an embodiment of the present disclosure;

[0020] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure;

[0021] FIG3A shows an exemplary flowchart of an autonomous driving method implemented by using an autonomous driving model according to an embodiment of the present disclosure;

[0022] FIG3B shows an exemplary process for assisting autonomous driving according to an embodiment of the present disclosure.

[0023] FIG4 shows a structural block diagram of an automatic driving device according to an embodiment of the present disclosure; and

[0024] FIG5 shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION

[0025] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0026] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.

[0027] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.

[0028] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0029] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] FIG1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein may be implemented according to an embodiment of the present disclosure. Referring to FIG1 , the system 100 includes a motor vehicle 110, a server 120, and one or more communication networks 130 coupling the motor vehicle 110 to the server 120.

[0031] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to perform a method according to an embodiment of the present disclosure.

[0032] The server 120 may run one or more services or software applications that enable autonomous driving. In some embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In the configuration shown in Figure 1, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of the motor vehicle 110 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0033] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0034] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0035] In some embodiments, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from motor vehicle 110. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.

[0036] The network 130 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 130 may be a satellite communication network, a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including, for example, Bluetooth, WiFi), and / or any combination of these and other networks.

[0037] The system 100 may also include one or more databases 150. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repository 150 can reside in a variety of locations. For example, the data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repository 150 can be of different types. In some embodiments, the data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0038] In some embodiments, one or more of the databases 150 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0039] Motor vehicle 110 may include sensors 111 for sensing its surroundings. Sensors 111 may include one or more of the following: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radar, and laser radar (LiDAR). Different sensors offer different detection accuracy and range. Cameras may be mounted on the front, rear, or other locations of the vehicle. Visual cameras can capture real-time information about the vehicle's interior and exterior and present it to the driver and / or passengers. Furthermore, by analyzing the images captured by the visual cameras, information such as traffic light indications, intersection conditions, and the operating status of other vehicles can be obtained. Infrared cameras can detect objects in night vision conditions. Ultrasonic sensors can be mounted on all sides of the vehicle, utilizing the strong directionality of ultrasonic waves to measure the distance of external objects from the vehicle. Millimeter-wave radars can be mounted on the front, rear, or other locations of the vehicle, utilizing the properties of electromagnetic waves to measure the distance of external objects from the vehicle. LiDARs can be mounted on the front, rear, or other locations of the vehicle, detecting object edges and shapes for object recognition and tracking. Due to the Doppler effect, radar devices can also measure changes in the speed of the vehicle and moving objects.

[0040] The motor vehicle 110 may also include a communication device 112. The communication device 112 may include a satellite positioning module that can receive satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generate coordinates based on these signals. The communication device 112 may also include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as GSM / GPRS, CDMA, LTE, and other current or evolving wireless communication technologies (e.g., 5G technology). The communication device 112 may also have a vehicle-to-everything (V2X) module that is configured to implement vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may also include a module configured to communicate with a user terminal 145 (including but not limited to a smartphone, tablet computer, or wearable device such as a watch) via a wireless local area network or Bluetooth using the IEEE 802.11 standard, for example. Using the communication device 112, the motor vehicle 110 may also access the server 120 via the network 130.

[0041] The motor vehicle 110 may also include a control device 113. The control device 113 may include a processor that communicates with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, and braking system of the motor vehicle 110 (not shown) via multiple actuators in response to input from multiple sensors 111 or other input devices to control acceleration, steering, and braking, respectively, without human intervention or limited human intervention. Some processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an on-board processor, while other processing may be performed using computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.

[0042] The system 100 of FIG. 1 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with this disclosure.

[0043] End-to-end autonomous driving models can continuously achieve better performance based on massive data, but explainability and controllability are bottlenecks in the application of end-to-end autonomous driving models.

[0044] In order to improve the effect of the autonomous driving model, the present disclosure provides a new autonomous driving model.

[0045] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure.

[0046] As shown in FIG2 , the autonomous driving model 200 includes an input layer 210 , an encoding layer 220 , an autoregressive inference layer 230 , and a decoding layer 240 .

[0047] The input layer 210 is configured to receive historical decision information 201 , perception information 202 , traffic information 203 , and interaction information 204 at a current moment.

[0048] The encoding layer 220 is configured to encode historical decision information, perception information, traffic information and interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing historical decision information, perception information, traffic information and interaction information respectively.

[0049] The autoregressive inference layer 230 is configured to perform inference on an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for the next moment;

[0050] The decoding layer 240 is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.

[0051] Utilizing the autonomous driving model provided by the embodiments of the present disclosure, the autonomous driving model can understand the current driving environment by reasoning about perception information, traffic information, and interaction information. It can further combine reasoning about historical decision data to better understand the impact of historical operations on the autonomous driving process, thereby making the output results of the autonomous driving model more interpretable and controllable.

[0052] The principles of the present disclosure will be described in detail below.

[0053] The input layer 210 is configured to receive historical decision information 201 , perception information 202 , traffic information 203 , and interaction information 204 at a current moment.

[0054] The historical decision information 201 may include the autonomous driving decision information output by the decoding layer at at least one previous moment before the current moment t. In some implementations, the autonomous driving decision information may include information such as a planned trajectory or a control signal for the vehicle (e.g., a signal for controlling the throttle, brake, steering amplitude, etc.). That is, the historical decision information may include a sequence of historical trajectories and / or historical control signals output by the autonomous driving model before the current moment. In some examples, the historical decision information may include historical decision information at all moments since the start of the autonomous driving process (e.g., for moment t, the historical decision information includes autonomous driving decision information from moments 0 to t-1), or historical decision information within a predetermined time period before the current moment t (e.g., autonomous driving decision information from moment tk to moment t, where k represents a predetermined time range).

[0055] The perception information 202 may include sensor input collected by at least one sensor installed on the autonomous driving vehicle. The perception information for the vehicle's surrounding environment may include at least one of the following items: perception information of one or more cameras, perception information of one or more lidars, and perception information of one or more millimeter-wave radars. The perception information 202 may include sensor input collected by the sensor at the current time t, and may also include historical perception information of sensor input collected by the sensor at at least one previous time before the current time t. In some examples, the historical perception information may include historical perception information of all times since the start of the autonomous driving process (for example, from time 0 to time t-1), and may also include historical perception information within a predetermined time period before the current time t (for example, from time tk to time t, where k represents a predetermined time range).

[0056] Traffic information 203 may include at least one of speed limit information, map information, and navigation information for the current route. For example, map information may include lane information, stop line information, traffic light information, etc. In the example, traffic information may include lane-level or road-level traffic information. Traffic information 203 may include traffic information at the current time t, and may also include historical traffic information at at least one previous time before the current time t. In some examples, historical traffic information may include historical traffic information at all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), and may also include historical traffic information within a predetermined time period before the current time t (e.g., from time tk to time t, where k represents a predetermined time range).

[0057] Interaction information 204 may include at least one of traffic control information, passenger interaction information, and safety officer interaction information. Traffic control information may include gestures and / or speech from outside the vehicle used for traffic control purposes. Passenger interaction information and safety officer interaction information may include gestures and / or speech collected from inside the vehicle while the passenger or safety officer is inside, gestures and / or speech used to communicate with the vehicle while the passenger or safety officer is outside the vehicle, and instructions sent from the passenger or safety officer to the vehicle via a remote communication device. Interaction information 204 may be information collected by sensors such as cameras and microphones, or information received remotely via a communication device. In some implementations, interaction information 204 may include interaction information acquired at the current time t, or historical interaction information acquired prior to the current time t. Historical interaction information may include historical interaction information from all times since the start of the autonomous driving process (e.g., from time 0 to time t-1), or historical interaction information within a predetermined time period prior to the current time t (e.g., from time tk to time t, where k represents a predetermined time range).

[0058] The encoding layer 220 can be configured to encode historical decision information, perception information, traffic information and interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing historical decision information, perception information, traffic information and interaction information, respectively, wherein the first tensor, the second tensor, the third tensor and the fourth tensor have the same spatial representation.

[0059] In some embodiments, the encoding layer may include a recurrent neural network or a Transformer network, and may be configured to encode historical decision information using the recurrent neural network or the Transformer network. The historical trajectory of the vehicle may be determined based on the historical decision information, and the coordinates of the trajectory points at each historical moment in the vehicle coordinate system at the current moment t may be determined. Each trajectory coordinate in the historical trajectory may be input into the recurrent neural network or the Transformer network to obtain a one-dimensional vector or a two-dimensional tensor for representing the coordinates of the trajectory point. Furthermore, the one-dimensional vectors or two-dimensional tensors corresponding to the coordinates of the trajectory points at each historical moment may be stacked in chronological order to obtain a two-dimensional or three-dimensional vector with an added time dimension. As the first tensor used to represent historical decision information.

[0060] In some embodiments, the coding layer may further include a coding network for mapping information into a bird's-eye view BEV space, such as a BEVFormer, and configured to map the perception information into the BEV space to obtain a BEV representation of the perception information. The perception information collected by the sensor at each moment can be input into the BEVFormer network, and a BEV representation of the perception information at that moment can be obtained. In the example, the BEV representation can be a three-dimensional vector. In the case where the input information includes perception information at multiple moments, the BEV representations of the perception information at each moment can be stacked in chronological order to obtain a four-dimensional tensor with an added time dimension. As the second tensor used to represent the perceptual information.

[0061] The encoding layer can also be configured to map traffic information into BEV space to obtain a bird's-eye view BEV representation of traffic information. For example, the traffic information at each moment can be vectorized and encoded using BEVFormer to obtain the BEV representation of the traffic information at that moment. In the case where the input information includes traffic information at multiple moments, the BEV representations of the traffic information at each moment can be stacked in chronological order to obtain a four-dimensional tensor with an added time dimension. As the third tensor used to represent traffic information.

[0062] It is understandable that the BEVFormer for processing perception information and the BEVFormer for processing traffic information can be configured separately according to actual conditions.

[0063] In some embodiments, the encoding layer may further include a pre-trained language model (PLM). The pre-trained language model may be used to vectorize natural language to convert natural language information into information that can be processed by a machine. The pre-trained language model may be any model that can process input natural language information. In some examples, a large language model (LLM) may also be used to vectorize natural language. When the input interaction information includes action information represented by an image, a suitable image recognition algorithm may be used to convert the information in the image into natural language information or vectorized information that can be processed by a machine. The interaction information may be encoded using a pre-trained language model to convert the natural language information into a fourth tensor containing multiple dimensions (such as 2 dimensions). The natural language information may also include timestamp information.

[0064] Although the first tensor, the second tensor, the third tensor, and the fourth tensor are the results of output through different encoding methods, the above tensors can have the same length and width dimensions. For example, the first tensor used to represent historical decision information and the fourth tensor used to represent interaction information, although not BEV representation, can have the same spatial representation as the second tensor and the third tensor represented in the BEV space. In this way, the first tensor, the second tensor, the third tensor, and the fourth tensor can be processed uniformly in subsequent model processing, so that the model can uniformly process and reason on different input information when making inference decisions, so as to make autonomous driving decisions while considering input information of different modalities.

[0065] The autoregressive inference layer 230 is configured to perform inference on an input tensor formed of the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for the next moment.

[0066] In some embodiments, the autoregressive reasoning layer 230 can be implemented by a world model. In some implementations, the autoregressive reasoning layer can be implemented by a recurrent neural network (such as a long short-term memory network LSTM, a gated recurrent unit GRU), a Transformer, or a recursive structure with a memory mechanism (such as a recursive memory Transformer network (Recurrent Memory Transformer)) or a diffusion model. Using the autoregressive method, the input signal at time t can be used to predict the output result at time t+1. The combination of the first tensor, the second tensor, the third tensor and the fourth tensor output by the encoding layer can be used as the input of the autoregressive reasoning layer 230. For example, the first tensor, the second tensor, the third tensor and the fourth tensor can be flattened into a two-dimensional tensor sequence, and the above two-dimensional tensor sequence can be used as the input of the autoregressive reasoning layer.

[0067] Taking the recursive memory Transformer network as an example, the memory tensor M0 at the initial moment can be pre-set. Among them, the parameters in M0 can be randomly initialized parameters. The memory tensor M at time t can be t It is input into the Transformer layer together with the two-dimensional tensor sequence, so that each vector in the memory tensor is used to process the two-dimensional tensor sequence based on the attention mechanism to obtain the memory tensor M at the next moment. t+1 As the hidden state of the next moment. Among them, the memory tensor M t It is equivalent to the query parameter (Q) of the input Transformer layer, and the two-dimensional tensor sequence is equivalent to the key (K) and value (V) parameters of the input Transformer layer.

[0068] The decoding layer 240 is configured to decode based on the hidden state at the next moment to obtain interaction information for the next moment and autonomous driving decision information for the next moment.

[0069] In some embodiments, the decoding layer may include a Transformer network. The Transformer can be used to transform the hidden state M at time t+1. t+1 Decoding is performed to obtain the autonomous driving decision information at time t+1, such as the control signals of throttle, brake, steering amplitude, etc. at time t+1.

[0070] You can also use Transformer to calculate the hidden state M at time t+1 t+1 , to obtain the interactive information output at time t+1. The interactive information output at time t+1 may include a natural language description, which can be used to respond to the interactive information at time t in natural language form.

[0071] You can also use the hidden state M at time t+1 t+1 Generate future prediction information. The future prediction information may include a future prediction image at time t+1, indicating the obstacle position at time t+1 or perception information at future times. The hidden state M that can be represented by BEV t+1 Perform spatial mapping to convert it to the sensor (such as camera) coordinate system. Image diffusion and deconvolution can be further used to transform M in the sensor coordinate system. t+1 The image is processed to obtain the future prediction image at time t+1. The future prediction image can be used to train the autonomous driving model in a self-supervised manner, so that the autonomous driving model has accurate future prediction capabilities, thereby improving the accuracy of the decision information output by the autonomous driving model.

[0072] According to one aspect of the present disclosure, a method for autonomous driving is also provided.

[0073] Figure 3A shows an exemplary flowchart of an autonomous driving method 300 according to an embodiment of the present disclosure. Autonomous driving method 300 can be implemented using autonomous driving model 200 described in conjunction with Figure 2 . The advantages of autonomous driving model 200 described in conjunction with Figure 2 also apply to autonomous driving method 300 and are not further elaborated here.

[0074] In step S302, input information may be obtained, wherein the input information includes historical decision information, perception information, traffic information, and interaction information at the current moment.

[0075] In step S304, historical decision information, perception information, traffic information and interaction information may be encoded to obtain a first tensor, a second tensor, a third tensor and a fourth tensor respectively used to represent historical decision information, perception information, traffic information and interaction information.

[0076] In step S306 , autoregressive inference may be performed on the input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a first moment after the current moment.

[0077] In step S308, decoding can be performed based on the hidden state at the first moment to obtain interaction information for the first moment and autonomous driving decision information for the first moment.

[0078] The interaction information at the first moment may include a signal indicating that the autonomous driving process requires assistance.

[0079] In some cases, the autonomous driving decision information output by the autonomous driving model is invalid. For example, the autonomous driving decision information output by the autonomous driving model may prevent the autonomous vehicle from continuing along the predetermined navigation route. In such cases, the interactive information output by the autonomous driving method 300 may include a signal indicating that assistance is required during the autonomous driving process. This signal may be a request for human or other model intervention in the autonomous driving process, and the signal may be in natural language, so that the request signal accurately and completely describes the current situation of the autonomous driving vehicle and is easy to understand.

[0080] In some embodiments, the autonomous driving method 300 may further include sending the interaction information (i.e., assistance request) at the first moment and the current driving state of the autonomous driving vehicle to a remote server, and receiving response information from the remote server, wherein the response information may include prompt interaction information for prompting the autonomous driving process. For example, a signal indicating that the autonomous driving process requires assistance, the perception information at the first moment, and the natural language description of the current driving scene of the autonomous driving vehicle at the first moment may be sent to the remote server. Using this method, a solution can be obtained from the remote server when a problem occurs with the autonomous driving model. The remote server may be a cloud server.

[0081] In some implementations, the current driving state of the autonomous vehicle may include a natural language description and / or image of the driving scene at a first moment. The natural language description of the current driving scene may be output by an autonomous driving model deployed on the autonomous vehicle. A large language model may be deployed on a remote server. The large language model may be used to process the current driving state of the autonomous vehicle to obtain response information for the assistance request. The large language model may provide at least one response information based on the current driving state of the autonomous vehicle to assist the autonomous driving model. The at least one response information may be evaluated, and the response information with the best evaluation result may be sent to the autonomous vehicle. In other implementations, a cloud-based autonomous driving model with a larger number of parameters may be deployed on the remote server. The larger number of parameters here refers to the number of parameters deployed on the autonomous driving vehicle. Thus, the cloud-based autonomous driving model can handle more complex autonomous driving processes. The cloud-based autonomous driving model may be capable of processing multimodal input information (e.g., images, videos, and language). The current driving state of the autonomous vehicle may include sensor input collected by the autonomous vehicle's sensors at the second moment, traffic information at the second moment, and other information. The autonomous vehicle's current driving state and / or response information can be input into a cloud-based autonomous driving model with a larger set of parameters to obtain response information. In some examples, the response information can be natural language information rather than control signals used to directly control the autonomous vehicle. For example, the natural language prompt interaction information in the response information can include instructions for guiding vehicle driving. Depending on the actual situation, the response information can also include any other form of information that can be processed by the autonomous driving model to guide vehicle driving. In this way, the autonomous driving vehicle can obtain remote driving assistance for the current driving scenario. By inputting the driving instructions in the received response information along with historical decision information, perception information, and traffic information at the current moment into the autonomous driving model, the autonomous driving model can generate autonomous driving decisions based on the instructions given by the remote server, thereby enabling the autonomous vehicle to autonomously escape from a difficult situation. In this process, the autonomous driving vehicle can rely on the intelligence of the autonomous driving model to escape from a difficult situation without the need for human control of the vehicle.

[0082] In some embodiments, the autonomous driving method 300 may further include encoding the perception information and response information obtained at the second moment after receiving the response information to obtain a fifth tensor and a sixth tensor for representing the perception information and response information at the second moment; performing autoregressive inference on the input tensor formed by the fifth tensor and the sixth tensor to obtain a hidden state at a third moment after the second moment, and decoding based on the hidden state at the third moment to obtain autonomous driving decision information for the third moment.

[0083] It can be understood that when using the input information at the second moment to infer the autonomous driving decision information at the third moment, the autonomous driving model can also simultaneously receive the historical decision information at the second moment and the traffic information at the second moment, and use the autonomous driving model described in combination with Figure 2 to obtain the autonomous driving decision information at the third moment.

[0084] Different from directly sending control signals from a remote location to take over the autonomous driving process, using the implementation method of the present disclosure, the autonomous driving vehicle uses the response information fed back remotely as part of the interactive information received by the autonomous driving model for reasoning on autonomous driving decisions. In this way, the response information output by the model located on the remote server does not need to be directly applied to the autonomous driving vehicle, so the model deployed at the remote server can be a model with general problem-solving capabilities. By using a model with general problem-solving capabilities, more comprehensive reference information can be considered when generating response information. The autonomous driving model uses the response information as part of the interactive information input during model reasoning, and can generate appropriate autonomous driving decision information based on the current actual driving situation, the historical trajectory of the current autonomous driving process, and the content of the response information.

[0085] In some embodiments, the historical decision information includes autonomous driving decision information output by the decoding layer at at least one previous moment before the current moment.

[0086] In some embodiments, the perception information includes sensor input collected by at least one sensor mounted on the autonomous vehicle.

[0087] In some embodiments, the traffic information includes at least one of speed limit information, map information, and navigation information of the current travel route.

[0088] In some embodiments, the interactive information includes at least one of traffic control information, interactive information from passengers, and interactive information from security officers.

[0089] In some embodiments, a recurrent neural network or a Transformer network may be used to encode historical decision information.

[0090] In some embodiments, the perceptual information may be mapped to a bird's-eye view (BEV) space to obtain a bird's-eye view (BEV) representation of the perceptual information.

[0091] In some embodiments, traffic information may be mapped to a bird's-eye view BEV space to obtain a bird's-eye view BEV representation of the traffic information.

[0092] In some embodiments, the interaction information may be encoded using a pre-trained language model.

[0093] In some embodiments, the autoregressive inference layer can be implemented by a recurrent neural network (such as a long short-term memory network LSTM, a gated recurrent unit GRU), a Transformer, or a recursive structure with a memory mechanism (such as a recurrent memory Transformer network (Recurrent Memory Transformer)) or a diffusion model.

[0094] By using the autonomous driving method provided by the embodiments of the present disclosure, the prompt information of natural language interaction can be used to guide the autonomous driving model to control the vehicle, and realize the rapid generation of escape plans for the autonomous driving vehicle's escape, intervention and other control. The autonomous driving model can actively initiate an interaction request to request remote assistance when appropriate. Cloud guidance for a certain number (tens of autonomous driving vehicles) of autonomous driving vehicles can be achieved by using a remotely deployed model with general capabilities. Thus, users of autonomous driving vehicles only need to enable the autonomous driving model on the vehicle side. Even if difficulties arise during the autonomous driving process, the autonomous driving model can autonomously generate an assistance signal and send it to the remote server to obtain an escape plan without the need for manual control of the vehicle to resolve difficulties during driving.

[0095] FIG3B shows an exemplary process for assisting autonomous driving according to an embodiment of the present disclosure.

[0096] As shown in Figure 3B, in block 301, the autonomous driving model of the autonomous vehicle encounters a problem and outputs an assistance request. In block 303, the cloud-deployed model receives the assistance request and / or the current autonomous driving state of the autonomous vehicle. In block 305, three escape plan instructions in natural language are generated based on the assistance request and / or the current autonomous driving state of the autonomous vehicle. Escape Plan 1 includes: 1) backing up one meter while ensuring safety; 2) attempting a left turn into the adjacent lane and quickly exiting if there are no oncoming vehicles; and 3) returning to the lane. Similarly, Escape Plans 2 and 3 include escape instructions with different contents in natural language. A response message may include one of Escape Plans 1, 2, and 3 (e.g., Escape Plan 1) selected from Escape Plan 1, 2, and 3 and sent to the autonomous vehicle. In block 307, the autonomous vehicle utilizes the autonomous driving model described in conjunction with Figure 2, inputs the response message as interactive information into the autonomous driving model, and generates subsequent autonomous driving decisions to achieve autonomous escape.

[0097] According to another aspect of the present disclosure, an automatic driving device based on an automatic driving model is provided.

[0098] FIG4 illustrates a block diagram of an autonomous driving device 400 according to an embodiment of the present disclosure. As shown in FIG4 , autonomous driving device 400 includes an acquisition unit 410, an encoding unit 420, an inference unit 430, and a decoding unit 440. Autonomous driving device 400 can be implemented based on autonomous driving model 200 described in conjunction with FIG2 .

[0099] The acquisition unit 410 may be configured to acquire input information, where the input information includes historical decision information, perception information, traffic information, and interaction information at the current moment.

[0100] The encoding unit 420 can be configured to encode historical decision information, perception information, traffic information and interaction information to obtain a first tensor, a second tensor, a third tensor and a fourth tensor for representing historical decision information, perception information, traffic information and interaction information respectively.

[0101] The inference unit 430 may be configured to perform inference on an input tensor formed of the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a first time step after the current time step.

[0102] The decoding unit 440 may be configured to perform decoding based on the hidden state at the first moment to obtain interaction information for the first moment and autonomous driving decision information for the first moment, wherein the interaction information at the first moment includes a signal indicating that assistance is required for the autonomous driving process.

[0103] In some embodiments, the autonomous driving device 400 may further include a communication unit configured to: transmit the interaction information at the first moment and the current driving state of the autonomous driving vehicle to a remote server; and receive response information from the remote server. For example, a signal indicating that assistance is required for the autonomous driving process, the perception information at the first moment, and a natural language description of the current driving scene of the autonomous driving vehicle at the first moment may be transmitted to the remote server.

[0104] In some embodiments, the encoding unit is further configured to, at a second moment after receiving the response information, encode the perception information acquired at the second moment and the response information to obtain a fifth tensor and a sixth tensor for representing the perception information and the response information at the second moment. The inference unit is further configured to perform autoregressive inference on the input tensor formed by the fifth tensor and the sixth tensor to obtain a hidden state at a third moment after the second moment. The decoding unit is further configured to decode based on the hidden state at the third moment to obtain autonomous driving decision information for the third moment.

[0105] In some embodiments, the response information is natural language information.

[0106] In some embodiments, the response information is the result of processing the current driving state of the autonomous vehicle using a large language model.

[0107] It should be understood that the various modules or units of the apparatus 400 shown in FIG4 may correspond to the various steps in the method 300 described with reference to FIG3A . Thus, the operations, features, and advantages described above for the method 300 are also applicable to the apparatus 400 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.

[0108] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various units discussed herein may be separated into multiple units, and / or at least some functionality of multiple units may be combined into a single unit.

[0109] It should also be understood that various technologies can be described herein in the general context of software hardware elements or program modules. The various units described above with respect to Figure 4 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of units 410 to 440 can be implemented together in a system on chip (SoC). SoC can include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and can optionally execute the received program code and / or include embedded firmware to perform a function.

[0110] According to another aspect of the present disclosure, an electronic device is also provided, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the autonomous driving method according to an embodiment of the present disclosure.

[0111] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, where the computer instructions are used to enable the computer to execute the automatic driving method according to an embodiment of the present disclosure.

[0112] According to another aspect of the present disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the autonomous driving method according to an embodiment of the present disclosure when executed by a processor.

[0113] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising an autonomous driving device according to an embodiment of the present disclosure and one of the above-mentioned electronic devices.

[0114] With reference to Figure 5, a block diagram of an electronic device 500 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0115] As shown in Figure 5, electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In RAM 503, various programs and data required for the operation of electronic device 500 can also be stored. Computing unit 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to bus 504.

[0116] Multiple components within electronic device 500 are connected to I / O interface 505, including an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. Input unit 506 can be any type of device capable of inputting information into electronic device 500. Input unit 506 can receive input numeric or character information and generate key signal input related to user settings and / or function control of the electronic device. It may include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 508 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. It may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0117] The computing unit 501 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the method (or process) 300. For example, in some embodiments, the method (or process) 300 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method (or process) 300 described above can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the method (or process) 300 in any other appropriate manner (eg, by means of firmware).

[0118] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0119] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0120] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0122] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0123] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0124] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0125] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.

Claims

1. An autonomous driving method for autonomous escape, comprising: Acquiring input information, wherein the input information includes historical decision information, perception information, traffic information, and interaction information at a current moment; Encoding the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor respectively used to represent the historical decision information, the perception information, the traffic information, and the interaction information; Performing autoregressive inference on an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a first moment after the current moment; as well as Decoding is performed based on the hidden state at the first moment to obtain interaction information for the first moment and autonomous driving decision information for the first moment; The interaction information at the first moment includes a signal indicating that the autonomous driving process requires assistance.

2. The automatic driving method according to claim 1, further comprising: Sending the signal indicating that the autonomous driving process requires assistance, the perception information at the first moment, and the natural language description information of the current driving scene of the autonomous driving vehicle at the first moment to a remote server; and A response message is received from the remote server.

3. The automatic driving method according to claim 2, further comprising: At a second moment after receiving the response information, encoding the perception information and the response information acquired at the second moment to obtain a fifth tensor and a sixth tensor for representing the perception information and the response information at the second moment; Performing autoregressive inference on an input tensor formed by the fifth tensor and the sixth tensor to obtain a hidden state at a third moment after the second moment; Decoding is performed based on the hidden state at the third moment to obtain the autonomous driving decision information for the third moment.

4. The autonomous driving method of claim 2, wherein the response information is natural language information.

5. The automatic driving method according to claim 4, wherein: The response information is a result of processing the current driving state of the autonomous driving vehicle using a large language model.

6. The automatic driving method according to any one of claims 1 to 5, wherein: Encoding the historical decision information includes: encoding the historical decision information using a recurrent neural network or a Transformer network.

7. The automatic driving method according to claim 6, wherein: The historical decision information includes autonomous driving decision information output by the decoding layer at least one previous moment before the current moment.

8. The automatic driving method according to any one of claims 1 to 7, wherein: Encoding the perception information includes: mapping the perception information to a bird's eye view (BEV) space to obtain a bird's eye view (BEV) representation of the perception information.

9. The autonomous driving method of claim 8, wherein the perception information comprises sensor input collected by at least one sensor mounted on the autonomous driving vehicle.

10. The automatic driving method according to any one of claims 1 to 9, wherein: Encoding the traffic information includes mapping the traffic information to a bird's-eye view BEV space to obtain a bird's-eye view BEV representation of the traffic information.

11. The automatic driving method according to claim 10, wherein: The traffic information includes at least one of speed limit information, map information and navigation information of a current travel route.

12. The automatic driving method according to any one of claims 1 to 11, wherein: Encoding the interaction information includes: encoding the interaction information using a pre-trained language model.

13. The automatic driving method according to claim 12, wherein: The interactive information includes at least one of traffic control information, interactive information from passengers and interactive information from security personnel.

14. The automatic driving method according to any one of claims 1 to 13, wherein: Performing autoregressive reasoning on the input tensor includes: using one of a recurrent neural network, a Transformer, a recursive memory Transformer, and a diffusion model to implement autoregressive reasoning on the input tensor.

15. An automatic driving device, comprising: an acquisition unit configured to acquire input information, wherein the input information includes historical decision information, perception information, traffic information, and interaction information at a current moment; an encoding unit configured to encode the historical decision information, the perception information, the traffic information, and the interaction information to obtain a first tensor, a second tensor, a third tensor, and a fourth tensor respectively used to represent the historical decision information, the perception information, the traffic information, and the interaction information; an inference unit configured to infer an input tensor formed by the first tensor, the second tensor, the third tensor, and the fourth tensor to obtain a hidden state for a first moment after a current moment; as well as a decoding unit configured to perform decoding based on the hidden state at the first moment to obtain interaction information for the next moment and automatic driving decision information for the first moment, The interaction information at the first moment includes a signal indicating that the autonomous driving process requires assistance.

16. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 14.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-14.

18. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 1 to 14 when executed by a processor.

19. An autonomous driving vehicle comprising: One of the automatic driving device according to claim 15 and the electronic device according to claim 16.

Citation Information

Patent Citations

  • Automatic driving method and device, electronic equipment and storage medium

    CN114194211A

  • Content tendency evaluation and prediction method based on adaptive context inference mechanism

    CN115563989A

  • Time sequence autoregression simultaneous decision-making and prediction automatic driving model and training method thereof

    CN116859724A

  • Automatic driving model capable of performing natural language interaction and training method thereof

    CN117010265A

  • Automatic driving method and device capable of achieving autonomous escape by following instructions and vehicle

    CN117539253A