Autonomous driving model based on temporal recursive autoregressive inference, and method and vehicle

By adopting the timing recursive autoregressive reasoning method in the autonomous driving model, combined with the functions of the coding layer, trajectory planning layer and inference layer, the problem of difficulty in combining historical information and future prediction in the existing technology is solved, and more accurate and reliable autonomous driving decisions are achieved.

WO2025118534A1PCT designated stage expired Publication Date: 2025-06-12BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/099428
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-07
Filing Date
2024-06-14
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

The existing autonomous driving technology is difficult to effectively combine historical information and future predictions, resulting in a lack of accuracy and reliability of models in driving decisions.

Method used

The autonomous driving model based on timing recursive autoregressive reasoning is adopted, and the sensor information is encoded through the encoding layer. The trajectory planning layer determines the driving trajectory based on the historical scenario, and the inference layer combines current and historical information to make future predictions and decisions.

Benefits of technology

The model is realized to make future predictions while learning history, improve the prediction effect and driving ability of the autonomous driving model, and enhance the accuracy and reliability of decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024099428_12062025_PF_FP_ABST
    Figure CN2024099428_12062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, and particularly relates to the technical fields of autonomous driving and artificial intelligence. Provided are an autonomous driving model based on temporal recursive autoregressive inference, and a method, an apparatus and a vehicle. In the autonomous driving model, a coding layer is configured to code sensor information of a current moment, so as to obtain a current scenario representation; a trajectory planning layer is configured to determine a driving trajectory from the current moment to a future moment on the basis of a historical scenario representation of the current moment; and an inference layer is configured to determine a predicted scenario representation of at least one future moment and a historical scenario representation of the future moment on the basis of the current scenario representation, the historical scenario representation and current prompt information, wherein the prompt information at least comprises the driving trajectory. Therefore, the autonomous driving model can use an integrated inference layer to realize the learning of historical information and the prediction of future information, so that the model can perform future prediction while learning the history, and thus the prediction effect of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Autonomous driving model, method and vehicle based on temporal recursive autoregressive inference

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to Chinese patent application No. 202311676405.2 filed on December 7, 2023, the entire contents of which are incorporated herein by reference in their entirety. Technical Field

[0003] The present disclosure relates to the field of computer technology, in particular to the field of autonomous driving and artificial intelligence technology, and specifically to an autonomous driving model, a training method for an autonomous driving model, an autonomous driving method implemented using an autonomous driving model, an autonomous driving device based on the autonomous driving model, a training device for the autonomous driving model, an electronic device, a computer-readable storage medium, a computer program product, and an autonomous driving vehicle. Background Art

[0004] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0005] Autonomous driving technology integrates technologies such as recognition, decision-making, positioning, communication security, and human-computer interaction. Artificial intelligence learning can assist in generating autonomous driving strategies.

[0006] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.

[0007] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art.

[0008] Summary of the Invention

[0009] The present disclosure provides an autonomous driving model, method, device, and vehicle based on temporal recursive autoregressive reasoning.

[0010] According to one aspect of the present disclosure, an autonomous driving model is provided, comprising an encoding layer, a trajectory planning layer, and an inference layer, wherein the encoding layer is configured to encode sensor information at a current moment to obtain a current scene representation; the trajectory planning layer is configured to determine a driving trajectory from the current moment to a future moment based on a historical scene representation at the current moment, wherein the historical scene representation at the current moment indicates a scene at at least one previous moment of the current moment; and the inference layer is configured to determine a predicted scene representation at at least one future moment and a historical scene representation at the future moment based on the current scene representation, the historical scene representation, and current prompt information, wherein the prompt information at least includes the driving trajectory.

[0011] According to another aspect of the present disclosure, a training method for an autonomous driving model is provided, comprising: obtaining sample sensor information within a sample period and an actual driving trajectory corresponding to the sample sensor information; encoding the sample sensor information at a sample moment using an encoding layer of the autonomous driving model to obtain a sample scene representation; determining a predicted driving trajectory from the current moment to a future moment based on a historical scene representation at the sample moment using a trajectory planning layer of the autonomous driving model, wherein the historical scene representation at the sample moment indicates a scene at at least one previous moment of the sample moment; determining a predicted scene representation at at least one future moment and a historical scene representation at the future moment based on the sample scene representation, the historical scene representation, and sample prompt information using an inference layer of the autonomous driving model, wherein the prompt information includes at least the predicted driving trajectory from the current moment to the future moment; and adjusting parameters of the autonomous driving model based on the difference between the predicted driving trajectory from the current moment to the future moment and the actual driving trajectory from the current moment to the future moment.

[0012] According to another aspect of the present disclosure, there is provided an autonomous driving method implemented using an autonomous driving model, comprising: encoding sensor information at a current moment to obtain a current scene representation; determining a driving trajectory from the current moment to a future moment based on the historical scene representation of the current moment, wherein the historical scene representation of the current moment indicates a scene of at least one previous moment of the current moment; and determining a predicted scene representation of at least one future moment and a historical scene representation of the future moment based on the current scene representation, the historical scene representation, and current prompt information, wherein the prompt information includes at least the driving trajectory.

[0013] According to another aspect of the present disclosure, an autonomous driving device based on an autonomous driving model is provided, comprising: an encoding unit configured to encode sensor information at a current moment to obtain a current scene representation; a trajectory determination unit configured to determine a driving trajectory from the current moment to a future moment based on a historical scene representation at the current moment, wherein the historical scene representation at the current moment indicates a scene at at least one previous moment of the current moment; and a future prediction unit configured to determine a predicted scene representation at at least one future moment and a historical scene representation at the future moment based on the current scene representation, the historical scene representation, and current prompt information, wherein the prompt information includes at least the driving trajectory.

[0014] According to another aspect of the present disclosure, a training device for an autonomous driving model is provided, comprising: a sample acquisition unit configured to acquire sample sensor information within a sample period and a real driving trajectory corresponding to the sample sensor information; an encoding unit configured to encode the sample sensor information at a sample moment using an encoding layer of the autonomous driving model to obtain a sample scene representation; a trajectory prediction unit configured to determine a predicted driving trajectory from the current moment to a future moment based on a historical scene representation at the sample moment using a trajectory planning layer of the autonomous driving model, wherein the historical scene representation at the sample moment indicates a scene at at least one previous moment of the sample moment; an inference unit configured to determine a predicted scene representation at at least one future moment and a historical scene representation at the future moment based on the sample scene representation, the historical scene representation, and sample prompt information using the inference layer of the autonomous driving model, wherein the prompt information includes at least the predicted driving trajectory from the current moment to the future moment; and a parameter adjustment unit configured to adjust the parameters of the autonomous driving model based on the difference between the predicted driving trajectory from the current moment to the future moment and the real driving trajectory from the current moment to the future moment.

[0015] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can perform the above method.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.

[0017] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0018] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising: an autonomous driving device according to an embodiment of the present disclosure, or one of an electronic device.

[0019] Utilizing the embodiments of the present disclosure, an integrated reasoning layer can be used to learn historical information and predict future information, thereby enabling the model to make future predictions while learning history, improving the model's prediction effectiveness. Utilizing the autonomous driving model provided by the present disclosure, temporal autoregressive reasoning can be implemented by inferring the prediction results at time t+1 using information at time t. When the reasoning layer adopts a recursive structure, the autonomous driving model provided by the embodiments of the present disclosure can implement temporal recursive autoregressive reasoning and determine autonomous driving decisions based on the reasoning results.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0022] FIG1 shows a schematic diagram of an exemplary system in which various methods described herein may be implemented according to an embodiment of the present disclosure;

[0023] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure;

[0024] FIG3 shows an exemplary flowchart of a training method for an autonomous driving model according to an embodiment of the present disclosure;

[0025] FIG4 shows an exemplary flowchart of an autonomous driving method implemented by using an autonomous driving model according to an embodiment of the present disclosure;

[0026] FIG5 shows an example of a reasoning process at time t according to an embodiment of the present disclosure;

[0027] FIG6 shows a structural block diagram of an automatic driving device according to an embodiment of the present disclosure;

[0028] FIG7 shows a structural block diagram of an autonomous driving training device according to an embodiment of the present disclosure; and

[0029] FIG8 shows a block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION

[0030] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0031] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.

[0032] The terms used in the descriptions of the various examples described in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in this disclosure encompasses any one and all possible combinations of the listed items.

[0033] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0034] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0035] FIG1 shows a schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein may be implemented according to an embodiment of the present disclosure. Referring to FIG1 , the system 100 includes a motor vehicle 110, a server 120, and one or more communication networks 130 coupling the motor vehicle 110 to the server 120.

[0036] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to perform a method according to an embodiment of the present disclosure.

[0037] The server 120 may run one or more services or software applications that enable autonomous driving. In some embodiments, the server 120 may also provide other services or software applications, which may include non-virtual environments and virtual environments. In the configuration shown in Figure 1, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. The user of the motor vehicle 110 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may be different from the system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0038] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0039] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0040] In some embodiments, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from motor vehicle 110. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.

[0041] The network 130 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 130 may be a satellite communication network, a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including, for example, Bluetooth, WiFi), and / or any combination of these and other networks.

[0042] The system 100 may also include one or more databases 150. In some embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repository 150 can reside in a variety of locations. For example, the data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repository 150 can be of different types. In some embodiments, the data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0043] In some embodiments, one or more of the databases 150 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0044] Motor vehicle 110 may include sensors 111 for sensing its surroundings. Sensors 111 may include one or more of the following: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radar, and laser radar (LiDAR). Different sensors offer different detection accuracy and range. Cameras may be mounted on the front, rear, or other locations of the vehicle. Visual cameras can capture real-time information about the vehicle's interior and exterior and present it to the driver and / or passengers. Furthermore, by analyzing the images captured by the visual cameras, information such as traffic light indications, intersection conditions, and the operating status of other vehicles can be obtained. Infrared cameras can detect objects in night vision conditions. Ultrasonic sensors can be mounted on all sides of the vehicle, utilizing the strong directionality of ultrasonic waves to measure the distance of external objects from the vehicle. Millimeter-wave radars can be mounted on the front, rear, or other locations of the vehicle, utilizing the properties of electromagnetic waves to measure the distance of external objects from the vehicle. LiDARs can be mounted on the front, rear, or other locations of the vehicle, detecting object edges and shapes for object recognition and tracking. Due to the Doppler effect, radar devices can also measure changes in the speed of the vehicle and moving objects.

[0045] The motor vehicle 110 may also include a communication device 112. The communication device 112 may include a satellite positioning module that can receive satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generate coordinates based on these signals. The communication device 112 may also include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as GSM / GPRS, CDMA, LTE, and other current or evolving wireless communication technologies (e.g., 5G technology). The communication device 112 may also have a vehicle-to-everything (V2X) module that is configured to implement vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may also include a module configured to communicate with a user terminal 145 (including but not limited to a smartphone, tablet computer, or wearable device such as a watch) via a wireless local area network or Bluetooth using the IEEE 802.11 standard, for example. Using the communication device 112, the motor vehicle 110 may also access the server 120 via the network 130.

[0046] The motor vehicle 110 may also include a control device 113. The control device 113 may include a processor that communicates with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, and braking system of the motor vehicle 110 (not shown) via multiple actuators in response to input from multiple sensors 111 or other input devices to control acceleration, steering, and braking, respectively, without human intervention or limited human intervention. Some processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an on-board processor, while other processing may be performed using computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.

[0047] The system 100 of FIG. 1 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with this disclosure.

[0048] End-to-end model-based autonomous driving solutions have the advantages of not requiring the definition of intermediate representations and being able to continuously improve model performance through data-driven approaches. To improve the performance of autonomous driving models, this disclosure provides a new autonomous driving model.

[0049] FIG2 shows an exemplary block diagram of an autonomous driving model according to an embodiment of the present disclosure.

[0050] As shown in FIG2 , the autonomous driving model 200 includes a coding layer 210 , a trajectory planning layer 220 , and an inference layer 230 .

[0051] The encoding layer 210 is configured to encode the sensor information at the current moment to obtain a current scene representation.

[0052] The trajectory planning layer 220 is configured to determine a driving trajectory from a current moment to a future moment according to a historical scene representation at the current moment, wherein the historical scene representation at the current moment indicates a scene at at least one previous moment of the current moment.

[0053] The reasoning layer 230 is configured to determine at least one predicted scene representation at a future moment and a historical scene representation at a future moment based on the current scene representation, the historical scene representation, and the current prompt information, wherein the prompt information includes at least a driving trajectory.

[0054] By using the autonomous driving model provided by the embodiments of the present disclosure, an integrated reasoning layer can be used to realize the learning of historical information and the prediction of future information, so that the model can make future predictions while learning history, thereby improving the prediction effect of the model.

[0055] The principles of the present disclosure will be described in detail below.

[0056] The encoding layer 210 may be configured to encode sensor information at a current moment to obtain a current scene representation.

[0057] The sensor information may include sensor input collected by at least one sensor installed on the autonomous vehicle. For example, the sensor information may include at least one of the following: perception information from one or more cameras, perception information from one or more lidars, and perception information from one or more millimeter-wave radars.

[0058] The sensor information at the current moment may be the raw sensor information collected by the sensor at the current moment. The encoding layer may encode the raw sensor information collected by the sensor to obtain an implicit representation of the sensor information indicating the current scene as the current scene representation. For example, the sensor information may be mapped into an implicit representation of the bird's-eye view BEV space using a structure such as BEVFormer or BEVFusion, such as N w ×N H ×D-dimensional tensor. w ,N H Denote the length and width of the BEV space, and D is the dimension of the tensor. It should be understood that mapping sensor information into an implicit representation of the BEV space is merely an exemplary illustration of the present disclosure. Without departing from the principles of the present disclosure, any suitable method can be used to map raw sensor information in three-dimensional space into implicit representations of other spaces.

[0059] For each moment t, the current scene representation at that moment is based on the raw sensor information (RAW) collected by the sensor at moment t. t ) is encoded into an implicit representation (H t ).

[0060] The trajectory planning layer 220 may be configured to determine a driving trajectory from the current moment to a future moment based on a historical scene representation of the current moment, wherein the historical scene representation of the current moment indicates a scene at at least one previous moment of the current moment.

[0061] For the current time t, the historical scene representation M[:t-1] may include information about scenes in a historical period before time t. In some examples, the historical scene representation M[:t-1] may include information about scenes from an initial time 0 to time t-1.

[0062] In some embodiments, the current historical scene representation can be predicted based on the current scene representation and historical scene representations at previous moments, rather than encoded from sensor information collected at previous moments. This approach allows the autonomous driving model to learn and utilize scene information from previous moments.

[0063] The historical scene representation M[:0] at time t=1 can be initialized to a default value (e.g., 0). At time t=1, the sensor information RAW1 of the current moment can be obtained, and RAW1 can be encoded to obtain the scene representation H1 at time t=1. The historical scene representation M[:1] at the next moment (t=2) can be predicted based on the initial historical scene representation M[:0] and H1. For example, the reasoning layer of the autonomous driving model can be used to predict M[:0], H1, and the prompt information P1 including the driving trajectory at time t=1 to obtain the historical scene representation M[:1]. Similarly, the historical scene representation M[:t-1] at time t can be obtained through step-by-step reasoning.

[0064] The trajectory planning layer can be used to determine information for the current moment's driving action based on the historical scene representation at the current moment, and further determine the vehicle's trajectory after the current moment based on the predicted action information. For example, the trajectory planning layer can be configured to determine the current moment's driving decision action based on the historical scene representation at the current moment, and then determine a first driving trajectory from the current moment to a first moment based on the current moment's driving decision action. The first moment is the moment after the current moment. Based on the first driving trajectory, a prediction can be made for future moments after the first moment. The driving decision action can include the vehicle's steering wheel signal, throttle signal, etc.

[0065] In some examples, the historical scene information M[:t-1] at time t can be processed using, for example, a Transformer network to obtain the action information (A t Then, based on the principle of forward dynamics, the predicted motion information, and the current vehicle state (such as the vehicle's direction, speed, acceleration, current position, etc.), a trajectory after time t can be obtained, such as the trajectory S[t:t+1] from time t to time t+1.

[0066] The reasoning layer 230 is configured to determine at least one predicted scene representation at a future moment and a historical scene representation at a future moment based on the current scene representation, the historical scene representation, and the current prompt information, wherein the prompt information includes at least a driving trajectory.

[0067] The reasoning layer 230 can perform reasoning based on the information at the current moment t (for example, the scene representation at the current moment, the historical scene representation, and the current prompt information) to obtain the predicted scene representation at the next moment (t+1 moment) and the predicted results of the historical scene representation.

[0068] In some embodiments, the reasoning layer may determine a predicted scene representation at a first moment and a historical scene representation at the first moment based on the current scene representation, the historical scene representation, and a first driving trajectory from the current moment to the first moment.

[0069] Because the vehicle's position and posture vary at different times, the specific content of the scene representation at each moment is related to the coordinate system at that moment. Therefore, the scene representation and historical scene representation at time t should be represented in the coordinate system of time t, and the scene representation and historical scene representation at time t+1 should be represented in the coordinate system of time t+1. When predicting the scene representation and historical scene representation at the next moment, the inference layer can determine the coordinate system mapping relationship between the current and next moments based on the changes in the position information between the current and next moments.

[0070] In some embodiments, the reasoning layer can be configured to represent H based on the current scene t The predicted scene representation at the first moment in the coordinate system of the current moment and the historical scene representation at the first moment in the coordinate system of the current moment are determined by using the historical scene representation M[:t-1], and then the predicted scene representation at the first moment in the coordinate system of the current moment and the historical scene representation at the first moment in the coordinate system of the current moment are spatially transformed according to the first driving trajectory to obtain the predicted scene representation at the first moment in the coordinate system of the first moment and the historical scene representation M at the first moment in the coordinate system of the first moment pred In some examples, the vehicle position information at time t and time t+1 can be used to perform rotation and translation using a warp operation to achieve a transformation from the coordinate system at time t to the coordinate system at time t+1. It is understood that any other suitable method can also be used to achieve the transformation of the coordinate system.

[0071] Using the inference results for time t+1 obtained based on the information at time t, the scene representation at time t+2 can be predicted.

[0072] In some embodiments, the reasoning layer is further configured to determine the predicted scene representation at the second moment and the historical scene representation at the second moment based on the predicted scene representation at the first moment, the historical scene representation at the first moment, and the second driving trajectory from the first moment to the second moment, wherein the second moment is a future moment of the first moment. The above method can be used to continuously use the prediction results output by the reasoning layer to reason forward (for example, K steps) to predict the scene within a certain time period in the future (for example, to time t+K). The above method can be used to implement an autonomous driving model based on temporal reasoning. Based on the information at time t, the scene representation from time t+1 to time t+K can be obtained. Predicted historical scenario representation M pred [:t]……M pred [:t+K-1]. Using the trajectory planning layer in the autonomous driving model, M can be expressed based on the predicted historical scenarios. pred [:t]……M pred [:t+K-1] Determine the predicted driving decision action at each moment sequence.

[0073] In some implementations, the trajectory planning layer can be configured to determine a driving decision action at the first moment based on a historical scene representation at the first moment, and to determine a second driving trajectory based on the driving decision action at the first moment. As previously described, the historical scene representation at the first moment can be generated during the inference process at time t based on the current scene representation at time t.

[0074] In other implementations, a time-delayed reasoning process can be used. A driving decision action at a first moment can be determined based on a predicted historical scene representation at the first moment, where the predicted historical scene representation is predicted based on a scene representation at a moment before the current moment and historical scene representations at the moment before the current moment. A second driving trajectory from the first moment to the second moment can be determined based on the driving decision action at the first moment.

[0075] In this case, the current scene representation (such as H t-1 ) to determine the driving decision action at time t+1. Taking the time delay as Δ as an example, in the reasoning process at each moment, only the predicted driving decision action after step Δ+1 can take effect. That is, for the reasoning process at time t, only The action (and its corresponding trajectory) will be executed. In the previous Δ step reasoning process, the predicted historical scene representation M can be output for the t to t+Δ time using the reasoning process of the previous moment pred(-1) [:t]……M pred(-1)[:t+Δ] corresponds to the predicted driving decision action, i.e. The corresponding predicted trajectory is The superscript (-1) indicates that this is the result output in the reasoning step at time t-1.

[0076] Therefore, in the inference process at time t taking into account the time delay, the trajectory actually predicted by the model is

[0077] The second moment can be the next moment after the first moment, that is, moment t+2. In some examples, the time step used for reasoning can be fixed, that is, the time step between moment t and moment t+1 is the same as the time step between moment t+1 and moment t+2. In some examples, the time step used for reasoning can be dynamic, that is, the time step between moment t and moment t+1 is different from the time step between moment t+1 and moment t+2. Therefore, the time step between the second moment and the first moment is different from the time step between the first moment and the current moment. Those skilled in the art can set the time step used in model reasoning according to actual conditions.

[0078] In some embodiments, the inference layer may include at least one deformable cross attention layer. For example, the inference layer may include a recursive structure formed by connecting multiple deformable cross attention layers in series. The current scene representation, the historical scene representation at the current moment, and the prompt information at the current moment may be used as input. After processing the structure of the multiple deformable cross attention layers in series, the predicted scene representation at the future moment in the coordinate system of the current moment and the historical scene representation at the future moment in the coordinate system of the current moment are obtained. The multiple deformable cross attention layers may have different parameters.

[0079] In some implementations, the prompt information may further include at least one of the following information: map information, current vehicle driving status information, navigation information, destination information, and any instruction information related to the driving process. The prompt information may be encoded using a neural network to obtain an implicit representation of the prompt information (P t ) as the input of the inference layer 230.

[0080] When processing the input of the inference layer 230 using the deformable cross-attention layer, the historical scene representation at the current moment and the current scene representation can be used as the query (Q) input of the deformable cross-attention layer, the historical scene representation at the current moment, the prompt information, and the current scene representation can be used as the key (K) input of the deformable cross-attention layer, and the historical scene representation at the current moment, the prompt information, and the current scene representation can be used as the value (V) input of the deformable cross-attention layer. Based on the output of the deformable cross-attention layer, the predicted scene representation at the first moment in the coordinate system of the current moment and the historical scene representation at the first moment in the coordinate system of the current moment can be determined. When multiple deformable cross-attention layers are connected in series, starting with the second layer, each layer uses the historical scene representation and predicted scene representation output by the previous layer as its query (Q) input, the historical scene representation, hint information, and predicted scene representation output by the previous layer as its key (K) input, and the historical scene representation, hint information, and predicted scene representation output by the previous layer as its value (V) input. The output of the last deformable cross-attention layer can be used as the predicted scene representation and historical scene representation of the first moment in the current coordinate system. Recursive reasoning can be achieved by repeatedly reasoning forward based on the output of the previous attention layer.

[0081] The predicted scene representation of at least one future moment output by the inference layer can be used to predict sensor information of at least one future moment. In some embodiments, the autonomous driving model may further include a prediction layer. The prediction layer may be configured to determine predicted sensor information of at least one future moment based on the scene representation of at least one future moment. In some examples, the prediction layer may be implemented by, for example, a cross attention network or any other suitable neural network. The prediction layer may predict N future moments. w ×N H The implicit representation of ×D is reversely mapped to space, and operations such as deconvolution are used to convert the predicted perception representation back to the imaging space of the sensor and then into the image collected by the sensor at the future moment.

[0082] Using the above autonomous driving model, during the reasoning process at time t, H can be represented according to the scene at time t. t , the historical scenario predicted at time t-1 is represented by M pred[:t-1] is used as the historical scene representation M[:t-1] at time t and the driving trajectory S[t:t+1] determines the predicted scene representation at time t+1 and history-aware representation M pred [:t+1]. Furthermore, based on the predicted scene representation and historical scene representation at time t+1, the inference layer can further predict K steps forward and predict the scene representation and historical scene representation at subsequent future moments (t+2, t+3...t+K) to obtain the predicted historical perception sequence (M pred [:t],M pred [:t+1],…,M pred [:t+K-1]), predicted perceptual representation Predicting Actions and predicted trajectory Each step starting at time t+1 is an inference of the next step based on future predictions. During the reasoning process at time t, the historical scene representation at time t+1 can be saved and used for the reasoning process at time t+1. Similarly, predictions can be made continuously for future scenarios.

[0083] The autonomous driving model provided by this disclosure can implement time-series autoregressive reasoning by inferring the prediction result at time t+1 using information at time t. When the reasoning layer adopts a recursive structure, the autonomous driving model provided by the embodiments of this disclosure can implement time-series recursive autoregressive reasoning and determine autonomous driving decisions based on the reasoning results.

[0084] FIG3 shows an exemplary flow chart of a training method for an autonomous driving model according to an embodiment of the present disclosure. The training method 300 shown in FIG3 can be used to train the autonomous driving model 200 described in conjunction with FIG2 .

[0085] In step S302, the sample sensor information within the sample period and the actual driving trajectory corresponding to the sample sensor information may be obtained. In some embodiments, the sample data used to train the autonomous driving model may include The training data time series of RAW t represents the sample sensor information at time t, represents the actual driving trajectory at time t, and T represents the time period to which the sample data belongs. The sample data can be collected by an expert driver driving the vehicle. In some examples, the sample data can also include sample prompt information.

[0086] The above training data time series can be used to obtain the input required by the autonomous driving model, such as the sample sensor information RAW at time t t .

[0087] In step S304, the encoding layer of the autonomous driving model can be used to encode the sample sensor information at the sample moment to obtain a sample scene representation. In step S306, the trajectory planning layer of the autonomous driving model can be used to determine the predicted driving trajectory from the current moment to the future moment based on the historical scene representation at the sample moment. The historical scene representation at the sample moment indicates the scene of at least one previous moment of the sample moment. In step S308, the reasoning layer of the autonomous driving model can be used to determine the predicted scene representation of at least one future moment and the historical scene representation of the future moment based on the sample scene representation, the historical scene representation, and the sample prompt information, wherein the prompt information includes at least the predicted driving trajectory from the current moment to the future moment. The autonomous driving model 200 can be used to process the sample data to implement the above steps S304 to S308, which will not be repeated here.

[0088] In step S310 , the parameters of the autonomous driving model may be adjusted based on the difference between the predicted driving trajectory and the actual driving trajectory from the current moment to the future moment.

[0089] For example, the trajectory planning layer can be used to determine the predicted driving decision action at time t based on the historical scene representation M[:t-1] at time t According to the principle of forward dynamics Determine the first driving trajectory from time t to time t+1 The actual driving trajectory from time t to time t+1 can be determined based on the sample data and can be minimized and The difference between them adjusts the parameters of the autonomous driving model.

[0090] In some embodiments, step S306 may be used to determine the driving decision action at the current sample moment and further determine the driving trajectory from the current sample moment to the next moment. Step S306 may include determining a predicted driving decision action at the sample moment based on the historical scene representation at the sample moment, and determining a first predicted driving trajectory from the sample moment to the next moment based on the predicted driving decision action at the current moment.

[0091] The parameters of the autonomous driving model can be adjusted based on the difference between the actual driving decision action at the current sample moment and the predicted driving decision action determined in step S306. In some embodiments, a first actual driving trajectory from the sample moment to the next moment can be obtained based on the sample data, and the actual driving decision action at the sample moment can be determined based on the first actual driving trajectory. The actual driving decision action can be determined from the actual driving trajectory based on the principles of inverse dynamics. The parameters of the autonomous driving model can then be adjusted based on the difference between the predicted driving decision action at the sample moment and the actual driving decision action at the sample moment.

[0092] For example, the trajectory planning layer can be used to determine the predicted driving decision action at time t based on the historical scene representation M[:t-1] at time t According to the principle of inverse dynamics, the actual driving trajectory from time t to time t+1 can be Determine the actual driving decision action at time t By minimizing and The difference between them adjusts the parameters of the autonomous driving model.

[0093] By using forward / inverse dynamic transformations between driving maneuvers and vehicle trajectories, the unification of model encoding and decision prediction can be achieved.

[0094] In some embodiments, the parameters of the autonomous driving model may also be adjusted based on the difference between the actual sensor information and the predicted sensor information output by the autonomous driving model.

[0095] The autonomous driving model may also include a prediction layer. The prediction layer may be used to determine predicted sensor information for at least one future moment based on a predicted scenario representation for at least one future moment. Actual sensor information for at least one future moment may be determined based on the sample sensor information. Parameters of the autonomous driving model may then be adjusted based on the difference between the actual sensor information for at least one future moment and the predicted sensor information for at least one future moment.

[0096] For example, taking the future moment as time t+1 as an example, the predicted sensor information at the future moment can be determined based on the predicted scene representation at time t+1 output by the autonomous driving model. The sample sensor information RAW at time t+1 can be obtained from the training sample data t+1 , and can be minimized by and RAW t+1 The parameters of the autonomous driving model are adjusted based on the difference between them.

[0097] During the training process of an autonomous driving model, when using the inference layer to output historical and predicted scene representations for future moments, the real driving trajectories in the training data can be used to implement coordinate system transformation. For example, the inference layer of the autonomous driving model can determine the predicted scene representation for the next moment in the coordinate system of the sample moment and the historical scene representation for the next moment in the coordinate system of the sample moment based on the sample scene representation and the historical scene representation, and perform a spatial transformation on the predicted scene representation for the next moment in the coordinate system of the sample moment and the historical scene representation for the next moment in the coordinate system of the sample moment based on the first real driving trajectory from the sample moment to the next moment, to obtain the predicted scene representation for the next moment in the coordinate system of the next moment and the historical scene representation for the next moment in the coordinate system of the next moment. This can improve the accuracy of the model's prediction of future scenes and improve the efficiency of the model's learning for historical scenes.

[0098] In some examples, each segment of the training data time series [1, T] can be divided into smaller batches for processing. By using the autonomous driving model to make forward predictions on the training data, the predicted sequence can be obtained. By minimizing The corresponding real sequence The error between them is used to adjust the parameters of the autonomous driving model.

[0099] FIG4 illustrates an exemplary flow chart of an autonomous driving method 400 implemented using an autonomous driving model according to an embodiment of the present disclosure. The autonomous driving model described in conjunction with FIG2 can be used to implement autonomous driving method 400. The advantages of the autonomous driving model described in conjunction with FIG2 also apply to autonomous driving method 400 and are not further elaborated here.

[0100] In step S402 , the sensor information at the current moment may be encoded to obtain a current scene representation.

[0101] In step S404 , a driving trajectory from the current moment to a future moment may be determined based on a historical scene representation at the current moment, wherein the historical scene representation at the current moment indicates a scene at at least one previous moment of the current moment.

[0102] In step S406, at least one predicted scene representation at a future moment and a historical scene representation at a future moment may be determined based on the current scene representation, the historical scene representation, and the current prompt information, wherein the prompt information includes at least the driving trajectory.

[0103] By utilizing the autonomous driving method provided by the embodiments of the present disclosure, the autonomous driving model can utilize an integrated reasoning layer to realize the learning of historical information and the prediction of future information, thereby enabling the model to make future predictions while learning history, thereby improving the driving ability of the autonomous driving model.

[0104] FIG5 shows an example of an inference process at time t according to an embodiment of the present disclosure.

[0105] As shown in Figure 5 , the encoding layer 510 can be used to encode raw sensor information 510 to obtain a current scene representation 502 at time t. The historical scene representation 504 at time t can be determined based on the inference results at time t-1. The trajectory planning layer 520 of the autonomous driving model can include an action prediction layer 521 and a forward dynamics layer 522. The action prediction layer can process the historical scene representation 504 at time t to obtain a predicted driving decision action 505 at time t. Furthermore, the forward dynamics layer 522 can process the action 505 to obtain a predicted trajectory 506 at time t. The predicted trajectory 506 can be input into the inference layer 530 as the prompt information 503 at time t.

[0106] The reasoning layer 530 of the autonomous driving model may include multiple deformable attention layers 531 and a spatial transformation layer 522. The multiple deformable attention layers 531 may process the current scene representation 502, the historical scene representation 504, and the prompt information 503 to obtain an updated historical scene representation 543 for time t+1 and an updated current representation 544 for time t+1 in the coordinate system at time t. The spatial transformation layer 532 may transform the updated historical scene representation 543 and the updated current representation 544 from the coordinate system at time t to the coordinate system at time t+1. Spatial transformation may be achieved using, for example, a warp operation, using the vehicle positions at time t and time t+1 indicated in the trajectory 506. The spatial transformation layer 532 may output a historical scene representation 509 for time t+1 and a current scene representation 507 for time t+1 in the coordinate system at time t+1. Similarly, reasoning may continue based on the historical scene representation 509 and the current scene representation 507 in the same manner. 5 , the multiple deformable attention layers 531 for processing the historical scene representation 509 and the current scene representation 507 and the multiple deformable attention layers 531 for processing the historical scene representation 504 and the current scene representation 502 can have the same parameters. Using the inference layers with the same parameters can maintain the stability of temporal reasoning.

[0107] FIG6 shows a block diagram of an automatic driving device 600 according to an embodiment of the present disclosure. As shown in FIG6 , the automatic driving device 600 includes an encoding unit 610 , a trajectory determination unit 420 , and a future prediction unit 630 .

[0108] The encoding unit 610 may be configured to encode the sensor information at a current moment to obtain a current scene representation.

[0109] The trajectory determination unit 620 may be configured to determine a driving trajectory from a current moment to a future moment according to a historical scene representation at the current moment, wherein the historical scene representation at the current moment indicates a scene at at least one previous moment of the current moment.

[0110] The future prediction unit 630 may be configured to determine at least one predicted scene representation at a future moment and a historical scene representation at a future moment based on the current scene representation, the historical scene representation, and the current prompt information, wherein the prompt information includes at least a driving trajectory.

[0111] It should be understood that the various modules or units of the apparatus 600 shown in FIG6 may correspond to the various steps in the method 400 described with reference to FIG4 . Thus, the operations, features, and advantages described above for the method 400 are also applicable to the apparatus 600 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.

[0112] FIG7 shows a block diagram of an autonomous driving training device 700 according to an embodiment of the present disclosure. As shown in FIG7 , the autonomous driving device 700 includes a sample acquisition unit 710 , an encoding unit 720 , a trajectory prediction unit 730 , an inference unit 740 , and a parameter adjustment unit 750 .

[0113] The sample acquisition unit 710 may be configured to acquire sample sensor information within a sample period and a real driving trajectory corresponding to the sample sensor information.

[0114] The encoding unit 720 can be configured to encode the sample sensor information at the sample moment using the encoding layer of the autonomous driving model to obtain a sample scene representation.

[0115] The trajectory prediction unit 730 can be configured to determine a predicted driving trajectory from a current moment to a future moment based on a historical scene representation at a sample moment using a trajectory planning layer of the autonomous driving model, wherein the historical scene representation at the sample moment indicates a scene at at least one previous moment of the sample moment.

[0116] The reasoning unit 740 may be configured to determine, using the reasoning layer of the autonomous driving model, at least one predicted scene representation at a future moment and a historical scene representation at a future moment based on the sample scene representation, the historical scene representation, and the sample prompt information, wherein the prompt information includes at least a predicted driving trajectory from the current moment to the future moment; and

[0117] The parameter adjustment unit 750 may be configured to adjust the parameters of the autonomous driving model based on the difference between the predicted driving trajectory from the current moment to the future moment and the actual driving trajectory from the current moment to the future moment.

[0118] It should be understood that the various modules or units of the apparatus 700 shown in FIG7 may correspond to the various steps in the method 300 described with reference to FIG3 . Thus, the operations, features, and advantages described above for the method 300 are also applicable to the apparatus 700 and the modules and units included therein. For the sake of brevity, certain operations, features, and advantages are not described in detail herein.

[0119] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various units discussed herein may be separated into multiple units, and / or at least some functionality of multiple units may be combined into a single unit.

[0120] It should also be understood that various technologies can be described herein in the general context of software hardware elements or program modules. The various units described above with respect to Figure 6 and Figure 7 can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program codes / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of units 610 to 630 and units 710 to 750 can be implemented together in a system on chip (SoC). SoC can include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and can optionally execute the received program code and / or include embedded firmware to perform functions.

[0121] According to another aspect of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute a method according to an embodiment of the present disclosure.

[0122] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, where the computer instructions are used to cause the computer to execute the method according to the embodiment of the present disclosure.

[0123] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the method according to the embodiment of the present disclosure when executed by a processor.

[0124] According to another aspect of the present disclosure, an autonomous driving vehicle is provided, comprising an autonomous driving device according to an embodiment of the present disclosure and one of the above-mentioned electronic devices.

[0125] With reference to Figure 8, a block diagram of an electronic device 800 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0126] As shown in Figure 8, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In RAM 803, various programs and data required for the operation of electronic device 800 can also be stored. Computing unit 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0127] Multiple components within electronic device 800 are connected to I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. Input unit 806 can be any type of device capable of inputting information into electronic device 800. Input unit 806 can receive input numeric or character information and generate key signal input related to user settings and / or function control of the electronic device. It may include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 807 can be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. It may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0128] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the methods (or processes) 300 and 400. For example, in some embodiments, the methods (or processes) 300 and 400 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the methods (or processes) 300 and 400 described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the methods (or processes) 300 and 400 in any other appropriate manner (eg, by means of firmware).

[0129] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0131] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0133] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0134] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0135] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0136] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. In addition, the steps may be performed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples may be combined in various ways. It is important that as technology evolves, many of the elements described herein may be replaced by equivalent elements that appear after this disclosure.

Claims

1. An autonomous driving model, comprising a coding layer, a trajectory planning layer and an inference layer, wherein: The encoding layer is configured to encode sensor information at a current moment to obtain a current scene representation; The trajectory planning layer is configured to determine a driving trajectory from the current moment to a future moment according to a historical scene representation of the current moment, wherein the historical scene representation of the current moment indicates a scene of at least one previous moment of the current moment; The reasoning layer is configured to determine at least one predicted scene representation at a future moment and a historical scene representation at the future moment based on the current scene representation, the historical scene representation, and current prompt information, wherein the prompt information includes at least the driving trajectory.

2. The automatic driving model according to claim 1, wherein: The historical scene representation at the current moment is predicted based on the current scene representation at the previous moment and the historical scene representation at the previous moment.

3. The automatic driving model according to claim 1 or 2, wherein: The trajectory planning layer is configured as: Determining a driving decision action at the current moment according to the historical scene representation at the current moment; A first driving trajectory from the current moment to a first moment is determined according to the driving decision action at the current moment, wherein the first moment is a moment next to the current moment.

4. The automatic driving model as claimed in claim 3, wherein: The inference layer is configured to: A predicted scene representation at the first moment and a historical scene representation at the first moment are determined according to the current scene representation, the historical scene representation, and a first driving trajectory from the current moment to the first moment.

5. The automatic driving model as claimed in claim 4, wherein: Determining the predicted scene representation at the first moment and the historical scene representation at the first moment according to the current scene representation, the historical scene representation, and the first driving trajectory from the current moment to the first moment includes: Determine, according to the current scene representation and the historical scene representation, a predicted scene representation at a first moment in the coordinate system of the current moment and a historical scene representation at the first moment in the coordinate system of the current moment; According to the first driving trajectory, the predicted scene representation at the first moment in the coordinate system of the current moment and the historical scene representation at the first moment in the coordinate system of the current moment are spatially transformed to obtain the predicted scene representation at the first moment in the coordinate system of the first moment and the historical scene representation at the first moment in the coordinate system of the first moment.

6. The automatic driving model as claimed in claim 4, wherein: The inference layer is also configured to: The predicted scenario representation at the second moment and the historical scenario representation at the second moment are determined according to the predicted scenario representation at the first moment, the historical scenario representation at the first moment, and a second driving trajectory from the first moment to the second moment, wherein the second moment is a future moment of the first moment.

7. The automatic driving model according to claim 6, wherein: The second moment is a moment next to the first moment.

8. The automatic driving model as claimed in claim 7, wherein: The time step between the second moment and the first moment is different from the time step between the first moment and the current moment.

9. The automatic driving model according to claim 6, wherein: The trajectory planning layer is also configured to: determining a driving decision action at the first moment according to the historical scene representation at the first moment; The second driving trajectory is determined according to the driving decision action at the first moment.

10. The automatic driving model according to claim 6, wherein: The trajectory planning layer is also configured to: Determining the driving decision action at the first moment according to the predicted historical scene representation at the first moment, wherein the predicted historical scene representation is predicted based on the scene representation at a moment before the current moment and the historical scene representation at the moment before the current moment; A second driving trajectory from the first moment to the second moment is determined according to the driving decision action at the first moment.

11. The autonomous driving model as described in any one of claims 1-10, wherein the reasoning layer includes at least one deformable cross-attention layer.

12. The automatic driving model according to claim 11, wherein: The prompt information also includes at least one of the following information: map information, current vehicle driving status information, navigation information, destination information and driving instruction information.

13. The automatic driving model according to claim 11, wherein: Determining at least one predicted scene representation at a future moment and the historical scene representation at the future moment according to the current scene representation, the historical scene representation, and the current prompt information includes: Taking the historical scene representation at the current moment and the current scene representation as the query (Q) input of the deformable cross attention layer; Using the historical scene representation at the current moment, the prompt information and the current scene representation as a key (K) input of the deformable cross attention layer; Input the historical scene representation at the current moment, the prompt information, and the current scene representation as the value (V) of the deformable cross attention layer; The predicted scene representation of the future moment and the historical scene representation of the future moment in the coordinate system of the current moment are determined according to the output of the deformable cross-attention layer.

14. The autonomous driving model as described in any one of claims 1-13 further includes a prediction layer, which is configured to determine the predicted sensor information of at least one future moment based on the predicted scene representation of at least one future moment.

15. A training method for an autonomous driving model, comprising: Acquire sample sensor information within a sample period and a real driving trajectory corresponding to the sample sensor information; Encoding the sample sensor information at the sample time using the encoding layer of the autonomous driving model to obtain a sample scene representation; Determining, using a trajectory planning layer of the autonomous driving model, a predicted driving trajectory from the current moment to a future moment based on a historical scene representation at the sample moment, wherein the historical scene representation at the sample moment indicates a scene at at least one previous moment of the sample moment; Determine, using the inference layer of the autonomous driving model, at least one predicted scene representation at a future moment and a historical scene representation at the future moment according to the sample scene representation, the historical scene representation, and sample prompt information, wherein the prompt information at least includes a predicted driving trajectory from the current moment to the future moment; Adjust the parameters of the autonomous driving model based on the difference between the predicted driving trajectory from the current moment to the future moment and the actual driving trajectory from the current moment to the future moment.

16. The training method according to claim 15, wherein: Determining the predicted driving trajectory from the current moment to the future moment according to the historical scene representation of the sample moment includes: Determining a predicted driving decision action at the sample moment according to the historical scene representation at the sample moment; A first predicted driving trajectory from the sample moment to the next moment is determined according to the predicted driving decision action at the current moment.

17. The training method according to claim 16, wherein: Determining at least one predicted scene representation at a future moment and the historical scene representation at the future moment according to the sample scene representation, the historical scene representation, and the sample prompt information includes: Determine, according to the sample scene representation and the historical scene representation, a predicted scene representation of the next moment in the coordinate system of the sample moment and a historical scene representation of the next moment in the coordinate system of the sample moment; According to the first real driving trajectory from the sample moment to the next moment, the predicted scene representation of the next moment in the coordinate system of the sample moment and the historical scene representation of the next moment in the coordinate system of the sample moment are spatially transformed to obtain the predicted scene representation of the next moment in the coordinate system of the next moment and the historical scene representation of the next moment in the coordinate system of the next moment.

18. The training method according to claim 16, further comprising: Obtaining a first real driving trajectory from the sample moment to the next moment; Determining a real driving decision action at the sample moment according to the first real driving trajectory; Adjust the parameters of the autonomous driving model according to the difference between the predicted driving decision action at the sample moment and the actual driving decision action at the sample moment.

19. The training method according to any one of claims 15 to 18, wherein: The autonomous driving model also includes a prediction layer, The training method further comprises: The prediction layer is used to determine the predicted sensor information for the at least one future time according to the predicted scene representation for the at least one future time. Determine the real sensor information at at least one future moment according to the sample sensor information; Adjust the parameters of the autonomous driving model according to the difference between the actual sensor information of at least one future moment and the predicted sensor information of at least one future moment.

20. An automatic driving method implemented by using the automatic driving model according to any one of claims 1 to 14, comprising: Encode the sensor information at the current moment to obtain the current scene representation; determining a driving trajectory from the current moment to a future moment according to a historical scene representation of the current moment, wherein the historical scene representation of the current moment indicates a scene of at least one previous moment of the current moment; as well as At least one predicted scene representation at a future moment and a historical scene representation at the future moment are determined according to the current scene representation, the historical scene representation and the current prompt information, wherein the prompt information at least includes the driving trajectory.

21. An automatic driving device based on the automatic driving model according to any one of claims 1 to 14, comprising: an encoding unit configured to encode sensor information at a current moment to obtain a current scene representation; a trajectory determination unit configured to determine a driving trajectory from the current moment to a future moment according to a historical scene representation of the current moment, wherein the historical scene representation of the current moment indicates a scene of at least one previous moment of the current moment; as well as A future prediction unit is configured to determine at least one predicted scene representation at a future moment and a historical scene representation at the future moment based on the current scene representation, the historical scene representation and current prompt information, wherein the prompt information at least includes the driving trajectory.

22. A training device for an autonomous driving model, comprising: A sample acquisition unit, configured to acquire sample sensor information within a sample period and a real driving trajectory corresponding to the sample sensor information; an encoding unit configured to encode the sample sensor information at the sample time using the encoding layer of the autonomous driving model to obtain a sample scene representation; a trajectory prediction unit configured to determine a predicted driving trajectory from the current moment to a future moment according to a historical scene representation at the sample moment using a trajectory planning layer of the autonomous driving model, wherein the historical scene representation at the sample moment indicates a scene at at least one previous moment of the sample moment; The reasoning unit is configured to determine at least one predicted scene representation at a future moment and the predicted scene representation at the future moment according to the sample scene representation, the historical scene representation and the sample prompt information using the reasoning layer of the autonomous driving model. A historical scene representation, wherein the prompt information at least includes a predicted driving trajectory from the current moment to the future moment; as well as A parameter adjustment unit is configured to adjust the parameters of the autonomous driving model based on the difference between the predicted driving trajectory from the current moment to the future moment and the actual driving trajectory from the current moment to the future moment.

23. An electronic device, comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 15 to 20.

24. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 15-20.

25. A computer program product comprising a computer program, wherein: The computer program implements the method according to any one of claims 15 to 20 when executed by a processor.

26. An autonomous driving vehicle comprising: One of the automatic driving device according to claim 21 and the electronic device according to claim 23.

Citation Information

Patent Citations

  • Automatic driving method based on similar scene mining and vehicle

    CN115675528A

  • Scene coding model training method, trajectory planning method and device

    CN115861953A

  • Method for determining planned trajectory, model training method and autonomous vehicle

    CN116394977A

  • Automatic driving model for predicting position trajectory and training method thereof

    CN116560377A

  • Automatic driving model capable of autonomously interacting with personnel outside vehicle and training method

    CN116776151A