An autonomous driving model capable of natural language interaction and its training method

By introducing generators and multimodal coding layers into the autonomous driving model, planning is directly based on perceived information and outputting interactive information, the problem of lack of explanation of autonomous driving strategy information is solved, and user experience and trust are improved.

CN117010265BActive Publication Date: 2025-06-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310403812.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-14
Publication Date
2025-06-13
Estimated Expiration
2043-04-14

AI Technical Summary

Technical Problem

During the process of autonomous driving, the lack of explanation of the information on autonomous driving strategy has led to users' doubts and distrust of their autonomous driving behavior.

Method used

An autonomous driving model is provided, which includes a generator and a multimodal coding layer, which can generate target autonomous driving policy information and interactive information based on input perceptual information and historical interaction information, and plan directly based on perceptual information.

Benefits of technology

By outputting interactive information, users can interact with the autonomous driving model, improve the passenger user experience during autonomous driving, and reduce users' doubts about autonomous driving behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117010265B_ABST
    Figure CN117010265B_ABST
Patent Text Reader

Abstract

The present disclosure provides an autonomous driving model capable of natural language interaction and a training method thereof, relating to the field of computer technologies, and particularly to the field of autonomous driving technologies. The autonomous driving model includes a generator configured to generate target autonomous driving policy information and target interaction information based on input first input information, wherein the first input information is related to perception information regarding the vehicle surrounding environment and includes historical interaction information. By using the embodiments of the present disclosure, while obtaining autonomous driving policy information using the input perception information, interaction information can also be output, so that a user can realize interaction with the autonomous driving model through natural language expression, improving the user experience of passengers during the autonomous driving process. The autonomous driving policy can be adjusted according to the interaction with the user, and the user trust can be increased by providing the user with system response information or autonomous driving policy-related information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular to the field of autonomous driving technologies. Specifically, the present disclosure relates to an autonomous driving model capable of interaction, an autonomous driving method implemented by using the autonomous driving model, a training method of the autonomous driving model, an autonomous driving device, a training device, an electronic device, a computer-readable storage medium, a computer program product, and an autonomous driving vehicle. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), including both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] Autonomous driving technology integrates technologies in many aspects such as recognition, decision-making, positioning, communication security, and human-computer interaction. During the autonomous driving process, the autonomous driving model obtains an autonomous driving strategy for controlling the vehicle behavior based on the input information. In this process, the lack of explanation for the autonomous driving strategy information may lead to doubts and distrust of users regarding the autonomous driving behavior.

[0004] The methods described in this section are not necessarily methods that have been previously conceived or adopted. Unless otherwise specified, any method described in this section should not be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides an autonomous driving model capable of interaction, an autonomous driving method implemented by using the autonomous driving model, a training method of the autonomous driving model, an autonomous driving device, a training device, an electronic device, a computer-readable storage medium, a computer program product, and an autonomous driving vehicle.

[0006] According to one aspect of the present disclosure, there is provided an autonomous driving model including a generator configured to generate target autonomous driving strategy information and target interaction information based on input first input information, wherein the first input information is related to perception information about the vehicle surrounding environment and includes historical interaction information, and wherein the autonomous driving model further includes a multi-modal encoding layer configured to output an implicit representation corresponding to the second input information based on the input second input information, the second input information including perception information about the vehicle surrounding environment obtained by using sensors, and wherein the first input information includes the implicit representation corresponding to the second input information.

[0007] According to another aspect of the present disclosure, there is provided an autonomous driving method implemented by using an autonomous driving model, the autonomous driving model including a generator, the method including: obtaining first input information, wherein the first input information is related to perception information about the vehicle surrounding environment and includes historical interaction information; and inputting the first input information into the generator of the autonomous driving model to generate target autonomous driving strategy information and target interaction information, wherein the autonomous driving model further includes a multi-modal encoding layer configured to output an implicit representation corresponding to the second input information based on the input second input information, the second input information including perception information about the vehicle surrounding environment obtained by using sensors, and wherein the first input information includes the implicit representation corresponding to the second input information.

[0008] According to another aspect of the present disclosure, there is provided a training method for an autonomous driving model, the autonomous driving model including a generator, the method including: obtaining first sample input information, wherein the first sample input information is related to perception information about the vehicle surrounding environment and includes sample historical interaction information; obtaining true autonomous driving strategy information and true interaction information corresponding to the first sample input information; inputting the first sample input information into the generator of the autonomous driving model to generate predicted autonomous driving strategy information and predicted interaction information; and adjusting parameters of the autonomous driving model based on a difference between the predicted autonomous driving strategy information and the true autonomous driving strategy information and a difference between the predicted interaction information and the true interaction information, wherein the autonomous driving model further includes a multi-modal encoder, the multi-modal encoding layer configured to output a sample implicit representation corresponding to the second sample input information based on the input second sample input information, the second sample input information including perception information about the vehicle surrounding environment obtained by using sensors, and wherein the first sample input information includes the sample implicit representation corresponding to the second sample input information.

[0009] According to another aspect of the present disclosure, there is provided an autonomous driving device based on an autonomous driving model, including: an input information acquisition unit configured to acquire first input information, where the first input information is related to the perception information of the vehicle surrounding environment and includes historical interaction information; and a generation unit configured to input the first input information into a generator of the autonomous driving model to generate target autonomous driving policy information and target interaction information, where the autonomous driving model further includes a multimodal encoding layer configured to output an implicit representation corresponding to the second input information based on the input second input information, the second input information including the perception information of the vehicle surrounding environment obtained by using sensors, and where the first input information includes the implicit representation corresponding to the second input information.

[0010] According to another aspect of the present disclosure, there is provided a training device for an autonomous driving model, the autonomous driving model including a multimodal encoding layer and a generator, the training device including: a sample information acquisition unit configured to acquire first sample input information, where the sample input information is related to the perception information of the vehicle surrounding environment and includes sample historical interaction information; a ground truth information acquisition unit configured to acquire the ground truth autonomous driving policy information and the ground truth interaction information corresponding to the first sample input information; a generator training unit configured to input the first sample input information into the generator of the autonomous driving model to generate predicted autonomous driving policy information and predicted interaction information; and a parameter adjustment unit configured to adjust parameters of the autonomous driving model based on the difference between the predicted autonomous driving policy information and the ground truth autonomous driving policy information and the difference between the predicted interaction information and the ground truth interaction information, where the autonomous driving model further includes a multimodal encoding layer configured to output an implicit representation corresponding to the second input information based on the input second input information, the second input information including the perception information of the vehicle surrounding environment obtained by using sensors, and where the first input information includes the implicit representation corresponding to the second input information.

[0011] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.

[0012] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause the computer to execute the method as described above.

[0013] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method as described above.

[0014] According to another aspect of the present disclosure, there is provided a self-driving vehicle including: the device for training the self-driving model as described above or the electronic device as described above.

[0015] According to one or more embodiments of the present disclosure, while obtaining the self-driving strategy information by using the input perception information, interaction information can also be output, so that the user can interact with the self-driving model, improving the user experience of passengers during the self-driving process.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0017] The drawings exemplarily show embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0018] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an embodiment of the present disclosure;

[0019] Figure 2 A schematic diagram of a self-driving model according to an embodiment of the present disclosure;

[0020] Figure 3 A schematic diagram of another self-driving model according to an embodiment of the present disclosure;

[0021] Figure 4 An exemplary flowchart of a self-driving method according to an embodiment of the present disclosure;

[0022] Figure 5 An exemplary flowchart of a method for training a self-driving model according to an embodiment of the present disclosure;

[0023] Figure 6 An exemplary block diagram of a self-driving device based on a self-driving model according to an embodiment of the present disclosure;

[0024] Figure 7Shows an exemplary block diagram of a training device for an autonomous driving model according to an embodiment of the present disclosure; and

[0025] Figure 8 Shows a block diagram of an exemplary electronic device that can be used to implement an embodiment of the present disclosure. Detailed implementation manners

[0026] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0027] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0028] In the description of various examples in the present disclosure, the terms used are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element may be one or more. In addition, the term "and / or" used in the present disclosure covers any one of the listed items and all possible combinations.

[0029] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0030] Figure 1 Shows a schematic diagram of an exemplary system 100 in which the various methods and devices described herein can be implemented according to an embodiment of the present disclosure. Referring Figure 1 , the system 100 includes a motor vehicle 110, a server 120, and one or more communication networks 130 that couple the motor vehicle 110 to the server 120.

[0031] In an embodiment of the present disclosure, the motor vehicle 110 may include a computing device according to an embodiment of the present disclosure and / or be configured to execute a method according to an embodiment of the present disclosure.

[0032] Server 120 can run one or more services or software applications that enable autonomous driving. In some embodiments, Server 120 can also provide other services or software applications, which can include non-virtual environments and virtual environments. In Figure 1 the configuration shown, Server 120 can include one or more components that implement the functions performed by Server 120. These components can include software components, hardware components, or a combination thereof that can be executed by one or more processors. A user of motor vehicle 110 can in turn utilize one or more client applications to interact with Server 120 to utilize the services provided by these components. It should be understood that various different system configurations are possible, which can be different from System 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0033] Server 120 can include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, Server 120 can run one or more services or software applications that provide the functions described below.

[0034] The computing units in Server 120 can run one or more operating systems including any of the above operating systems as well as any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0035] In some implementations, Server 120 can include one or more applications to analyze and merge data feeds and / or event updates received from motor vehicle 110. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of motor vehicle 110.

[0036] Network 130 can be any type of network well-known to those skilled in the art, which can support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 130 can be a satellite communication network, a local area network (LAN), an Ethernet-based network, token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (including for example Bluetooth, WiFi) and / or any combination of these with other networks.

[0037] System 100 may also include one or more databases 150. In certain embodiments, these databases can be used to store data and other information. For example, one or more of the databases 150 can be used to store information such as audio files and video files. The data repositories 150 can reside at various locations. For example, a data repository used by the server 120 can be local to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network-based or dedicated connection. The data repositories 150 can be of different types. In certain embodiments, a data repository used by the server 120 can be a database, such as a relational database. One or more of these databases can store, update, and retrieve data to and from the database in response to commands.

[0038] In certain embodiments, one or more of the databases 150 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.

[0039] The motor vehicle 110 may include sensors 111 for sensing the surrounding environment. The sensors 111 may include one or more of the following sensors: visual cameras, infrared cameras, ultrasonic sensors, millimeter-wave radars, and lidars (LiDARs). Different sensors may provide different detection accuracies and ranges. The cameras may be installed in the front, rear, or other positions of the vehicle. The visual cameras may capture the situations inside and outside the vehicle in real time and present them to the driver and / or passengers. In addition, by analyzing the images captured by the visual cameras, information such as traffic signal indications, intersection situations, and the operating states of other vehicles can be obtained. The infrared cameras may capture objects in night vision conditions. The ultrasonic sensors may be installed around the vehicle and are used to measure the distance between the vehicle and external objects by taking advantage of the strong directivity of ultrasonic waves. The millimeter-wave radars may be installed in the front, rear, or other positions of the vehicle and are used to measure the distance between the vehicle and external objects by using the characteristics of electromagnetic waves. The lidars may be installed in the front, rear, or other positions of the vehicle and are used to detect object edge and shape information for object recognition and tracking. Due to the Doppler effect, the radar device can also measure the speed change between the vehicle and a moving object.

[0040] The motor vehicle 110 may further include a communication device 112. The communication device 112 may include a satellite positioning module capable of receiving satellite positioning signals (e.g., Beidou, GPS, GLONASS, and GALILEO) from satellites 141 and generating coordinates based on these signals. The communication device 112 may further include a module for communicating with a mobile communication base station 142. The mobile communication network may implement any suitable communication technology, such as currently or continuously evolving wireless communication technologies like GSM / GPRS, CDMA, LTE (e.g., 5G technology). The communication device 112 may also have a vehicle networking or vehicle-to-everything (V2X) module configured to enable vehicle-to-vehicle (V2V) communication with other vehicles 143 and vehicle-to-infrastructure (V2I) communication with infrastructure 144, for example. In addition, the communication device 112 may have a module configured to communicate with a user terminal 145 (including but not limited to smartphones, tablets, or wearable devices such as watches) via a wireless local area network or Bluetooth using the IEEE802.11 standard, for example. With the communication device 112, the motor vehicle 110 can also access the server 120 via the network 130.

[0041] The motor vehicle 110 may further include a control device 113. The control device 113 may include a processor communicating with various types of computer-readable storage devices or media, such as a central processing unit (CPU) or a graphics processing unit (GPU), or other dedicated processors, etc. The control device 113 may include an autonomous driving system for automatically controlling various actuators in the vehicle. The autonomous driving system is configured to control the powertrain, steering system, braking system, etc. of the motor vehicle 110 (not shown) via a plurality of actuators in response to inputs from a plurality of sensors 111 or other input devices to respectively control acceleration, steering, and braking without human intervention or with limited human intervention. Some of the processing functions of the control device 113 may be implemented through cloud computing. For example, some processing may be performed using an in-vehicle processor, while other processing may utilize computing resources in the cloud. The control device 113 may be configured to execute the method according to the present disclosure. In addition, the control device 113 may be implemented as an example of a computing device on the motor vehicle side (client) according to the present disclosure.

[0042] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and devices described according to the present disclosure.

[0043] Before deploying the autonomous driving model to a real vehicle for real vehicle training, in order to preliminarily train the autonomous driving model, the autonomous driving model can be trained in a simulation environment. However, due to the differences between the simulation data generated in the simulation environment and the data collected by sensors in the real environment, policy failures may occur during the process of directly migrating the autonomous driving strategy from the simulation environment to the real environment.

[0044] To address the issue of the difference between the real environment and the simulation environment, the following methods are mainly adopted in related technologies to achieve the migration of autonomous driving strategies from the simulation environment to the real environment: (1) Domain Adaption method. In this method, a mapping function that maps the common states of the simulation environment and the real environment to the latent variable space is learned. During the training of the algorithm in the simulation environment, the mapped state space is used. Then, when the model is migrated to the real environment, after mapping the states to the latent space in the same way, the model trained in the simulation environment can be directly applied; (2) Progressive Network. In this method, the migration from the simulation environment to the real environment is carried out by gradually transitioning from simple tasks to complex tasks. The agent first trains on simple tasks in the simulation environment, and then gradually increases the difficulty of the tasks until the training of complex tasks is finally completed; (3) Inverse Dynamic Model. In this method, an inverse transition probability matrix is learned in the real environment to directly apply the model trained in the simulation environment in the real environment. The inverse transition probability matrix is used to map the states of the real environment to the states in the simulation environment, thereby achieving transfer learning of the model; (4) Domain Randomization. In this method, visual information or physical parameters are randomized in the simulation environment. For example, in an obstacle avoidance task, the agent learns in a simulation environment where parameters such as wall color, floor color, or friction, and atmospheric pressure randomly change, so that the model trained in the simulation environment is more robust and has better generalization ability, and can better adapt to the changes in the real environment.

[0045] The methods used in related technologies usually result in the loss of perceptual information in order to map the simulation environment and the real environment to a unified feature space. For example, mapping both simulation images and real images to image semantic features (such as semantic features based on image segmentation) loses a large amount of original image features, leading to a decline in the effect of the final autonomous driving strategy. The Inverse Dynamic Model uses the inverse transition probability matrix to map the states of the real environment to the states in the simulation environment, but this cannot accurately model the real environment. This overly simplified environmental modeling method will cause the strategies learned in the simulation environment to not achieve ideal effects in the real environment. Domain Randomization can partially solve the problem of migrating autonomous driving strategies from the simulation environment to the real environment, but this method is based on the premise that the simulator's modeling of the environmental elements in the real environment is relatively accurate. In the field of autonomous driving, it is usually very difficult for the simulator to accurately model the human traffic driving environment. Once the perceptual information in the simulation environment differs greatly from that in the real environment, Domain Randomization cannot achieve the desired effect.

[0046] To solve the above problems, the present disclosure provides a new method for training an autonomous driving model.

[0047] According to one aspect of the present disclosure, an autonomous driving model is provided. Figure 2 A schematic diagram of an autonomous driving model 200 according to an embodiment of the present disclosure is shown.

[0048] As Figure 2 shown, the autonomous driving model 200 includes a multi-modal encoding layer 210 and a decision control layer 220. The multi-modal encoding layer 210 and the decision control layer 220 are connected to form an end-to-end neural network model, so that the decision control layer 220 directly obtains autonomous driving policy information based on the output of the multi-modal encoding layer 210. The first input information of the multi-modal encoding layer 210 includes the navigation information In1 of the target vehicle and the perception information of the surrounding environment of the target vehicle obtained by using sensors (for example, but not limited to, including In2, In3, and In4. In the following content, the perception information including In2, In3, and In4 is taken as an example for description). The perception information includes the current perception information and historical perception information of the surrounding environment of the target vehicle during the driving process of the target vehicle. The multi-modal encoding layer 210 is configured to obtain an implicit representation e t corresponding to the first input information In1 to In4. The second input information of the decision control layer 220 includes the implicit representation e t , and the decision control layer 220 is configured to obtain the target autonomous driving policy information based on the second input information.

[0049] As described above, in the related art, prediction can be first performed based on the perception information to obtain future prediction information, and then the decision control layer plans based on the future prediction information. That is to say, the decision control layer 220 does not directly plan based on the perception information, but directly plans based on the future prediction information. In the embodiment of the present application, the decision control layer 220 can directly obtain the autonomous driving policy information based on the output of the multi-modal encoding layer 210, and the multi-modal encoding layer 210 is used to perform encoding calculation on the perception information, which is equivalent to the decision control layer 220 can directly plan based on the perception information to obtain the autonomous driving policy information. In other words, in the embodiment of the present application, the perception is directly responsible for the decision.

[0050] In the example, the autonomous driving model 200 may adopt a Transformer network structure with an Encoder and a Decoder. It can be understood that the autonomous driving model 200 may also be other neural network models based on the Transformer network structure, which is not limited herein. The Transformer architecture can calculate the implicit representations of the model input and output through the self-attention mechanism. In other words, the Transformer architecture can be an Encoder-Decoder model constructed based on this self-attention mechanism.

[0051] In the example, the navigation information In1 of the target vehicle in the first input information may include vectorized navigation information and vectorized map information, and the vectorized navigation information and vectorized map information may be obtained by performing vectorization operations on one or more of lane-level or road-level navigation information and rough positioning information.

[0052] According to some embodiments of the present application, the perception information In2, In3, and In4 of the surrounding environment of the target vehicle may include the perception information In2 of one or more cameras, the perception information In3 of one or more lidars, and the perception information In4 of one or more millimeter-wave radars. It can be understood that the perception information of the surrounding environment of the target vehicle is not limited to the above form. For example, it may only include the perception information In2 of multiple cameras, without including the perception information In3 of one or more lidars and the perception information In4 of one or more millimeter-wave radars. The perception information In2 obtained through the camera may be perception information in the form of pictures or videos, and the perception information In3 obtained through the lidar may be perception information in the form of radar point clouds (such as three-dimensional point clouds). In the example, the above different forms of information (pictures, videos, point clouds, etc.) can be directly input into the multi-modal encoding layer 210 without preprocessing. In addition, the perception information includes the current perception information x of the surrounding environment of the target vehicle during the driving process of the vehicle t and the historical perception information x corresponding to multiple historical moments t-Δt , where there may be a time span with a preset duration between t and Δt.

[0053] In the example, the multi-modal encoding layer 210 may perform encoding calculations on the first input information to generate corresponding implicit representations e t . The implicit representation e tFor example, it can be an implicit representation in the Bird's Eye View (BEV) space. For example, the perception information In2 of the camera can be first input into a shared backbone network to extract the data features of each camera. Then, the perception information In2 of multiple cameras is fused and transformed into the BEV space. Next, cross-modal fusion can be performed within the BEV space to fuse pixel-level visual data and lidar point clouds. Finally, temporal fusion is performed to form an implicit representation e in the BEV space t .

[0054] In one example, a Transformer Encoder structure that fuses spatio-temporal information can be used to project the input information of multiple cameras into an implicit representation e in the BEV space t . For example, spatio-temporal information can be utilized through a BEV query mechanism with pre-set parameter grid partitioning (BEV queries). The spatial cross-attention mechanism (i.e., the BEV query mechanism extracts the required spatial features from multi-camera features through the attention mechanism) enables the BEV query mechanism to extract features from the multi-camera perspectives it is interested in, thereby aggregating spatial information; in addition, historical information is fused through the temporal self-attention mechanism (i.e., the BEV features generated at each moment obtain the required temporal information from the BEV features of the previous moment), thereby aggregating temporal information

[0055] Correspondingly, the decision and control layer 220 obtains target autonomous driving policy information based on the input implicit representation e t . The target autonomous driving policy information can, for example, include a planned trajectory Out1 or a control signal Out2 for the vehicle (such as signals for controlling the throttle, brakes, steering angle, etc.). In the example, the control strategy module in the autonomous vehicle can be used to interpret the trajectory planning Out1 to obtain the control signal Out2 for the vehicle; or a neural network can be used to directly output the control signal Out2 for the vehicle based on the implicit representation e t

[0056] In the example, the decision and control layer 220 can include a decoder in the Transformer

[0057] In Figure 2 , the solid arrows between the multi-modal encoding layer 210 and the decision and control layer 220, and between the decision and control layer 220 and the trajectory planning Out1 represent differentiable operations. In other words, during model training, the gradient can be backpropagated through the above solid arrows

[0058] ​It can be seen that in the autonomous driving model 200 according to the embodiments of the present disclosure, the multi-modal encoding layer 210 and the decision control layer 220 are connected to form an end-to-end neural network model. Therefore, the perception information can be directly responsible for the decision-making, and the coupling problem between prediction and planning can be solved. In addition, the introduction of implicit representation can overcome the problem that the algorithm is prone to failure due to the representation defect of structured information. In addition, since the perception is directly responsible for the decision-making, the perception can capture the information that is crucial for the decision-making and reduce the error accumulation caused by perception errors. Moreover, since the perception is directly responsible for the decision-making, the autonomous driving technology with emphasis on perception and light on mapping is realized, and thus the problem of decision failure caused by untimely update and regional limitation of high-precision maps can be overcome. Since the dependence on high-precision maps is eliminated, the update cost of high-precision maps can be saved.

[0059] According to some embodiments, with continued reference to Figure 2 , the autonomous driving model 200 may further include a future prediction layer 230, which is configured to predict future prediction information Out3 for the surrounding environment of the target vehicle based on the input implicit representation e t , and at least a part of the future prediction information Out3 may further be included in the second input information of the decision control layer 220. For example, the future prediction information Out3 may include the obstacle position at a future moment or the sensor input information at a future moment predicted based on the implicit representation e t . At least a part of the future prediction information Out3 may be input into the decision control layer 220 as auxiliary information A, and the decision control layer 220 may predict the target autonomous driving strategy information based on the implicit representation e t and the auxiliary information A.

[0060] In an example, the future prediction layer 230 may be a decoder in a Transformer. In some examples, the future prediction layer 230 and the decision control layer 220 may share the same network structure. The future prediction information involved in the following description may be output by the future prediction layer or the decision control layer.

[0061] In an example, the future prediction information Out3 may output structured prediction information. Correspondingly, the dashed arrows between the future prediction information Out3 to the auxiliary information A and the auxiliary information A to the decision control layer 220 represent non-differentiable operations. In other words, during model training, the gradient cannot be backpropagated through the above dashed arrows. However, since the operations between the multi-modal encoding layer 210 to the future prediction layer 230 and the future prediction layer 230 to the future prediction information Out3 are differentiable operations, the gradient can still be backpropagated in the direction indicated by the solid arrows. In other words, the future prediction layer 230 can also be trained separately.

[0062] Thus, by introducing the future prediction layer 230 into the autonomous driving model 200, at least a part of the information predicted by the future prediction layer 230 is input into the decision control layer 220 as auxiliary information to assist in decision-making, which can improve the accuracy and safety of decision-making. In addition, when training the model, based on the decision control layer 220, the multi-modal encoding layer 210 can be further trained through the future prediction layer 230, so that the encoding of the multi-modal encoding layer 210 is more accurate, and thus the decision control layer 220 can predict more optimized target autonomous driving strategy information.

[0063] According to some embodiments, the future prediction information Out3 may include at least one of the following: future prediction perception information about the environment around the target vehicle (e.g., sensor information at a future moment The sensor information at a future moment includes camera input information or radar input information at a future moment), a future prediction implicit representation corresponding to the future prediction perception information (e.g., an implicit representation in the BEV space corresponding to the sensor information at a future moment), and future prediction detection information about the environment around the target vehicle (e.g., the position of obstacles at a future moment ). Moreover, the future prediction detection information may include the types of multiple obstacles in the environment around the target vehicle and their future prediction state information (including the size of the obstacles and various long-tail information).

[0064] According to some embodiments, continuing to refer to Figure 2 , the autonomous driving model 200 may further include a perception detection layer 240, and the perception detection layer 240 may be configured to obtain target detection information Out4 about the environment around the target vehicle based on the input implicit representation e t . The target detection information Out4 includes current detection information and historical detection information. The current detection information includes the types of multiple road surface elements and obstacles in the environment around the target vehicle and their current state information, and the historical detection information includes the types of multiple obstacles in the environment around the target vehicle and their historical state information. And at least a part of the target detection information Out4 may further be included in the second input information of the decision control layer 220.

[0065] The road surface elements may be stationary objects, while the obstacles may be moving objects, so the historical state information of the road surface elements may not be detected.

[0066] In the example, the target detection information Out4 can be a bounding box in the three-dimensional space for an obstacle, and can indicate the classification, status, etc. of the corresponding obstacle in the bounding box. For example, it can indicate the size, position of the obstacle in the bounding box, as well as the vehicle type, the current state of the vehicle (such as whether the turn signal, high beam, etc., long-tail information are turned on), the position and length of the lane line, etc. It will be understood that the classification of the corresponding obstacle in the bounding box can be one or more of a plurality of predefined categories.

[0067] In addition, the target detection information Out4 (current detection information and historical detection information) can be structured information. Accordingly, the dashed arrows between the target detection information Out4 and the auxiliary information A, and between the auxiliary information A and the decision control layer 220 represent non-differentiable operations. In other words, during model training, the gradient cannot be backpropagated through the above dashed arrows. However, since the operations between the multi-modal encoding layer 210 and the perception detection layer 240, and between the perception detection layer 240 and the target detection information Out4 are differentiable operations, the gradient can still be backpropagated in the direction indicated by the solid arrows. In other words, the perception detection layer 240 can also be trained separately.

[0068] In the example, the perception detection layer 240 can include a decoder in the Transformer.

[0069] Thus, by introducing the perception detection layer 240 into the autonomous driving model 200, at least a part of the information predicted by the perception detection layer 240 is input into the decision control layer 220 as auxiliary information to assist in decision-making, which can enable the detection information for the current and historical periods of time around the vehicle to be used to assist in decision-making, thereby improving the accuracy and safety of decision-making. In addition, during model training, based on the decision control layer 220, the multi-modal encoding layer 210 can be further trained through the perception detection layer 240, so that the encoding of the multi-modal encoding layer 210 is more accurate, and thus the decision control layer 220 can predict more optimized target autonomous driving strategy information.

[0070] According to some embodiments, continuing to refer to Figure 2 , the autonomous driving model 200 may further include an evaluation feedback layer 250, and the evaluation feedback layer 250 may be configured to obtain evaluation feedback information Out5 for the target autonomous driving strategy information based on the input implicit representation e t .

[0071] In the example, the evaluation feedback layer 250 can be a decoder in the Transformer.

[0072] Thus, by introducing the evaluation feedback layer 250 into the autonomous driving model 200, it can be indicated whether the current driving behavior is from a human driver or the model, whether the current driving is comfortable, whether the current driving violates traffic rules, and whether the current driving is dangerous driving, etc., thereby improving the user experience.

[0073] It will be understood that the solid arrows from the multi-modal encoding layer 210 to the evaluation feedback layer 250 and from the evaluation feedback layer 250 to the evaluation feedback information Out5 represent differentiable operations. In other words, during model training, the gradient can be backpropagated through the above solid arrows. Thus, during model training, based on the decision control layer 220, the multi-modal encoding layer 210 can be further trained through the evaluation feedback layer 250, so that the encoding of the multi-modal encoding layer 210 is more accurate, and thus the decision control layer 220 can predict a more optimized target autonomous driving strategy information.

[0074] According to some embodiments, as Figure 2 shown by the dashed arrow of the auxiliary information A including the future prediction information Out3 and the target detection information Out4 in Figure 2 pointing to the evaluation feedback layer 250, when the autonomous driving model 200 includes the future prediction layer 230 and the perception detection layer 240, the evaluation feedback layer 250 can be configured to be based on at least a part of one or both of the input future prediction information Out3 and the target detection information Out4, and the implicit representation e t obtain the evaluation feedback information Out5 for the target autonomous driving strategy information. Thus, the detection information and future prediction information of the current and a historical period of time around the vehicle can be used to assist in the evaluation, improving the accuracy of the evaluation.

[0075] According to some embodiments, the evaluation feedback layer 250 can be configured to be based on the input implicit representation e t and the target autonomous driving strategy information (such as the planned trajectory Out1) to obtain the evaluation feedback information for the target autonomous driving strategy information. Thus, based on the autonomous driving strategy information to assist in the evaluation feedback can further improve the accuracy of the evaluation.

[0076] According to other embodiments of the present application, the evaluation feedback layer 250 can be configured to be based on at least a part of one or both of the input future prediction information Out3 and the target detection information Out4, the target autonomous driving strategy information, and the implicit representation e t obtain the evaluation feedback information Out5 for the target autonomous driving strategy information, so as to further improve the accuracy of the evaluation.

[0077] According to some embodiments, further referring to Figure 2, the autonomous driving model 200 may further include an interpretation layer 260, which may be configured to obtain, based on the input implicit representation e t interpretation information Out6 for the target autonomous driving policy information, and the interpretation information Out6 can characterize the decision classification of the target autonomous driving policy information. Thus, during the autonomous driving process, interpretation information related to the target autonomous driving policy information can be provided to passengers, improving the interpretability of the autonomous driving policy and thus enhancing the user experience.

[0078] In an example, the interpretation layer 260 may classify the target autonomous driving policy information, and each classification may be mapped to a preset natural language statement. For example, the interpretation information Out6 may include natural language statements such as: currently need to change lanes, there is a traffic light ahead so need to slow down, surrounding vehicles may need to cut in, etc. In addition, the interpretation layer 260 may include a decoder in the Transformer to decode natural language for the interpretation of driving policies.

[0079] According to some embodiments, when the autonomous driving model 200 includes a future prediction layer 230 and a perception detection layer 240, the interpretation layer 260 may be configured to obtain, based on at least a part of one or both of the input future prediction information and target detection information, and the implicit representation e t interpretation information Out6 for the target autonomous driving policy information. Thus, the target detection information and future prediction information for the current and a historical period of time around the vehicle can be used to assist in the interpretation, further improving the accuracy and rationality of the interpretation.

[0080] According to some embodiments, continuing to refer to Figure 2 , the interpretation layer 260 may be configured to obtain, based on the input implicit representation e t and the target autonomous driving policy information (such as the planned trajectory Out1), interpretation information for the target autonomous driving policy information. Thus, using the autonomous driving policy information to assist in the interpretation can further improve the accuracy of the interpretation.

[0081] According to other embodiments of the present application, the interpretation layer 260 may be configured to obtain, based on at least a part of one or both of the input future prediction information Out3 and target detection information Out4, the target autonomous driving policy information, and the implicit representation e t interpretation information Out6 for the target autonomous driving policy information, thereby being able to further improve the accuracy of the interpretation.

[0082] According to some embodiments, the sensor may include a camera, and the perception information may include a two-dimensional image captured by the camera. Further, the multimodal encoding layer 210 may be further configured to: obtain an implicit representation e corresponding to the first input information based on the first input information including the two-dimensional image, as well as the internal and external parameters of the camera t .

[0083] In the example, the internal parameters of the camera (i.e., the parameters related to the characteristics of the camera itself, such as the focal length and pixel size of the camera) and the external parameters (i.e., the parameters in the world coordinate system, such as the position and rotation direction of the camera) may be input into the modality encoding layer 210 as hyperparameters of the autonomous driving model 200. The internal and external parameters of the camera can be used to perform the transformation of the input two-dimensional image into, for example, the BEV space

[0084] In addition, the perception information may be a sequence of two-dimensional images captured by multiple cameras respectively

[0085] According to some embodiments, the first input information may further include a lane-level map, and the navigation information may include road-level navigation information and / or lane-level navigation information. Different from the high-precision map, the lane-level map has better availability and smaller space occupancy. Thus, by using the lane-level map and lane-level navigation information, the dependence on the high-precision map can be overcome

[0086] The navigation map may include a road-level map (SD Map), a lane-level map (LD Map), and a high-precision map (HDMap). The road-level map is mainly composed of coarse-grained road topology information, with relatively low navigation positioning accuracy (for example, the accuracy is about 15 meters), and is mainly used to assist the driver in navigation, which cannot meet the requirements of autonomous driving. While the lane-level map and the high-precision map can be used for autonomous driving. The lane-level map adds lane-level topology information, has higher accuracy than the road-level map, generally at the sub-meter level, and may include road information (such as lane lines) and ancillary facility information related to the lane (such as traffic lights, road signs, parking spaces, etc.), and can be used to assist autonomous driving. Compared with the lane-level map, the high-precision map has higher map data accuracy (the accuracy reaches the centimeter level), richer map data types, and higher map update frequency, and can be used for autonomous driving. Among these three navigation maps, the high-precision map has the richest information and the highest accuracy, but also has higher usage and update costs. Since the solution in the embodiments of the present application directly takes responsibility for decision-making through perception, an autonomous driving technology with heavy perception and light map can be realized, so the dependence on the high-precision map can be eliminated and efficient decision-making can be ensured. Further, using the lane-level map as auxiliary information for decision-making can improve the decision-making effect

[0087] According to some embodiments, the perception information may include at least one of the following: images collected by a camera, information collected by a lidar, and information collected by a millimeter-wave radar. It will be understood that the images obtained by the camera may be in the form of pictures or videos, and the information obtained by the lidar may be radar point clouds (e.g., three-dimensional point clouds).

[0088] According to some embodiments, the multimodal encoding layer 210 is configured to map the first input information to a preset space to obtain an intermediate representation, and process the intermediate representation using a temporal attention mechanism and / or a spatial attention mechanism to obtain an implicit representation e corresponding to the first input information t 。

[0089] In an example, the preset space may be a BEV space. Since processes such as perception, prediction, decision-making, and planning are all carried out in a three-dimensional space, and the image information captured by the camera is only a projection of the real physical world in a perspective view, the information obtained from the image needs to undergo complex processing to be used, so there will be a certain amount of information loss. Mapping visual information to the BEV space can more conveniently connect perception and planning control.

[0090] In an example, the first input information (e.g., the image information in the first input information) may be first input into a backbone network (e.g., backbone networks such as ResNet and EfficientNet), and multi-layer image features are extracted as the intermediate representation. In addition, the data of the lidar and the millimeter-wave radar can be directly converted to the BEV space. Subsequently, a spatial self-attention mechanism can be used to extract the required spatial features from the image features, thereby aggregating spatial information; in addition, a temporal self-attention mechanism can be used to fuse historical information, thereby aggregating temporal information.

[0091] Thus, through temporal and spatial fusion, the implicit representation e t can characterize rich temporal and spatial information, thereby further improving the accuracy and safety of decision-making.

[0092] According to some embodiments, the target autonomous driving strategy information may include a target planned trajectory Out1.

[0093] In the related art, an autonomous driving model determines the autonomous driving strategy of an autonomous driving vehicle based on perception information collected by sensors provided on the vehicle. In this process, it is difficult for passengers in the autonomous driving vehicle to interact with the autonomous driving vehicle in a convenient manner. Even if a natural language model capable of communicating with users is deployed on the autonomous driving vehicle, it is difficult to directly consider the information of user interaction in the decision-making process of the autonomous driving strategy.

[0094] To solve the above problems, the present disclosure provides a new autonomous driving model.

[0095] Figure 3 FIG. shows a schematic diagram of another autonomous driving model 300 according to an embodiment of the present disclosure.

[0096] As Figure 3 shown, the autonomous driving model 300 includes a generator 310. The generator 310 is configured to generate target autonomous driving policy information and target interaction information based on the input first input information. Among them, the first input information is related to the navigation information of the vehicle and the perception information of the surrounding environment of the vehicle and includes historical interaction information.

[0097] Using the autonomous driving model provided by the embodiment of the present disclosure, while obtaining the autonomous driving policy information by using the input perception information, the interaction information can also be output, so that the user can interact with the autonomous driving model through natural language expression, improving the user experience of the passengers during the autonomous driving process. The autonomous driving policy can be adjusted according to the interaction with the user, and the user's trust can be increased by providing the user with system response information or autonomous driving policy related information.

[0098] The principle of the present disclosure will be described in detail below.

[0099] As Figure 3 shown, the first input information as the input of the generator 310 may include historical interaction information 301 and information 302 related to the navigation information of the vehicle and the perception information of the surrounding environment of the vehicle.

[0100] The historical interaction information 301 may include the interaction content between the user (such as the passenger of the autonomous driving vehicle) and the model during the autonomous driving process, including the instructions issued by the user and the replies made by the model. The historical interaction information 301 may be generated based on the natural language interaction between the user and the model during the autonomous driving process. The natural language interaction history between the user and the model can be appropriately processed to obtain the form of information that the generator can process, such as the vector c 1 ……c N corresponding to at least a part of the content in the natural language information.

[0101] The information 302 is related to the perception information of the surrounding environment of the vehicle equipped with the autonomous driving model. In some embodiments, the information 302 may be determined based on the encoding result of the perception information by the multi-modal encoding layer. In this case, the autonomous driving model 300 may further include a multi-modal encoding layer. The multi-modal encoding layer in the autonomous driving model 300 can be implemented by using the multi-modal encoding layer 210 described in combination with Figure 2 description.

[0102] The multimodal encoding layer can be configured to output an implicit representation corresponding to the second input information based on the second input information. The second input information includes the perception information of the vehicle surrounding environment obtained by using sensors, and the perception information includes the current perception information and historical perception information of the vehicle surrounding environment. By using the implicit representation output by the multimodal encoding layer as the input of the generator, the autonomous driving model can generate an output result based on the multimodal perception information.

[0103] The information 302 can be an implicit representation corresponding to the second input information or a variant thereof. For example, the information 302 can be a vectorized form of the implicit representation. As Figure 3 shown, the implicit representation e t can be represented as e 1,t ……e N,t in the form of N vectors.

[0104] In some embodiments, as Figure 3 shown, the first input information may further include a time identifier and a type identifier. For example, the first input information may include the time identifier of the historical interaction information and the time identifier of the implicit representation. Among them, the time identifier of the historical interaction information can be used to represent the time point when the interaction event corresponding to the historical interaction information occurs, and the time identifier of the implicit representation can be used to represent the time point when the perception information corresponding to the implicit representation is collected. Any suitable way can be used to represent the above time identifier. For example, the time identifier can be represented as the relative time to the current moment t (such as represented as the time interval from the current moment t), or the time identifier can be represented as the time value of the absolute clock. The first input information may further include the type identifier of the historical interaction information and the type identifier of the implicit representation. By assigning different values to the type identifier of the historical interaction information and the type identifier of the implicit representation, the model can distinguish which part of the input information belongs to the interaction information and which part belongs to the perception information.

[0105] The generator 310 can be configured to generate the target autonomous driving policy information 303 based on the information 301 and the information 302 in the first input information. Further, the generator 310 can generate the target interaction information 304 based on the target autonomous driving policy information 303. In the case where the historical interaction information is natural language information, the target interaction information can also be natural language information.

[0106] The generator 310 can be implemented using any suitable generative model. In some examples, the generator 310 can be implemented as a transformer network structure. For example, the generator 310 can adopt an encoder-decoder integrated transformer network structure, where both encoding and decoding use a unidirectional self-attention mechanism. When generating an output vector using the generator 310, the previously generated vector can be used as an input to generate the next vector. For example, after the generator 310 outputs the target autonomous driving policy information 303, the target autonomous driving policy information 303 can be used as an input to generate subsequent outputs. Using this method, interactive content can be generated based on the generated autonomous driving policy information, making the interaction between the autonomous driving model and the user relevant to the current driving behavior.

[0107] In some embodiments, the target interaction information 304 can include an interaction feedback flag g t . When the interaction feedback flag g t is true (such as 1), the target interaction information further includes interaction content r t,1 , etc., and the decoding of the target interaction information will use special symbols (such as <e>) as the end - bit flag. Among them, the interaction content can be natural - language information. And when the interaction feedback identifier g t is false (such as 0), the decoding process of the target interaction information for this time will end, and the target interaction information will not include specific interaction content. Using the above - mentioned method, the autonomous - driving model provided by the embodiments of the present disclosure can identify the timing when interaction content needs to be output and adaptively output the target interaction information when needed. Thus, a single model can be used to simultaneously implement the decision - making of the autonomous - driving strategy and the interaction with the user, and the historical interaction content between the autonomous - driving model and the user will help the autonomous - driving model to implement the decision - making of the autonomous - driving strategy. Therefore, the autonomous - driving strategy information can be provided as feedback for the user interaction. In some examples, the interaction content output by the autonomous - driving model can include feedback on the user's instructions, such as clarifying and explaining the driving behavior of the autonomous vehicle.

[0108] After the autonomous - driving model outputs the interaction content, the interaction content can be presented to the user in various ways, such as presenting the content output by the autonomous - driving model to the user through voice, image, text, etc. The feedback of the interaction content can be provided to the user in the form of a system response message.

[0109] Figure 4 shows an exemplary flowchart of an autonomous - driving method according to an embodiment of the present disclosure. The autonomous - driving method shown in Figure 3 can be implemented by using the autonomous - driving model described in Figure 4 .

[0110] In step S402, the first input information can be obtained. Among them, the first input information is related to the perception information of the vehicle's surrounding environment and includes historical interaction information.

[0111] In step S404, the first input information can be input into the generator of the autonomous - driving model to generate the target autonomous - driving strategy information and the target interaction information.

[0112] Using the autonomous - driving method provided by the embodiments of the present disclosure, while obtaining the autonomous - driving strategy information by using the input perception information, interaction information can also be output, so that the user can interact with the autonomous - driving model, improving the user experience of passengers during the autonomous - driving process.

[0113] In some embodiments, the autonomous driving model used in the autonomous driving method may further include a multi-modal encoding layer. The multi-modal encoding layer may be configured to output an implicit representation corresponding to the second input information based on the input second input information. The second input information includes the perception information of the vehicle surrounding environment obtained by using sensors. Among them, the first input information includes the implicit representation corresponding to the second input information. In some embodiments, the perception information may include the current perception information and historical perception information of the vehicle surrounding environment during the driving process of the target vehicle.

[0114] By using the implicit representation output by the multi-modal encoding layer as the input of the generator, the autonomous driving model can generate an output result based on the multi-modal perception information. In some implementation manners, the first input information may further include the time identifier of the historical interaction information, the type identifier of the historical interaction information, the time identifier of the implicit representation, and the type identifier of the implicit representation. By assigning different values to the type identifier of the historical interaction information and the type identifier of the implicit representation, the model can distinguish which part of the input information belongs to the interaction information and which part belongs to the perception information.

[0115] In some embodiments, the target interaction information may include an interaction feedback identifier. When the interaction feedback identifier is true, the target interaction information further includes interaction content. In some implementation manners, the interaction content may include natural language information. By using the above method, the autonomous driving model provided by the embodiments of the present disclosure can identify the timing when interaction content needs to be output and adaptively output the target interaction information when needed. Thus, a single model can be used to simultaneously implement the decision-making of the autonomous driving strategy and the interaction with the user, and the historical interaction content between the autonomous driving model and the user will help the autonomous driving model to implement the decision-making of the autonomous driving strategy.

[0116] In some embodiments, the generator may be configured to generate target autonomous driving strategy information based on the first input information and generate target interaction information based on the target autonomous driving strategy information. By using this method, interaction content can be generated based on the generated autonomous driving strategy information, so that the interaction between the autonomous driving model and the user is related to the current driving behavior.

[0117] Figure 5 An exemplary flowchart of a method for training an autonomous driving model according to an embodiment of the present disclosure is shown. The method described in combination with Figure 5 can be used to train the autonomous driving model described in combination with Figure 3 description.

[0118] In step S502, first sample input information may be obtained. The first sample input information is related to the perception information of the vehicle surrounding environment and includes sample historical interaction information.

[0119] In step S504, the true autonomous driving policy information and the true interaction information corresponding to the first sample input information can be obtained.

[0120] In step S506, the first sample input information can be input into the generator of the autonomous driving model to generate predicted autonomous driving policy information and predicted interaction information.

[0121] In step S508, the parameters of the autonomous driving model can be adjusted based on the difference between the predicted autonomous driving policy information and the true autonomous driving policy information and the difference between the predicted interaction information and the true interaction information.

[0122] Using the training method of the embodiments of the present disclosure, the losses of both the autonomous driving policy information and the interaction information can be considered simultaneously, so that the autonomous driving model can learn the ability to output the autonomous driving policy and the interaction information.

[0123] In some embodiments, the autonomous driving model further includes a multi-modal encoder, and the multi-modal encoding layer is configured to output a sample implicit representation corresponding to the second sample input information based on the input second sample input information. The second sample input information includes the perception information of the vehicle surrounding environment obtained by using sensors. Among them, the first sample input information includes the sample implicit representation corresponding to the second sample input information. In some implementation manners, the first sample input information may further include the time identifier of the sample historical interaction information, the type identifier of the sample historical interaction information, the time identifier of the sample implicit representation, and the type identifier of the sample implicit representation.

[0124] In some embodiments, the generator can be configured to generate predicted autonomous driving policy information based on the first sample input information and generate predicted interaction information based on the predicted autonomous driving policy information.

[0125] In some embodiments, the true autonomous driving policy information may include true manual driving policy information, and the true interaction information may include true manual interaction information. The above true manual driving policy information (such as the vehicle trajectory and control data of manual driving ) and the true manual interaction information occurring between the passengers and the driver can be obtained by collecting the true data of human driving. Among them, the true manual interaction information may include a true interaction feedback identifier and true natural language content. At the moment when the driver's conversation occurs, the value of the interaction feedback identifier is determined to be true (such as 1), and at other moments, the value of the interaction feedback identifier is determined to be false (such as 0).

[0126] Step S508 may include: adjusting the parameters of the autonomous driving model by means of supervised learning based on the first difference between the predicted autonomous driving policy information and the real manual driving policy information and the second difference between the predicted interaction information and the real manual interaction information.

[0127] The objective function for supervised learning can be determined based on Equation (1):

[0128]

[0129] where L2 is the objective function of supervised learning. D represents a function for determining the difference between two variables, such as the mean squared error function, cross-entropy function, KL divergence function, etc. y t represents the predicted autonomous driving policy information output by the autonomous driving model at time t, represents the real manual driving policy information at time t, g t represents the interaction feedback identifier output by the autonomous driving model at time t, represents the value of the interaction feedback identifier in the real data at time t, r t represents the predicted natural language content in the predicted interaction information output by the autonomous driving model at time t, represents the real natural language content in the real manual interaction information at time t.

[0130] In some embodiments, the real autonomous driving policy information may include real model driving policy information, and the real interaction information may include real model interaction information.

[0131] After training the autonomous driving model by means of supervised learning, the trained autonomous driving model can be deployed on the vehicle, and further data can be collected during the on-vehicle test of the autonomous driving model, and the model can be further reinforced by using the data generated by the model. The data collected during the on-vehicle test phase of the model may include the real model driving policy information and real model interaction information output by the model. The first evaluation feedback information for the real model driving policy information and the second evaluation feedback information for the real model interaction information can be obtained. The first evaluation feedback information and the second evaluation feedback information can be obtained by means of manual evaluation or by using the trained evaluation model.

[0132] Step S508 may include: adjusting the parameters of the autonomous driving model by means of reinforcement learning based on the first evaluation feedback information, the second evaluation feedback information, the third difference between the predicted autonomous driving policy information and the real model driving policy information, and the fourth difference between the predicted interaction information and the real model interaction information.

[0133] The objective function for supervised learning can be determined based on Equation (2):

[0134]

[0135] Among them, L2 is the objective function of reinforcement learning. D represents a function for determining the difference between two variables, such as the mean squared error function, cross-entropy function, KL divergence function, etc. y t represents the predicted autonomous driving policy information output by the autonomous driving model at time t, represents the true model driving policy information at time t, r t represents the predicted natural language content in the predicted interaction information output by the autonomous driving model at time t, represents the true natural language content in the true model interaction information at time t. represents the advantage function calculated based on the first evaluation feedback information at time t, represents the advantage function calculated based on the second evaluation feedback information at time t.

[0136] Figure 6 Shows an exemplary block diagram of an autonomous driving device based on an autonomous driving model according to an embodiment of the present disclosure.

[0137] As Figure 6 shown, the device 600 includes an input information acquisition unit 610 and a generation unit 620.

[0138] The input information acquisition unit 610 may be configured to acquire first input information, where the first input information is related to the perception information of the vehicle surrounding environment and includes historical interaction information.

[0139] The generation unit 620 is configured to input the first input information into the generator of the autonomous driving model to generate target autonomous driving policy information and target interaction information.

[0140] It should be understood that Figure 6 each module or unit of the device 600 shown in Figure 4 may correspond to each step in the method 400 described with reference to

[0141] Figure 7 Shows an exemplary block diagram of a training device for an autonomous driving model according to an embodiment of the present disclosure.

[0142] As Figure 7 shown, the device 700 includes a sample information acquisition unit 710, a true information acquisition unit 720, a generator training unit 730, and a parameter adjustment unit 740.

[0143] The sample information acquisition unit 710 is configured to acquire first sample input information, where the sample input information is related to the perception information of the vehicle surrounding environment and includes sample historical interaction information.

[0144] The ground truth information acquisition unit 720 is configured to acquire the ground truth autonomous driving policy information and the ground truth interaction information corresponding to the first sample input information;

[0145] The generator training unit 730 is configured to input the first sample input information into the generator of the autonomous driving model to generate predicted autonomous driving policy information and predicted interaction information; and

[0146] The parameter adjustment unit 740 is configured to adjust the parameters of the autonomous driving model based on the difference between the predicted autonomous driving policy information and the ground truth autonomous driving policy information and the difference between the predicted interaction information and the ground truth interaction information.

[0147] It should be understood that Figure 7 each module or unit of the device 700 shown in Figure 5 can correspond to each step in the method 500 described with reference to

[0148] Accordingly, the operations, features, and advantages described above for the method 500 also apply to the device 700 and its included modules and units. For the sake of brevity, certain operations, features, and advantages are not described herein again.

[0149] It should also be understood that various technologies can be described herein in the general context of software-hardware elements or program modules. The above regarding Figure 6 , Figure 7 The described units can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of units 610 to 620 and units 710 to 740 can be implemented together in a System on Chip (SoC). The SoC can include an integrated circuit chip (which includes one or more components such as a processor (e.g., a Central Processing Unit (CPU), a microcontroller, a microprocessor, a Digital Signal Processor (DSP), etc.), a memory, one or more communication interfaces, and / or other circuits), and can optionally execute the received program code and / or include embedded firmware to perform functions.

[0150] According to another aspect of the present disclosure, there is also provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute a method for training an autonomous driving model according to an embodiment of the present disclosure.

[0151] According to another aspect of the present disclosure, there is also provided a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a method for training an autonomous driving model according to an embodiment of the present disclosure.

[0152] According to another aspect of the present disclosure, there is also provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements a method for training an autonomous driving model according to an embodiment of the present disclosure.

[0153] According to another aspect of the present disclosure, there is also provided an autonomous vehicle including: a device for training an autonomous driving model as described above or an electronic device as described above.

[0154] In the technical solution of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0155] Reference Figure 8 , a block diagram of an electronic device 800 that can be a server or a client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0156] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0157] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 can be any type of device that can input information into the electronic device 800. The input unit 806 can receive input digital or character information, and generate key signal inputs related to the user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 807 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 808 can include, but is not limited to, magnetic disks and optical disks. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0158] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as method 400, 500. For example, in some embodiments, methods 400, 500 can be implemented as computer software programs tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the methods 400, 500 described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute methods 400, 500 in any other suitable manner (e.g., by means of firmware).

[0159] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0160] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0161] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0162] For purposes of providing an interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0163] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0164] A computer system may include a client and a server. The client and the server are generally far away from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0165] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0166] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in this disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after this disclosure.< / e>

Claims

1. An autonomous driving model system, the system comprises: A generator, the generator includes a Transformer network, and the Transformer network is configured to process the input first input information to generate target autonomous driving policy information and target interaction information, wherein the first input information is related to the perception information of the vehicle surrounding environment and includes historical natural language interaction information, wherein, the autonomous driving model system further includes a multi-modal encoding layer, and the multi-modal encoding layer is configured to output an implicit representation corresponding to the second input information based on the input second input information, and the second input information includes the perception information of the vehicle surrounding environment obtained by using sensors, wherein, the first input information includes the implicit representation corresponding to the second input information, and the first input information further includes the time identifier of the historical natural language interaction information, the type identifier of the historical natural language interaction information, the time identifier of the implicit representation, and the type identifier of the implicit representation; wherein, the interaction information includes natural language interaction between the user and the autonomous driving model system, and the target interaction information includes an interaction feedback identifier. When the interaction feedback identifier is true, the target interaction information further includes interaction content. When the interaction feedback identifier is false, the target interaction information does not include interaction content.

2. The autonomous driving model system according to claim 1, wherein the interaction content includes natural language information.

3. The autonomous driving model system according to claim 1, wherein, the generator is configured to generate target autonomous driving policy information based on the first input information, and generate the target interaction information based on the target autonomous driving policy information.

4. The autonomous driving model system according to claim 1, wherein, the perception information includes current perception information and historical perception information of the vehicle surrounding environment during the driving of the target vehicle.

5. An autonomous driving method implemented by using an autonomous driving model, the autonomous driving model includes a generator, and the generator includes a Transformer network, the method comprises: Obtaining first input information, wherein the first input information is related to the perception information of the vehicle surrounding environment and includes historical natural language interaction information; and Inputting the first input information into the generator of the autonomous driving model, so that the Transformer network processes the first input information to generate target autonomous driving policy information and target interaction information, wherein, the autonomous driving model further includes a multi-modal encoding layer, and the multi-modal encoding layer is configured to output an implicit representation corresponding to the second input information based on the input second input information, and the second input information includes the perception information of the vehicle surrounding environment obtained by using sensors, Wherein, the first input information includes an implicit representation corresponding to the second input information, and the first input information further includes a time identifier of the historical natural language interaction information, a type identifier of the historical natural language interaction information, a time identifier of the implicit representation, and a type identifier of the implicit representation; Wherein, the interaction information includes natural language interaction between a user and the automatic driving model, and the target interaction information includes an interaction feedback identifier. When the interaction feedback identifier is true, the target interaction information further includes interaction content. When the interaction feedback identifier is false, the target interaction information does not include interaction content.

6. The automatic driving method according to claim 5, wherein the interaction content includes natural language information.

7. The automatic driving method according to claim 5, Wherein, The generator is configured to generate target automatic driving strategy information based on the first input information and generate the target interaction information based on the target automatic driving strategy information.

8. The automatic driving method according to claim 5, Wherein, The perception information includes current perception information and historical perception information about the surrounding environment of the target vehicle during the driving process of the target vehicle.

9. A training method for an automatic driving model, the automatic driving model includes a generator, the generator includes a Transformer network, and the method includes: Obtain first sample input information, where the first sample input information is related to perception information about the surrounding environment of the vehicle and includes sample historical natural language interaction information; Obtain true automatic driving strategy information and true interaction information corresponding to the first sample input information; Input the first sample input information into the generator of the automatic driving model, so that the Transformer network processes the first sample input information to generate predicted automatic driving strategy information and predicted interaction information; And Adjust the parameters of the automatic driving model based on the difference between the predicted automatic driving strategy information and the true automatic driving strategy information and the difference between the predicted interaction information and the true interaction information, Wherein, the automatic driving model further includes a multi-modal encoding layer, and the multi-modal encoding layer is configured to output a sample implicit representation corresponding to the second sample input information based on the input second sample input information. The second sample input information includes perception information about the surrounding environment of the vehicle obtained by using sensors. Wherein, the first sample input information includes a sample implicit representation corresponding to the second sample input information, and the first sample input information further includes a time identifier of the sample historical natural language interaction information, a type identifier of the sample historical natural language interaction information, a time identifier of the sample implicit representation, and a type identifier of the sample implicit representation; Among them, the interaction information includes natural language interaction between the user and the autonomous driving model. The predicted interaction information includes an interaction feedback identifier. When the interaction feedback identifier is true, the predicted interaction information further includes interaction content. When the interaction feedback identifier is false, the predicted interaction information does not include interaction content.

10. The method according to claim 9, wherein, the generator is configured to generate predicted autonomous driving strategy information based on the first sample input information, and generate the predicted interaction information based on the predicted autonomous driving strategy information.

11. The method according to claim 9, wherein, the true autonomous driving strategy information includes true manual driving strategy information, and the true interaction information includes true manual interaction information. Adjusting the parameters of the autonomous driving model based on a first difference between the predicted autonomous driving strategy information and the true autonomous driving strategy information and a second difference between the predicted interaction information and the true interaction information includes: Using a supervised learning method, adjusting the parameters of the autonomous driving model based on a first difference between the predicted autonomous driving strategy information and the true manual driving strategy information and a second difference between the predicted interaction information and the true manual interaction information.

12. The method according to any one of claims 9-11, wherein the true autonomous driving strategy information includes true model driving strategy information, the true interaction information includes true model interaction information, and the method further includes: Obtaining first evaluation feedback information for the true model driving strategy information and second evaluation feedback information for the true model interaction information; Using a reinforcement learning method, adjusting the parameters of the autonomous driving model based on the first evaluation feedback information, the second evaluation feedback information, a third difference between the predicted autonomous driving strategy information and the true model driving strategy information, and a fourth difference between the predicted interaction information and the true model interaction information.

13. An autonomous driving device based on an autonomous driving model, comprising: An input information acquisition unit configured to acquire first input information, wherein the first input information is related to perception information of the vehicle surrounding environment and includes historical natural language interaction information; and A generation unit configured to input the first input information into a generator of the autonomous driving model to generate target autonomous driving strategy information and target interaction information, wherein the generator includes a Transformer network, and the Transformer network processes the first input information to generate target autonomous driving strategy information and target interaction information. Among them, the autonomous driving model further includes a multi-modal encoding layer, and the multi-modal encoding layer is configured to output an implicit representation corresponding to the second input information based on the input second input information. The second input information includes perception information of the vehicle surrounding environment obtained by using sensors. Among them, the first input information includes an implicit representation corresponding to the second input information, and the first input information further includes a time identifier of the historical natural language interaction information, a type identifier of the historical natural language interaction information, a time identifier of the implicit representation, and a type identifier of the implicit representation; Among them, the interaction information includes natural language interaction between the user and the autonomous driving model, and the target interaction information includes an interaction feedback identifier. When the interaction feedback identifier is true, the target interaction information further includes interaction content. When the interaction feedback identifier is false, the target interaction information does not include interaction content.

14. A training device for an autonomous driving model, the autonomous driving model includes a multi-modal encoding layer and a generator, the generator includes a Transformer network, and the training device includes: A sample information acquisition unit configured to acquire first sample input information, where the sample input information is related to the perception information of the vehicle surrounding environment and includes sample historical natural language interaction information; A true information acquisition unit configured to acquire true autonomous driving policy information and true interaction information corresponding to the first sample input information; A generator training unit configured to input the first sample input information into the generator of the autonomous driving model, so that the Transformer network processes the first sample input information to generate predicted autonomous driving policy information and predicted interaction information; and A parameter adjustment unit configured to adjust the parameters of the autonomous driving model based on the difference between the predicted autonomous driving policy information and the true autonomous driving policy information and the difference between the predicted interaction information and the true interaction information, wherein the autonomous driving model further includes a multi-modal encoding layer, and the multi-modal encoding layer is configured to output a sample implicit representation corresponding to the second sample input information based on the input second sample input information, and the second sample input information includes the perception information of the vehicle surrounding environment obtained by using sensors, wherein the first sample input information includes a sample implicit representation corresponding to the second sample input information, and the first sample input information further includes a time identifier of the sample historical natural language interaction information, a type identifier of the sample historical natural language interaction information, a time identifier of the sample implicit representation, and a type identifier of the sample implicit representation; wherein the interaction information includes natural language interaction between the user and the autonomous driving model, and the predicted interaction information includes an interaction feedback identifier. When the interaction feedback identifier is true, the predicted interaction information further includes interaction content. When the interaction feedback identifier is false, the predicted interaction information does not include interaction content.

15. An electronic device, including: At least one processor; and A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of claims 5-12.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein, the computer instructions are for causing the computer to execute the method according to any one of claims 5-12.

17. A computer program product comprising a computer program, wherein, the computer program, when executed by a processor, implements the method according to any one of claims 5-12.

18. An autonomous vehicle, comprising: one of the autonomous driving device according to claim 13, the training device of the autonomous driving model according to claim 14, and the electronic device according to claim 15.

Citation Information

Patent Citations

  • Automatic driving active-type interaction vehicle-mounted system, equipment and method, and storing medium

    CN110103989A

  • Automatic driving control method and device, system, electronic equipment and storage medium

    CN110837258A

  • Intelligent vehicle track prediction system and method fusing peripheral vehicle interaction information

    CN113954864A

  • Automatic driving method, system and equipment of vehicle and storage medium

    CN115578876A