Control Signal Acquisition Method and Training Method for Autonomous Driving Model

By determining the encoding characteristics of multiple road levels and nodes and generating autonomous driving control signals, the problem of strategy decision optimization in complex environments in the prior art is solved, and end-to-end data-driven autonomous driving is realized.

CN118397862BActive Publication Date: 2025-07-22BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410494314.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-23
Publication Date
2025-07-22
Estimated Expiration
2044-04-23

AI Technical Summary

Technical Problem

Existing autonomous driving algorithms are difficult to effectively optimize strategy decisions in complex environments, which makes it difficult to maintain the system, poor algorithm applicability, and difficult to adapt to various scenarios.

Method used

By determining multiple road levels and road nodes, encoding is performed based on the correlation degree of vehicle planned routes, and autonomous driving control signals are generated using preset road query characteristics and road encoding characteristics to realize end-to-end data-driven autonomous driving.

Benefits of technology

It improves the effectiveness and prediction capabilities of comprehensive road characteristics, and can generate autonomous driving control signals that control vehicles to drive along planned routes, realizing end-to-end data-driven autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118397862B_ABST
    Figure CN118397862B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method for obtaining a control signal and a method for training an autonomous driving model, which relate to the field of artificial intelligence technology, specifically to technical fields such as computer vision and deep learning, and can be applied to scenarios such as autonomous driving. The method for obtaining a control signal includes: determining a plurality of road levels and the road nodes of each of the plurality of road levels, where the plurality of road levels have different degrees of association with the planned route of the vehicle, and the road nodes represent a section of road located within the same lane; encoding the position information of the road nodes of each of the plurality of road levels to obtain road encoding features of each of the plurality of road levels; obtaining a comprehensive road feature based on a preset road query feature and the road encoding features of each of the plurality of road levels; and generating an autonomous driving control signal based on the comprehensive road feature.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as computer vision and deep learning, and can be applied to scenarios such as autonomous driving. In particular, it relates to a method for obtaining control signals for autonomous driving, a method for training an autonomous driving model, a device for obtaining control signals for autonomous driving, a device for training an autonomous driving model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies how to make a computer simulate certain human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0003] Autonomous driving technology aims to achieve safe driving of vehicles without human operation by integrating advanced perception, decision-making, and control systems. The core technologies include but are not limited to sensor technology, machine vision, radar, lidar, and positioning, etc. These technologies work together on the vehicle's environmental perception ability to accurately identify surrounding objects, road signs, and traffic conditions. In addition, the autonomous driving system also relies on advanced algorithms and artificial intelligence technologies, such as machine learning and deep learning, to process a large amount of data and make fast and accurate driving decisions. The control system then converts these decisions into actions to precisely control the speed, direction, and path of the vehicle.

[0004] The methods described in this section are not necessarily methods that have been previously envisioned or adopted. Unless otherwise specified, no method described in this section should be considered prior art solely because it is included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to have been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides a method for obtaining control signals for autonomous driving, a method for training an autonomous driving model, a device for obtaining control signals for autonomous driving, a device for training an autonomous driving model, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, a method for obtaining control signals for autonomous driving is provided, including: determining a plurality of road levels and road nodes of each of the plurality of road levels, where the plurality of road levels have different degrees of association with the planned route of the vehicle, and the road nodes represent a section of road within the same lane; encoding the position information of the road nodes of each of the plurality of road levels to obtain road encoding features of each of the plurality of road levels; obtaining a comprehensive road feature based on a preset road query feature and the road encoding features of each of the plurality of road levels; and generating an autonomous driving control signal based on the comprehensive road feature.

[0007] According to another aspect of the present disclosure, a method for training an autonomous driving model is provided. The autonomous driving model includes a road encoding model, a plurality of road level models corresponding to a plurality of road levels, and a control signal generation model. The training method includes: determining sample road nodes of each of the plurality of road levels, where the sample road nodes represent a section of road within the same lane; using the road encoding model to encode the position information of the sample road nodes of each of the plurality of road levels to obtain road encoding features of each of the plurality of road levels; using the plurality of road level models to obtain a comprehensive road feature based on an initial road query feature and the road encoding features of each of the plurality of road levels; using the control signal generation model to generate an initial autonomous driving control signal based on the comprehensive road feature; and training the autonomous driving model based on the initial autonomous driving control signal and a true control signal to obtain a target autonomous driving model.

[0008] According to another aspect of the present disclosure, a device for obtaining control signals for autonomous driving is provided, including: a first road level determination unit configured to determine a plurality of road levels and road nodes of each of the plurality of road levels, where the plurality of road levels have different degrees of association with the planned route of the vehicle, and the road nodes represent a section of road within the same lane; a first road encoding unit configured to encode the position information of the road nodes of each of the plurality of road levels to obtain road encoding features of each of the plurality of road levels; a first feature extraction unit configured to obtain a comprehensive road feature based on a preset road query feature and the road encoding features of each of the plurality of road levels; and a first control signal generation unit configured to generate an autonomous driving control signal based on the comprehensive road feature.

[0009] According to another aspect of the present disclosure, there is provided a training device for an autonomous driving model, where the autonomous driving model includes a road encoding model, a plurality of road level models corresponding to a plurality of road levels, and a control signal generation model. The training device includes: a second road level determination unit configured to determine sample road nodes for each of the plurality of road levels, where the sample road nodes represent a section of road located within the same lane; a second road encoding unit configured to encode the position information of the sample road nodes for each of the plurality of road levels using the road encoding model to obtain road encoding features for each of the plurality of road levels; a second feature extraction unit configured to use the plurality of road level models to obtain a comprehensive road feature based on an initial road query feature and the road encoding features for each of the plurality of road levels; a second control signal generation unit configured to generate an initial autonomous driving control signal using the control signal generation model based on the comprehensive road feature; and a training unit configured to train the autonomous driving model based on the initial autonomous driving control signal and a true control signal to obtain a target autonomous driving model.

[0010] According to another aspect of the present disclosure, there is provided an electronic device including: at least one processor; and a memory communicatively connected to the at least one processor; where the memory stores instructions executable by the at least one processor, and when executed by the at least one processor, the instructions enable the at least one processor to execute the above method.

[0011] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the above method.

[0012] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where when the computer program is executed by a processor, the above method is implemented.

[0013] According to one or more embodiments of the present disclosure, by determining a plurality of road levels and corresponding road nodes according to different degrees of association with the planned route of the vehicle, and encoding their position information respectively, it is possible to effectively extract information about the planned route of the vehicle and other relevant roads near the vehicle. And by using a preset road query feature and the road encoding features of a plurality of road levels, the effectiveness and predictive ability of the obtained comprehensive road feature can be further improved. Furthermore, the autonomous driving model can generate an autonomous driving control signal for controlling the vehicle to follow the planned route based on the comprehensive road feature, realizing end-to-end data-driven autonomous driving.

[0014] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0015] The drawings exemplarily illustrate embodiments and form part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0016] Figure 1 A schematic diagram showing an exemplary system in which various methods described herein can be implemented according to an embodiment of the present disclosure;

[0017] Figure 2 A flowchart of a method for obtaining a control signal for autonomous driving according to an exemplary embodiment of the present disclosure is disclosed;

[0018] Figure 3 A schematic diagram showing a map according to an exemplary embodiment of the present disclosure;

[0019] Figure 4A A schematic diagram showing the determination of multiple road levels according to an exemplary embodiment of the present disclosure;

[0020] Figure 4B A schematic diagram showing the determination of a path within the field of view of a vehicle according to an exemplary embodiment of the present disclosure;

[0021] Figure 5 A flowchart showing the process of encoding multiple road levels according to an exemplary embodiment of the present disclosure;

[0022] Figure 6 A flowchart showing the process of encoding at least one path according to an exemplary embodiment of the present disclosure;

[0023] Figure 7 A schematic diagram showing a road encoding model according to an exemplary embodiment of the present disclosure;

[0024] Figure 8 A flowchart showing the process of feature processing for road query features and multiple road levels according to an exemplary embodiment of the present disclosure;

[0025] Figure 9 A flowchart showing the process of feature processing for the semantic features of the first road level and the remaining at least one road level according to an exemplary embodiment of the present disclosure;

[0026] Figure 10 A schematic diagram showing a road level model according to an exemplary embodiment of the present disclosure;

[0027] Figure 11Shows a schematic diagram of an autonomous driving model according to an exemplary embodiment of the present disclosure;

[0028] Figure 12 Shows a flowchart of a method for training an autonomous driving model according to an exemplary embodiment of the present disclosure;

[0029] Figure 13 Shows a flowchart of a process for encoding multiple road levels according to an exemplary embodiment of the present disclosure;

[0030] Figure 14 Shows a flowchart of a process for encoding at least one sample path according to an exemplary embodiment of the present disclosure;

[0031] Figure 15 Shows a flowchart of a process for feature processing of an initial road query feature and multiple road levels according to an exemplary embodiment of the present disclosure;

[0032] Figure 16 Shows a flowchart of a process for feature processing of the semantic feature of the first road level and the remaining at least one road level according to an exemplary embodiment of the present disclosure;

[0033] Figure 17 Discloses a structural block diagram of a control signal acquisition device for autonomous driving according to an exemplary embodiment of the present disclosure;

[0034] Figure 18 Discloses a structural block diagram of a training device for an autonomous driving model according to an exemplary embodiment of the present disclosure; and

[0035] Figure 19 Shows a structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure. Detailed Description of the Invention

[0036] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0037] In the present disclosure, unless otherwise specified, the terms "first", "second", etc. are used to describe various elements and are not intended to limit the positional relationship, timing relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, and in certain cases, based on the context description, they may also refer to different instances.

[0038] In the description of the various examples in this disclosure, the terms used are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically defined, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any and all possible combinations of the listed items.

[0039] In the related art, most of the existing control algorithms are algorithm-driven based on rule logic. These algorithms often require a large number of strategies for decision optimization when facing complex environments. When optimizing algorithms in an infinite environment, they often face problems such as difficult strategy development, difficult system maintenance, and difficulty in applying algorithms to various scenarios.

[0040] To solve the above problems, in this disclosure, multiple road levels and corresponding road nodes are determined according to different degrees of association with the planned route of the vehicle, and their location information is encoded respectively, so that information on the planned route of the vehicle and other relevant roads near the vehicle can be effectively extracted. By using the preset road query features and the road coding features of multiple road levels, the effectiveness and prediction ability of the obtained comprehensive road features can be further improved. Furthermore, the autonomous driving model can generate an autonomous driving control signal for controlling the vehicle to follow the planned route based on the comprehensive road features, realizing end-to-end data-driven autonomous driving.

[0041] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0042] Figure 1 FIG. shows a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein can be implemented according to an embodiment of the present disclosure. Referring Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more application programs.

[0043] In an embodiment of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of the methods of the present disclosure.

[0044] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual environments and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) network.

[0045] In Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or a combination thereof that may be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may be different from system 100. Thus, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to be limiting.

[0046] Users may use client devices 101, 102, 103, 104, 105, and / or 106 to perform human-computer interactions. The client device may provide an interface that enables the user of the client device to interact with the client device. The client device may also output information to the user via this interface. Although Figure 1 only six client devices are depicted, those skilled in the art will be able to understand that the present disclosure may support any number of client devices.

[0047] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computing devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices, etc. These computing devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT WindowsMobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, Internet-enabled gaming devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0048] Network 110 may be any type of network known to those skilled in the art, and it may support data communication using any one of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, token ring, wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (such as Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0049] Server 120 may include one or more general-purpose computers, dedicated server computers (such as PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for the server). In various embodiments, server 120 may run one or more services or software applications that provide the functions described below.

[0050] The computing units in server 120 can run one or more operating systems including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 can also run any one of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0051] In some embodiments, server 120 can include one or more applications to analyze and combine data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 can also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.

[0052] In some embodiments, server 120 can be a server of a distributed system or a server incorporating a blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services.

[0053] System 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of databases 130 can be used to store information such as audio files and video files. Databases 130 can reside in various locations. For example, the data repositories used by server 120 can be local to server 120, or can be remote from server 120 and can communicate with server 120 via a network-based or dedicated connection. Databases 130 can be of different types. In certain embodiments, the databases used by server 120 can be databases such as relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.

[0054] In certain embodiments, one or more of databases 130 can also be used by applications to store application data. The databases used by applications can be different types of databases, such as key-value repositories, object repositories, or conventional repositories supported by a file system.

[0055] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and devices described in this disclosure.

[0056] According to one aspect of the present disclosure, a method for obtaining control signals for autonomous driving is provided. Figure 2 The flowchart of the control signal acquisition method 200 according to an exemplary embodiment of the present disclosure is disclosed. As Figure 2 shown, the control signal acquisition method 200 includes: step S201, determining a plurality of road levels and the road nodes of each of the plurality of road levels, the plurality of road levels having different degrees of association with the planned route of the vehicle, and the road nodes representing a section of road located within the same lane; step S202, encoding the position information of the road nodes of each of the plurality of road levels to obtain the road coding features of each of the plurality of road levels; step S203, obtaining the comprehensive road features based on a preset road query feature and the road coding features of each of the plurality of road levels; and step S204, generating an autonomous driving control signal based on the comprehensive road features.

[0057] Thus, by determining a plurality of road levels and the corresponding road nodes according to different degrees of association with the planned route of the vehicle and encoding their position information respectively, it is possible to effectively extract information on the planned route of the vehicle and other relevant roads near the vehicle, and by using the preset road query feature and the road coding features of the plurality of road levels, the effectiveness and prediction ability of the obtained comprehensive road features can be further improved, so that the autonomous driving model can generate an autonomous driving control signal for controlling the vehicle to follow the planned route based on the comprehensive road features, realizing end-to-end data-driven autonomous driving.

[0058] Figure 3 A schematic diagram of a map according to an exemplary embodiment of the present disclosure is shown. The map used in the present disclosure can be a lane-level map or other maps including lane information. In the map, a section of road within the same lane can be expressed as a road node, and the connection relationship between two roads can be expressed by a directed edge. As Figure 3 shown, if the connection between two roads is a dotted line, then these two roads can be jointly adjusted, and at the same time, the successor roads of each road can be obtained according to the direction of the road.

[0059] In some embodiments, the planned route of the vehicle can be determined before or during the execution of step S201. The planned route can be, for example, the optimal path or the shortest path from the starting point to the ending point. In an exemplary embodiment, the planned route can be obtained through a shortest path solving algorithm such as the Dijkstra algorithm.

[0060] In step S201, multiple road levels can be determined, and road nodes of each road level can be determined respectively. Each road level can include at least one road node. Different road levels have different degrees of association with the planned route of the vehicle.

[0061] According to some embodiments, the multiple road levels can include: a planned level, including road nodes covered by the planned route; an alternative level, including road nodes covered by alternative routes having the same end point as the planned route. In other words, it is still possible to drive to the end point starting from the road nodes of the alternative level; and other levels, including other road nodes near the vehicle. The degrees of association of the planned level, the alternative level, and the other levels with the planned route decrease in sequence.

[0062] Thus, in the above manner, the road nodes are divided into recommended roads (planned level), drivable roads (alternative level), and view roads (other levels), so that relevant information of the above roads can be hierarchically extracted by using a road query model subsequently.

[0063] Figure 4A FIG. shows a schematic diagram of determining multiple road levels according to an exemplary embodiment of the present disclosure.

[0064] Among them, node 401 and node 402 respectively represent the starting point and the end point of the vehicle. In this embodiment, all routes from the starting point to the end point can be calculated according to the map connectivity relationship, and the covered road nodes are used as the road nodes 403 of the alternative level. For example, relevant search algorithms such as the depth-first algorithm can be used to determine the road nodes of this road level. Then, the other road nodes can be determined as the road nodes 404 of the other levels. Finally, the best path or the shortest path (i.e., the planned route) from the starting point to the end point can be calculated according to the map connectivity relationship, and the covered road nodes are used as the road nodes 405 of the planned level.

[0065] In some embodiments, the order of determining different road levels can be adjusted. It is also possible to first calculate the planned route (or according to the planned route determined before performing step S201), obtain the road nodes of the planned level, then determine the road nodes of the alternative level, and finally determine the road nodes of the other levels.

[0066] In some embodiments, there is no intersection between different road levels. For example, after determining the road nodes 405 of the above-mentioned planned level, these road nodes no longer belong to the alternative level. That is to say, the road nodes of the alternative level only include the nodes in 406.

[0067] In some embodiments, different road levels may also include the same road nodes. For example, the road nodes 403 of an alternative level may include the road nodes 405 of the planning level, and all the road nodes (403, 404) on the map may be road nodes of other levels.

[0068] It can be understood that multiple road levels with different degrees of association with the planned route of the vehicle can also be determined in other ways, which are not limited herein. In an exemplary embodiment, the multiple road levels may also include a level composed of road nodes covered by the planned route, a level composed of nodes having a connection relationship with the road nodes covered by the planned route, and a level composed of other nodes.

[0069] After obtaining the road nodes of each of the multiple road levels, the position information of these road nodes can be encoded in step S202 to obtain the road coding features of each of the multiple road levels. In some embodiments, the coordinates of the road corresponding to the road node can be used as the position information of the road node.

[0070] Figure 5 The flowchart of process 500 for encoding multiple road levels according to an exemplary embodiment of the present disclosure is shown. Process 500 can be used to implement step S202 in the above method 200. Process 500 includes: step S501, for each of the multiple road levels, based on the road connection relationship of the road level, determine at least one path of the road level, where each path includes one or more road nodes of the road level that can be continuously traveled by the vehicle; and step S502, encode the position information of the one or more road nodes included in each of the at least one path to obtain the path coding features of each of the at least one path, where the road coding features of the road level corresponding to the at least one path include the path coding features of each of the at least one path.

[0071] Since the road nodes included in each road level are scattered, by determining the paths that the vehicle can travel based on the road connection relationship in each road level and encoding the position information of the road nodes on the paths, more effective road coding features can be obtained.

[0072] In an exemplary embodiment, the multiple levels include the above-mentioned planning level, alternative level, and other levels. In step S501, for the alternative level, based on the road connectivity relationship of the alternative level, at least one path can be determined, where each path includes one or more road nodes, and these road nodes can form one or more routes that can be continuously traveled by a vehicle, such as the above-mentioned alternative routes. It should be noted that in the case where the road level only includes one road node, a section of the road corresponding to this road node can also be determined as a path. In other words, the road connectivity relationship can include the connectivity relationship between multiple road nodes (such as Figure 3 the directed edges in Figure 3 ), and can also include the road connectivity relationship represented by a section of the road in a single road node. A path can also include a road node that has a section of the road that can be continuously traveled by a vehicle; a path can also include multiple road nodes, and these multiple road nodes need to satisfy that the roads represented by them can be continuously traveled by a vehicle, that is, can connect these road nodes into a continuous line according to the relationship between the road nodes (such as

[0073] the directed edges in

[0074] ). For the planning level, since the road nodes included in the planning level are the road nodes covered by the planned route, these road nodes can be directly determined as the road nodes included in the paths of the planning level. In this exemplary embodiment, the planning level may only include one path corresponding to the planned route. For other levels, based on the connectivity relationship between the road nodes included in other levels, at least one path can be determined, where each path includes one or more road nodes, and these road nodes can be combined into a route that can be continuously traveled by a vehicle. Although these routes may not lead to the end point, the road information of these routes may be beneficial for the generation of control signals.

[0075] It can be understood that for multiple road levels determined by other means, at least one path for each different road level can also be determined in the above manner or other ways. Figure 3 the directed edges in

[0076] Thus, by determining the paths of each road level within the field of view of the vehicle, the range of road nodes included in the paths can be narrowed, thereby reducing the computational amount and enhancing the information validity of the subsequent obtained path coding features.

[0077] In an exemplary embodiment, the field of view can be set as an area centered on the vehicle with a preset length as the radius, such as an area within 60 meters around the vehicle.

[0078] Figure 4B FIG. shows a schematic diagram of determining a path within the field of view of a vehicle according to an exemplary embodiment of the present disclosure. In Figure 4B it, the three road levels of the planning level, the alternative level, and the other level are respectively represented by white, light gray, and dark gray. The vehicle is located at the black road node 407, and the area 408 can be the field of view of the vehicle. Then, within the field of view of the vehicle, there are three road nodes 407, 409, and 410 of the planning level, one road node 411 of the alternative level, and one road node 412 of the other level. For the planning level, a path can be determined based on the road connection relationship between the road nodes 407, 409, and 410, that is, the path obtained by connecting the road nodes 409→407→410 in sequence. For the alternative level, a path can be determined inside the road node 411. For the other level, a path can be determined inside the road node 412. In some embodiments, the road nodes 407, 409, and 410 can also be road nodes of the alternative level. Then, based on the connection relationship between the road nodes 407, 409, 410, and 411 of the road level, multiple paths can be constructed, including 409→407→410, 409→407→411, 411→407→410, and so on.

[0079] According to some embodiments, the coordinates of multiple road points can be based on a coordinate system with the vehicle's location as the origin and the vehicle head direction as the first direction. The first direction can be, for example, the x direction, and the horizontal direction perpendicular to the first direction can be the second direction or the y direction. In some embodiments, the coordinates of the road points can be three-dimensional space coordinates, and the direction perpendicular to the ground can be the third direction or the z direction.

[0080] Thus, by establishing a coordinate system centered on the vehicle to determine the coordinates of multiple road points, the influence of factors such as different geographical locations can be eliminated, making the subsequent obtained road point coding features independent of the geographical location and more helpful for generating vehicle control signals.

[0081] Figure 6FIG. 0 shows a flowchart of a process 600 for encoding at least one path according to an exemplary embodiment of the present disclosure. The process 600 can be used to implement step S502 in the above process 500. The process 600 includes: step S601, for each path in the at least one path, sampling one or more road nodes included in the path based on a preset number of samples to obtain a plurality of road points; step S602, encoding the coordinates of the plurality of road points to obtain a road point encoding feature of the path; step S603, aggregating the road point encoding features of the at least one path respectively to obtain a road point aggregation feature; and step S604, fusing the road point encoding features of the at least one path respectively with the road point aggregation feature to obtain a path encoding feature of each of the at least one path.

[0082] Thus, by sampling among the road nodes included in each path based on a preset number of samples and encoding the coordinates of the road points, different paths include the same number of road points, which facilitates encoding. And by aggregating the road point encoding features of different paths and then fusing the road point aggregation feature and the road point encoding feature of each path, the path encoding feature of each path includes both the position information of the path itself and the comprehensive information of all paths at the road level, thereby improving the prediction ability of the path encoding feature.

[0083] In some embodiments, step S502 can be performed using a road encoding model to encode at least one path and obtain a path encoding feature of each path. Figure 7 FIG. 8 shows a schematic diagram of a road encoding model 700 according to an exemplary embodiment of the present disclosure. The road encoding model 700 can further include a road point encoding sub-model 710, a road point aggregation sub-model 720, and a road point fusion sub-model 730, which are respectively used to perform step S602 - step S604, as will be described below. For different road levels, a corresponding road encoding model can be designed and trained for each road level to encode the paths at the corresponding road level, or the same road encoding model can be used to encode the paths at each road level, which is not limited herein.

[0084] In some embodiments, in step S601, for each path, sampling can be performed in a road segment corresponding to at least a part of the road nodes included in the path to obtain a plurality of road points. If the number of data points in the road node is less than the preset number of samples, zero padding can be performed. In an exemplary embodiment, the preset number of samples is 240 points.

[0085] In some embodiments, in step S602, the coordinates of multiple road points are encoded to obtain the road point encoding feature of the path. A road point encoder based on a multi-layer perceptron can be used as the road point encoding sub-model to encode the coordinates of multiple road points.

[0086] In some embodiments, the coordinates of multiple road points can be represented as a vector of length p_num * 3, where p_num is the preset sampling number, and 3 represents the values in three directions in the three-dimensional space coordinates. The vector is input into the road point encoder to obtain the road point encoding feature output by the road point encoder. The road point encoding feature can be represented as a vector of length h_dim, for example, where h_dim represents the encoded hidden dimension. In Figure 7 the illustrated embodiment, the size of the coordinates 702 of multiple road points of at least one path can be represented as [N, p_num * 3], where N is the number of paths. It can be understood that Figure 7 the solid rectangles in the road point coordinates 702, road point encoding feature 704, and path encoding feature 708 represent vectors, and the circles therein represent individual elements (coordinate values, feature values, etc.) in the vectors.

[0087] The solid box in the coordinates 702 indicates that each road level has at least one path, that is, N = 1, and the dashed box indicates that each road level may have multiple paths, that is, N > 1. The same applies to the road point encoding feature 704 and path encoding feature 708 hereinafter. The size of the road point encoding feature 704 can be represented as [N, h_dim].

[0088] In some embodiments, in step S603, the road point encoding features of different paths are aggregated to obtain the road point aggregation feature.

[0089] In an exemplary embodiment, a road point aggregation sub-model 720 can be used to pool (e.g., max-pooling) the road point encoding features of at least one path respectively to obtain the road point aggregation feature. After pooling the road point encoding features 704 (with a length of h_dim) of N paths, the obtained road point aggregation feature 706 has a size of [1, h_dim].

[0090] It can be understood that the road point encoding features of at least one path can also be aggregated in other ways to obtain the road point aggregation feature.

[0091] According to some embodiments, step S604, fusing the road point encoding features of each of the at least one path with the road point aggregation feature respectively to obtain the path encoding feature of each of the at least one path may include: splicing the road point encoding features of each of the at least one path with the road point aggregation feature along the hidden dimension to obtain the path encoding feature of each of the at least one path.

[0092] Thus, through the above method, it is possible to originally retain both the position information of the road points in each path and the comprehensive information of all paths corresponding to the road levels in the path encoding feature of each path, further improving the prediction ability of the path encoding feature.

[0093] In an exemplary embodiment, after the road point fusion sub-network 730 splices the road point encoding features of each of the at least one path with the road point aggregation feature along the hidden dimension, the length of the path encoding feature of each path obtained may be h_dim*2, and the size of the path encoding features 708 of the at least one path may be [N, h_dim*2]. Thus, the road encoding feature of the corresponding road level is obtained.

[0094] It can be understood that the road point encoding features of each of the at least one path can also be fused with the road point aggregation feature in other ways to obtain the path encoding feature of each path, thereby obtaining.

[0095] Figure 8 FIG. 800 shows a flowchart of a process 800 for feature processing of a road query feature and multiple road levels according to an exemplary embodiment of the present disclosure. The process 800 can be used to implement step S203 in the above method 200. The process 800 includes: step S801, constructing a query feature of a first road level among multiple road levels based on the road query feature; step S802, constructing a key feature and a value feature of the first road level based on the road encoding feature of the first road level; step S803, performing feature processing on the query feature, key feature, and value feature of the first road level based on the cross-attention mechanism to obtain the semantic feature of the first road level; and step S804, obtaining a road comprehensive feature based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level respectively.

[0096] Thus, by using the cross-attention mechanism, it is possible to effectively extract the road position information in the road encoding features of the first road level. In addition, by constructing the query features of the first road level based on the road query features, and constructing the value features and key features of the first road level based on the road encoding features of the first road level, the semantic features of the first road level obtained after passing through the cross-attention mechanism can maintain the form or size of the road query features, so as to perform further feature processing with the road encoding features of the remaining road levels subsequently.

[0097] The order of the road encoding features of multiple road levels during feature processing can be predetermined or random.

[0098] According to some embodiments, the road query features can be subjected to feature processing based on the cross-attention mechanism with the road encoding features of multiple road levels in ascending order of the degree of association with the planned route. That is to say, the first road level is the road level with the lowest degree of association with the planned route of the vehicle, and the degree of association of at least one subsequent road level with the planned route of the vehicle increases in turn.

[0099] Compared with directly performing feature processing on the features corresponding to the planned route, by first extracting the generalized information of the surrounding roads, gradually narrowing the scope, and finally extracting the information most relevant to the planned route, the difficulty of extracting the information related to the planned route is reduced, and richer road-related information can be obtained, further enhancing the prediction ability of the road comprehensive features.

[0100] In an exemplary embodiment, if the multiple road levels include the above-mentioned planning level, alternative level, and other levels, the first road level in step S801 can be the other level, and then in step S804, the semantic features of the other level and the alternative level and the planning level can be subjected to feature processing to finally obtain the road comprehensive features.

[0101] In some embodiments, the preset road query features can be obtained through training. For example, the initial road query features obtained by random initialization are trained and optimized through the training method 1200 of the autonomous driving model to be introduced below. Through training, the road query features can implicitly contain information that is independent of specific roads or maps but is helpful for the process of jointly performing feature processing with the road encoding features to achieve road information extraction. The road query features can also be determined in advance by heuristic methods, based on experience, or in other ways, which is not limited herein.

[0102] In some embodiments, each road level may have a corresponding cross-attention module, which includes a trained query matrix, key matrix, and value matrix. In step S801, the query matrix in the cross-attention module corresponding to the first road level can be used to map the road query features to the query features of the first road level. In step S802, the key matrix and value matrix in the cross-attention module corresponding to the first road level can be used to map the road encoding features of the first road level to the key features and value features of the first road level. In step S803, a cross-attention matrix can be calculated based on the query features and key features of the first road level, and a corresponding output result, i.e., the semantic features of the first road level, can be calculated based on the cross-attention matrix and the value features of the first road level. In step S804, the semantic features of the first road level and the cross-attention modules corresponding to the remaining at least one road level can be used to continue to complete the progressive feature processing based on the cross-attention mechanism, and finally the comprehensive road features can be obtained.

[0103] In some embodiments, the multi-head cross-attention mechanism can be used to perform feature processing on the query features, key features, and value features. Correspondingly, in step S801 and step S802, multiple groups of query features, multiple groups of key features, and multiple groups of value features corresponding to the multi-head can be constructed, and in step S803, feature processing can be performed based on the multi-head cross-attention mechanism to obtain more effective semantic features.

[0104] It can be understood that in step S801 and step S802, the query features, key features, and value features can also be constructed in other ways. For example, a small network such as a multi-layer perceptron can be used for construction, or the road query features can be directly used as the query features and / or the road encoding features can be directly used as the key features and value features, which are not limited herein.

[0105] In some embodiments, in step S802, the road encoding features of the first road level can be first processed based on the self-attention mechanism to obtain road spatial attention features, and then the query features of the first road level can be constructed based on the road spatial attention mechanism.

[0106] In some embodiments, in step S803, the semantic features of the first road level obtained after feature processing based on the cross-attention mechanism can be refined to further improve the prediction ability of the semantic features, and in step S804, the refined semantic features can be sequentially interacted with the road encoding features of the remaining road levels. The above work can be completed by a refinement module corresponding to the first road level. The refinement module can adopt a multi-layer perceptron or other structures, which are not limited herein.

[0107] In fact, the semantic features of the first road level obtained in step S803 are equivalent to the query features that extract the semantic information of the first road level. These query features are more effective than the road query features and have stronger prediction ability. Feature processing can be performed on these query features and the road encoding features of the remaining road levels in step S804 to extract the semantic information of these road levels.

[0108] Figure 9 FIG. 900 is a flowchart showing a process of feature processing on the semantic features of the first road level and at least one remaining road level according to an exemplary embodiment of the present disclosure. Process 900 can be used to implement step S804 in the above process 800. Process 900 includes: step S901, constructing key features and value features for each of at least one remaining road level based on the road encoding features of each of at least one remaining road level; step S902, for at least one remaining road level, sequentially constructing query features of the current road level based on the semantic features of the previous road level, and performing feature processing on the query features, key features, and value features of the current road level based on the cross-attention mechanism until the semantic features of the last road level are obtained; and step S903, determining the road comprehensive features based on the semantic features of the last road level.

[0109] Thus, in the above manner, progressive extraction of information of different road levels based on the cross-attention mechanism is achieved, enabling more effective road comprehensive features to be obtained. During this process, the form or size of the query features can remain unchanged to facilitate continuous feature processing based on the cross-attention mechanism.

[0110] In fact, the above processes 800 and 900 are equivalent to continuously "querying" the information of different road levels based on the cross-attention mechanism using a road query feature that does not contain specific road information, and sequentially extracting the semantic information related to the road level from the road encoding features of these road levels. Through progressive feature processing, the effectiveness and prediction ability of the query features (i.e., the semantic features corresponding to each road level) are gradually improved, and finally the road comprehensive features fully extract the information of multiple road levels related to the planned route and other surrounding roads.

[0111] In some embodiments, the construction operations of the key feature and the value feature in step S901 may refer to step S802 above. In step S902, query features of the second road level (i.e., the first road level among the remaining at least one road levels) may be constructed based on the semantic features of the first road level, and the query features, key features, and value features of the second road level may be processed based on the cross-attention mechanism to obtain the semantic features of the second road level. Furthermore, query features of the third road level (if the number of all road levels is greater than or equal to three) may be constructed based on the semantic features of the second road level until the semantic features of the last road level are obtained. The operation of constructing query features of the current road level based on the semantic features of the previous road level in step S902 may refer to step S801 above.

[0112] In some embodiments, as described above, each road level may have a corresponding cross-attention module, which includes a trained query matrix, key matrix, and value matrix. For the remaining at least one road levels, in step S901, key features and value features are constructed using the key matrix and value matrix included in the corresponding cross-attention module, and in step S902, query features may be constructed using the query matrix included in the corresponding cross-attention module.

[0113] In some embodiments, in step S902, the semantic features of each road level may be refined, and query features of the next road level may be constructed based on the refined features. In step S903, the semantic features of the last road level may be refined, and the refined features may be used as the road comprehensive features. Refinement can further improve the prediction ability of the semantic features. The above operations may be completed by the refinement modules corresponding to the remaining at least one road levels respectively. The refinement module may adopt a multi-layer perceptron or other structures, which is not limited herein.

[0114] In some embodiments, in step S903, the road comprehensive features may be determined based on the semantic features of the last road level in other ways, or the semantic features of the last road level may be directly used as the road comprehensive features, which is not limited herein.

[0115] In some embodiments, in step S804, the road comprehensive features may be obtained based on the semantic features of the first road level and the encoding features of the remaining at least one road levels in other ways. For example, the semantic features of the first road level and the encoding features of the second road level may be mapped respectively and then added together, or concatenated and then processed using a multi-layer perceptron or other neural network structures, or the process may be implemented in other ways. Similarly, in step S203, the road comprehensive features may be obtained based on the road query features and the road encoding features of multiple road levels in other ways, which is not limited herein.

[0116] Figure 10 A schematic diagram of a road hierarchy model 1000 according to an exemplary embodiment of the present disclosure is shown. The road hierarchy model 1000 includes a self-attention module 1010, a cross-attention module 1020, and a refinement module 1030. It should be noted that each of the multiple road hierarchies has a corresponding road hierarchy model for feature processing of road query features or semantic features of the previous road hierarchy and road encoding features of this road hierarchy. For a certain target road hierarchy, the road encoding features of the target road hierarchy (and the path encoding features 1002 of at least one path included therein) can be input into the self-attention module 1010 to obtain road spatial attention features.

[0117] Furthermore, if the target road hierarchy is the first road hierarchy, the road query feature 1004 and the road spatial attention features can be input into the cross-attention module 1020 to obtain the semantic features of the target road hierarchy. In the cross-attention module 1020, the road query feature is mapped to a query feature, and the road spatial attention features are mapped to key features and value features for performing cross-attention mechanism calculations. If the target road hierarchy is not the first road hierarchy, the semantic features 1004 output by the road hierarchy model of the previous road hierarchy and the road spatial attention features can be input into the cross-attention module 1020. In the cross-attention module 1020, the feature 1004 is mapped to a query feature, and the road spatial attention features are mapped to key features and value features for performing cross-attention mechanism calculations. After obtaining the calculation result of the cross-attention mechanism, the refinement module 1030 can be used to refine the semantic features of the target road hierarchy to obtain the refined semantic features 1006 as the output result of the road hierarchy model corresponding to the target road hierarchy.

[0118] In an exemplary embodiment, the size of the path encoding features 1002 of at least one path can be [N, 256], where N is the number of paths. The size of the path query feature or the semantic feature 1004 of the previous road hierarchy can be [1, 256]. The size of the refined semantic features 1006 output by the refinement module 1030 can also be [1, 256]. As can be seen from the above, the road hierarchy model can make the output semantic features maintain the size of the input road query features or the semantic features of the previous road hierarchy, thus facilitating progressive feature processing.

[0119] In some embodiments, for each road level, a road coding model and a road level model can be combined into a road level coding model for that road level. The road level coding model receives road query features or the result of feature processing of the previous road level, and at the same time receives the position information of the road nodes of the current road level, and outputs the result of feature processing corresponding to the current road level.

[0120] According to some embodiments, step S204, generating an autonomous driving control signal based on the comprehensive road features may include: inputting the navigation information, environmental perception information, and comprehensive road features of the vehicle into an autonomous driving model (for example, a control signal generation model) to obtain an autonomous driving control signal output by the autonomous driving model (for example, a control signal generation model), and the autonomous driving control signal is used to control the vehicle to perform autonomous driving.

[0121] Thus, through the above method, the comprehensive road features can be incorporated into the data-driven autonomous driving model, realizing that both the perception module and the decision-making module are data-driven, and capable of generating an autonomous driving control signal that follows the planned route.

[0122] In some embodiments, the autonomous driving model can be, for example, a Planning and Control (PNC) decision model (for example, a control signal generation model). The comprehensive road features obtained by the method of the present disclosure can be well compatible with the PNC decision model. The autonomous driving model can be trained by using the training method of the autonomous driving model to be described below.

[0123] Figure 11 FIG. shows a schematic diagram of an autonomous driving model 1100 according to an exemplary embodiment of the present disclosure. The autonomous driving model 1100 includes a plurality of road level coding models 1110 corresponding to a plurality of road levels and a control signal generation model 1120. The road level coding model of the first road level receives road query features 1112. Each road level coding model can include a road coding model and a road level model corresponding to the corresponding road level, and can receive the position information 1102 of the road nodes of the corresponding road level to output the result of feature processing corresponding to that road level. The road level coding model of the last road level outputs the comprehensive road features 1114, and these features can be further input into the control signal generation model 1120. In addition to receiving the comprehensive road features 1114, the control signal generation model 1120 can also receive the navigation information 1104 and environmental perception information 1106 of the vehicle, so as to generate an autonomous driving control signal 1108 for controlling the vehicle to follow the planned route for autonomous driving.

[0124] According to another aspect of the present disclosure, a method for training an autonomous driving model is provided. The autonomous driving model includes a road encoding model, a plurality of road level models corresponding to a plurality of road levels, and a control signal generation model. Figure 12 FIG. 1200 shows a flowchart of a method 1200 for training an autonomous driving model according to an exemplary embodiment of the present disclosure. The training method 1200 includes: Step S1201, determining respective sample road nodes of a plurality of road levels, where the sample road nodes represent a section of road located within the same lane; Step S1202, encoding the position information of the respective sample road nodes of the plurality of road levels using the road encoding model to obtain respective road encoding features of the plurality of road levels; Step S1203, using the plurality of road level models to obtain a comprehensive road feature based on an initial road query feature and the respective road encoding features of the plurality of road levels; Step S1204, generating an initial autonomous driving control signal using the control signal generation model based on the comprehensive road feature; and Step S1205, training the autonomous driving model based on the initial autonomous driving control signal and a true control signal to obtain a target autonomous driving model.

[0125] It can be understood that the operations and effects of Steps S1201 - S1204 in Method 1200 can refer to the descriptions of Steps S201 - S204 in Method 200 above, and will not be elaborated here.

[0126] Thus, through the above manner, the target road encoding model, the plurality of target road level models, and the target control signal generation model in the trained target autonomous driving model can generate an autonomous driving control signal for controlling the vehicle to follow the planned route based on the comprehensive road feature, thereby realizing end-to-end data-driven autonomous driving.

[0127] In some embodiments, the autonomous driving model further includes a road query feature. The road query feature can be obtained through random initialization. In Step S1205, the training of the autonomous driving model further includes training and optimizing the road query feature to obtain a target road query feature. The trained target road query feature can be used as the preset road query feature used in the above Method 200.

[0128] In some embodiments, the plurality of road levels may have different degrees of association with the sample planned route of the sample vehicle.

[0129] Figure 13FIG. 0 shows a flowchart of a process 1300 for encoding multiple road levels according to an exemplary embodiment of the present disclosure. The process 1300 can be used to implement step S1202 in the above method 1200. The process 1300 includes: step S1301, for each road level among the multiple road levels, based on the road connectivity relationship of the road level, determine at least one sample path of the road level, where each sample path includes one or more sample road nodes of the road level that can be continuously traveled by a vehicle; and step S1302, use a road encoding model to encode the position information of the one or more sample road nodes included in each of the at least one sample path, to obtain a path encoding feature for each of the at least one sample path, where the road encoding feature of the road level corresponding to the at least one path includes the path encoding feature for each of the at least one sample path.

[0130] It can be understood that the operations and effects of steps S1301 - S1302 in the process 1300 can refer to the description of steps S501 - S502 in the process 500 above, and will not be elaborated here.

[0131] Since the road nodes included in each road level are scattered, by determining the paths that can be traveled by a vehicle based on the connectivity relationship between road nodes in each road level and encoding the position information of the road nodes on the paths, more effective road encoding features can be obtained.

[0132] According to some embodiments, the road encoding model may include a road point encoding sub - model, a road point aggregation sub - model, and a road point fusion sub - model.

[0133] Figure 14 FIG. 13 shows a flowchart of a process 1400 for encoding at least one sample path according to an exemplary embodiment of the present disclosure. The process 1400 can be used to implement step S1302 in the above process 1300. The process 1400 includes: step S1401, for each of the at least one sample path, sample the one or more sample road nodes included in the sample path based on a preset sampling number to obtain a plurality of sample road points; step S1402, use the road point encoding sub - model to encode the coordinates of the plurality of sample road points to obtain a sample road point encoding feature of the sample path; step S1403, use the road point aggregation sub - model to aggregate the sample road point encoding features of each of the at least one sample path to obtain a sample road point aggregation feature; and step S1404, use the road point fusion sub - model to fuse the sample road point encoding features of each of the at least one sample path with the sample road point aggregation feature respectively to obtain a path encoding feature for each of the at least one sample path.

[0134] It can be understood that the operations and effects of steps S1401 - S1404 in process 1400 can be referred to the descriptions of steps S601 - S604 in process 600 above, and will not be elaborated here.

[0135] Thus, by sampling among the sample road nodes included in each path based on a preset number of samples and encoding the coordinates of the sample road points, different paths include the same number of road points, which facilitates encoding. And by aggregating the encoded features of the sample road points of different paths and then fusing the aggregated features of the sample road points and the encoded features of the sample road points of each path, the path encoding feature of each path includes both the position information of the path itself and the comprehensive information of all paths at the road level, thereby enhancing the prediction ability of the path encoding feature.

[0136] According to some embodiments, step S1404, using the road point fusion sub - model to fuse the encoded features of the sample road points of at least one sample path with the aggregated features of the sample road points respectively to obtain the path encoding features of at least one sample path respectively may include: using the road point fusion sub - model to splice the encoded features of the sample road points of at least one sample path with the aggregated features of the sample road points along the hidden dimension respectively to obtain the path encoding features of at least one sample path respectively.

[0137] Thus, through the above - mentioned method, it is possible to originally retain both the position information of the sample road points in each path and the comprehensive information of all paths at the corresponding road level in the path encoding feature, further enhancing the prediction ability of the path encoding feature.

[0138] Figure 15 FIG. shows a flowchart of a process 1500 for feature processing of an initial road query feature and multiple road levels according to an exemplary embodiment of the present disclosure. Process 1500 can be used to implement step S1203 in the above - mentioned method 1200. Process 1500 includes: step S1501, constructing a query feature of a first road level among multiple road levels based on the initial road query feature; step S1502, constructing a key feature and a value feature of the first road level based on the road encoding feature of the first road level; step S1503, using the road level model corresponding to the first road level to perform feature processing on the query feature, key feature, and value feature of the first road level based on the cross - attention mechanism to obtain the semantic feature of the first road level; and step S1504, using at least one road level model corresponding to the remaining at least one road level to obtain a road comprehensive feature based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level respectively.

[0139] It can be understood that the operations and effects of steps S1501 - S1404 in process 1500 can be referred to the description of steps S801 - S804 in process 800 above, and will not be elaborated here.

[0140] Thus, by using the cross - attention mechanism, it is possible to effectively extract the road position information in the road encoding features of the first road level. In addition, by constructing the query features of the first road level based on the initial road query features, and constructing the value features and key features of the first road level based on the road encoding features of the first road level, the semantic features of the first road level obtained after passing through the cross - attention mechanism can maintain the form or size of the road query features, so as to perform further feature processing with the road encoding features of the remaining road levels later.

[0141] Figure 16 FIG. shows a flowchart of a process 1600 for feature processing of the semantic features of the first road level and at least one remaining road level according to an exemplary embodiment of the present disclosure. Process 1600 can be used to implement step S1504 in the above - mentioned process 1500. Process 1600 includes: step S1601, constructing key features and value features for each of the at least one remaining road level based on the road encoding features of each of the at least one remaining road level; step S1602, for each of the at least one remaining road level, successively constructing the query features of the current road level based on the semantic features of the previous road level, and using the road - level model corresponding to the current road level to perform feature processing on the query features, key features, and value features of the current road level based on the cross - attention mechanism until the semantic features of the last road level are obtained; and step S1603, determining the comprehensive road features based on the semantic features of the last road level.

[0142] It can be understood that the operations and effects of steps S1601 - S1603 in process 1600 can be referred to the description of steps S1001 - S1003 in process 1000 above, and will not be elaborated here.

[0143] Thus, through the above - mentioned method, progressive extraction of information of different road levels based on the cross - attention mechanism is achieved, enabling more effective comprehensive road features to be obtained, and during this process, the form or size of the query features can remain unchanged to facilitate continuous feature processing based on the cross - attention mechanism.

[0144] According to some embodiments, the initial road query features are successively subjected to feature processing based on the cross - attention mechanism with the road encoding features of multiple road levels in ascending order of the degree of association with the sample planned route.

[0145] According to some embodiments, at step S1204, a ground truth control signal corresponding to the sample vehicle can be obtained. The ground truth control signal is not necessarily the signal actually used to control the vehicle, and can be label data in the form of a control signal obtained by other means. After obtaining the ground truth control signal, a loss value can be determined based on the initial autonomous driving control signal and the ground truth control signal, and the parameters of multiple models involved in the autonomous driving model can be adjusted using the loss value, and the initial road query feature can be optimized. It can be understood that the target road encoding model, the target road query feature, the multiple target road level models, and the target control signal generation model can respectively represent the road encoding model, the initial road query feature, the multiple road level models, and the control signal generation model that have been trained and optimized.

[0146] In some embodiments, the control signal generation model can be an end-to-end model, and the road encoding model, the initial road query feature, the multiple road level models, and the control signal generation model can be jointly trained end-to-end.

[0147] According to another aspect of the present disclosure, a control signal acquisition device for autonomous driving is provided. Figure 17 The structural block diagram of the control signal acquisition device 1700 according to an exemplary embodiment of the present disclosure is disclosed. As Figure 17 shown, the control signal acquisition device 1700 includes: a first road level determination unit 1710, configured to determine multiple road levels and the road nodes of each of the multiple road levels, the multiple road levels having different degrees of association with the planned route of the vehicle, and the road nodes representing a section of road located within the same lane; a first road encoding unit 1720, configured to encode the position information of the road nodes of each of the multiple road levels to obtain road encoding features of each of the multiple road levels; a first feature extraction unit 1730, configured to obtain a comprehensive road feature based on a preset road query feature and the road encoding features of each of the multiple road levels; and a first control signal generation unit 1740, configured to generate an autonomous driving control signal based on the comprehensive road feature.

[0148] It can be understood that the operations and effects of the units 1710-1740 in the device 1700 can refer to the descriptions of steps S201-S204 in the method 200 above, and will not be elaborated here.

[0149] According to some embodiments, the multiple road levels can include: a planning level, the planning level including the road nodes covered by the planned route; an alternative level, the alternative level including the road nodes covered by an alternative route having the same end point as the planned route; and other levels, the other levels including other road nodes near the vehicle. The degrees of association of the planning level, the alternative level, and the other levels with the planned route decrease in sequence.

[0150] According to some embodiments, the first road encoding unit may include: a first path determination subunit configured to determine, for each of a plurality of road levels, at least one path of the road level based on the road connectivity relationship of the road level, wherein each path includes one or more road nodes of the road level that can be continuously traveled by a vehicle; and a first path encoding subunit configured to encode the position information of one or more road nodes included in each of the at least one path to obtain a path encoding feature of each of the at least one path, wherein the road encoding feature of the road level corresponding to the at least one path includes the path encoding features of each of the at least one path.

[0151] According to some embodiments, the first path determination subunit may be configured to: for each of a plurality of road levels, determine at least one path of the road level based on the road connectivity relationship of the road nodes located within the field of view of the vehicle included in the road level.

[0152] According to some embodiments, the coordinates of a plurality of road points may be based on a coordinate system with the vehicle's location as the origin and the vehicle's head direction as the first direction.

[0153] According to some embodiments, the first path encoding subunit may include: a first road point sampling subunit configured to sample one or more road nodes included in each of the at least one path based on a preset number of samples to obtain a plurality of road points; a first road point encoding subunit configured to encode the coordinates of the plurality of road points to obtain a road point encoding feature of the path; a first feature aggregation subunit configured to aggregate the road point encoding features of each of the at least one path to obtain a road point aggregation feature; and a first feature fusion subunit configured to fuse the road point encoding features of each of the at least one path with the road point aggregation feature respectively to obtain a path encoding feature of each of the at least one path.

[0154] According to some embodiments, the first feature fusion subunit may be configured to: splice the road point encoding features of each of the at least one path with the road point aggregation feature along the hidden dimension respectively to obtain a path encoding feature of each of the at least one path.

[0155] According to some embodiments, the first feature extraction unit may include: a first feature construction subunit configured to construct query features of a first road level in multiple road levels based on road query features; a second feature construction subunit configured to construct key features and value features of the first road level based on road coding features of the first road level; a first cross-attention subunit configured to perform feature processing on the query features, key features, and value features of the first road level based on a cross-attention mechanism to obtain semantic features of the first road level; and a first feature extraction subunit configured to obtain comprehensive road features based on the semantic features of the first road level and road coding features of each of the remaining at least one road level.

[0156] According to some embodiments, the road query features may be subjected to feature processing based on the cross-attention mechanism with road coding features of each of the multiple road levels in ascending order of the degree of association with the planned route.

[0157] According to some embodiments, the first feature extraction subunit may include: a third feature construction subunit configured to construct key features and value features of each of the remaining at least one road level based on road coding features of each of the remaining at least one road level; a second cross-attention subunit configured to, for each of the remaining at least one road level, sequentially construct query features of the current road level based on semantic features of the previous road level and perform feature processing on the query features, key features, and value features of the current road level based on the cross-attention mechanism until semantic features of the last road level are obtained; and a second feature extraction subunit configured to determine comprehensive road features based on the semantic features of the last road level.

[0158] According to some embodiments, the first control signal generation unit may be configured to: input navigation information, environment perception information, and comprehensive road features of a vehicle into an autonomous driving model to obtain an autonomous driving control signal output by the autonomous driving model, and the autonomous driving control signal is used to control the vehicle to perform autonomous driving.

[0159] According to another aspect of the present disclosure, there is provided a training device for an autonomous driving model. The autonomous driving model includes a road coding model, multiple road level models corresponding to multiple road levels, and a control signal generation model. Figure 18 The structural block diagram of a training device 1800 for an autonomous driving model according to an exemplary embodiment of the present disclosure is disclosed. As Figure 18As shown, the training device 1800 includes: a second road level determination unit 1810 configured to determine sample road nodes for each of a plurality of road levels, where the sample road nodes represent a section of road located within the same lane; a second road encoding unit 1820 configured to encode the position information of the sample road nodes for each of the plurality of road levels using a road encoding model to obtain road encoding features for each of the plurality of road levels; a second feature extraction unit 1830 configured to use a plurality of road level models to obtain a comprehensive road feature based on an initial road query feature and the road encoding features for each of the plurality of road levels; a second control signal generation unit 1840 configured to generate an initial autonomous driving control signal based on the comprehensive road feature using a control signal generation model; and a training unit 1850 configured to train an autonomous driving model based on the initial autonomous driving control signal and a true control signal to obtain a target autonomous driving model.

[0160] It can be understood that the operations and effects of units 1810-1850 in device 1800 can be referred to the descriptions of steps S1201-S1205 in method 1200 above, and will not be elaborated here.

[0161] According to some embodiments, the second road encoding unit may include: a second path determination subunit configured to, for each of the plurality of road levels, determine at least one sample path for the road level based on the road connectivity relationship of the road level, where each sample path includes one or more sample road nodes that can be continuously traveled by a vehicle on the road level; and a second path encoding subunit configured to encode the position information of the one or more sample road nodes included in each of the at least one sample path using a road encoding model to obtain path encoding features for each of the at least one sample path, where the road encoding features of the road level corresponding to the at least one path include the path encoding features of each of the at least one sample path.

[0162] According to some embodiments, the road encoding model may include a road point encoding sub-model, a road point aggregation sub-model, and a road point fusion sub-model. The second path encoding sub-unit may include: a second road point sampling sub-unit configured to sample one or more sample road nodes included in each of at least one sample path based on a preset number of samplings to obtain a plurality of sample road points; a second road point encoding sub-unit configured to encode the coordinates of the plurality of sample road points by using the road point encoding sub-model to obtain the sample road point encoding features of the sample path; a second feature aggregation sub-unit configured to aggregate the sample road point encoding features of at least one sample path respectively by using the road point aggregation sub-model to obtain sample road point aggregation features; and a second feature fusion sub-unit configured to fuse the sample road point encoding features of at least one sample path respectively with the sample road point aggregation features by using the road point fusion sub-model to obtain the path encoding features of at least one sample path respectively.

[0163] According to some embodiments, the second feature fusion sub-unit may be configured to: fuse the sample road point encoding features of at least one sample path respectively with the sample road point aggregation features along the hidden dimension by using the road point fusion sub-model to obtain the path encoding features of at least one sample path respectively.

[0164] According to some embodiments, multiple road levels may have different degrees of association with the sample planning route of the sample vehicle. The second feature extraction unit may include: a fourth feature construction sub-unit configured to construct the query feature of the first road level in the multiple road levels based on the initial road query feature; a fifth feature construction sub-unit configured to construct the key feature and the value feature of the first road level based on the road encoding feature of the first road level; a third cross-attention sub-unit configured to perform feature processing on the query feature, the key feature, and the value feature of the first road level by using the road level model corresponding to the first road level based on the cross-attention mechanism to obtain the semantic feature of the first road level; and a third feature extraction sub-unit configured to obtain the comprehensive road feature by using the at least one road level model corresponding to the remaining at least one road level based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level respectively.

[0165] According to some embodiments, the third feature extraction subunit may include: a sixth feature construction subunit configured to construct key features and value features for each of the at least one remaining road level based on the road coding features of each of the at least one remaining road level; a fourth cross-attention subunit configured to, for each of the at least one remaining road level, sequentially construct a query feature for the current road level based on the semantic feature of the previous road level, and perform feature processing on the query feature, key feature, and value feature of the current road level based on the cross-attention mechanism using the road level model corresponding to the current road level until the semantic feature of the last road level is obtained; and a fourth feature extraction subunit configured to determine the comprehensive road feature based on the semantic feature of the last road level.

[0166] According to some embodiments, the initial road query feature is sequentially subjected to feature processing based on the cross-attention mechanism with the road coding features of each of the multiple road levels in ascending order of the degree of association with the sample planned route.

[0167] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0168] According to an embodiment of the present disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.

[0169] Reference Figure 19 , the structural block diagram of the electronic device 1900 that can be used as the server or client of the present disclosure will now be described. It is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0170] As Figure 19As shown, device 1900 includes a computing unit 1901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1902 or a computer program loaded from a storage unit 1908 into a random access memory (RAM) 1903. In the RAM 1903, various programs and data required for the operation of the device 1900 can also be stored. The computing unit 1901, the ROM 1902, and the RAM 1903 are connected to each other through a bus 1904. An input / output (I / O) interface 1905 is also connected to the bus 1904.

[0171] A plurality of components in the device 1900 are connected to the I / O interface 1905, including: an input unit 1906, an output unit 1907, a storage unit 1908, and a communication unit 1909. The input unit 1906 can be any type of device that can input information into the device 1900. The input unit 1906 can receive input digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include but are not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote control. The output unit 1907 can be any type of device that can present information, and can include but are not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1908 can include but are not limited to magnetic disks and optical disks. The communication unit 1909 allows the device 1900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include but are not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth TM device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0172] The computing unit 1901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning network algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1901 executes the various methods and processes described above, such as the control signal acquisition method for autonomous driving and / or the training method for the autonomous driving model. For example, in some embodiments, the control signal acquisition method for autonomous driving and / or the training method for the autonomous driving model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 1908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1900 via the ROM 1902 and / or the communication unit 1909. When the computer program is loaded into the RAM 1903 and executed by the computing unit 1901, one or more steps of the control signal acquisition method for autonomous driving and / or the training method for the autonomous driving model described above can be executed. Alternatively, in other embodiments, the computing unit 1901 can be configured to execute the control signal acquisition method for autonomous driving and / or the training method for the autonomous driving model in any other suitable manner (e.g., by means of firmware).

[0173] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0174] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.

[0175] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0176] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0177] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0178] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0179] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0180] Although embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above methods, systems, and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but is only defined by the authorized claims and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, the steps can be executed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. Importantly, with the evolution of technology, many of the elements described herein can be replaced by equivalent elements that emerge after the present disclosure.

Claims

1. A method for obtaining control signals for autonomous driving, comprising: Determining a plurality of road levels and the road nodes of each of the plurality of road levels, where the plurality of road levels have different degrees of association with the planned route of the vehicle, and the road nodes represent a section of road within the same lane; Encoding the position information of the road nodes of each of the plurality of road levels to obtain the road encoding features of each of the plurality of road levels; Based on a preset road query feature and the road encoding features of each of the plurality of road levels, obtaining a comprehensive road feature, including: Based on the road query feature, constructing a query feature of a first road level among the plurality of road levels; Based on the road encoding feature of the first road level, constructing a key feature and a value feature of the first road level; Based on a cross-attention mechanism, performing feature processing on the query feature, key feature, and value feature of the first road level to obtain the semantic feature of the first road level; and Based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level, obtaining the comprehensive road feature; and Generating an autonomous driving control signal based on the comprehensive road feature.

2. The method according to claim 1, wherein, Obtaining the comprehensive road feature based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level includes: Based on the road encoding features of the remaining at least one road level, constructing the key feature and the value feature of each of the at least one road level; For the remaining at least one road level, sequentially constructing a query feature of the current road level based on the semantic feature of the previous road level, and performing feature processing on the query feature, key feature, and value feature of the current road level based on the cross-attention mechanism until the semantic feature of the last road level is obtained; and Based on the semantic feature of the last road level, determining the comprehensive road feature.

3. The method according to claim 2, wherein The road query feature performs feature processing based on the cross-attention mechanism with the road encoding features of each of the plurality of road levels in ascending order of the degree of association with the planned route.

4. The method according to any one of claims 1 to 3, wherein, The plurality of road levels include: A planning level, where the planning level includes the road nodes covered by the planned route; An alternative level, where the alternative level includes the road nodes covered by an alternative route having the same end point as the planned route; and An other level, where the other level includes other road nodes near the vehicle, where the degree of association of the planning level, the alternative level, and the other level with the planned route decreases in sequence.

5. The method according to any one of claims 1 to 3, wherein Encoding the position information of the road nodes of each of the plurality of road levels to obtain the road encoding features of each of the plurality of road levels includes: For each road level among the plurality of road levels, based on the road connectivity relationship of the road level, determining at least one path of the road level, where each path includes one or more road nodes of the road level that can be continuously traveled by the vehicle; and Encode the location information of one or more road nodes included in each of the at least one path to obtain the path encoding feature of each of the at least one path, where the road encoding feature of the road level corresponding to the at least one path includes the path encoding feature of each of the at least one path.

6. The method according to claim 5, wherein Encoding the location information of one or more road nodes included in each of the at least one path to obtain the path encoding feature of each of the at least one path includes: For each path in the at least one path, sample one or more road nodes included in the path based on a preset number of samples to obtain a plurality of road points; Encode the coordinates of the plurality of road points to obtain the road point encoding feature of the path; Aggregate the road point encoding features of each of the at least one path to obtain a road point aggregation feature; and Fuse the road point encoding features of each of the at least one path with the road point aggregation feature respectively to obtain the path encoding feature of each of the at least one path.

7. The method according to claim 6, wherein, Fusing the road point encoding features of each of the at least one path with the road point aggregation feature respectively to obtain the path encoding feature of each of the at least one path includes: Concatenate the road point encoding features of each of the at least one path with the road point aggregation feature along the hidden dimension to obtain the path encoding feature of each of the at least one path.

8. The method according to claim 5, wherein, For each of the multiple road levels, determine at least one path of the road level based on the road connectivity relationship of the road level, including: For each of the multiple road levels, determine at least one path of the road level based on the road connectivity relationship of the road nodes located within the field of view of the vehicle included in the road level.

9. The method according to claim 6, wherein, The coordinates of the plurality of road points are based on a coordinate system with the vehicle's location as the origin and the vehicle's head direction as the first direction.

10. The method according to any one of claims 1-3, wherein Generating an autonomous driving control signal based on the road comprehensive feature includes: Input the vehicle's navigation information, environment perception information, and the road comprehensive feature into an autonomous driving model to obtain the autonomous driving control signal output by the autonomous driving model, and the autonomous driving control signal is used to control the vehicle to perform autonomous driving.

11. A training method for an autonomous driving model, the autonomous driving model including a road encoding model, a plurality of road level models corresponding to a plurality of road levels, and a control signal generation model, wherein, The training method includes: Determine the sample road nodes of each of the multiple road levels, where the sample road nodes represent a section of road within the same lane; Use the road encoding model to encode the location information of the sample road nodes of each of the multiple road levels to obtain the road encoding features of each of the multiple road levels; Use the multiple road level models to obtain a road comprehensive feature based on an initial road query feature and the road encoding features of each of the multiple road levels; Use the control signal generation model to generate an initial autonomous driving control signal based on the road comprehensive feature; and Train the autonomous driving model based on the initial autonomous driving control signal and the true control signal to obtain a target autonomous driving model. Among them, the multiple road levels have different degrees of association with the sample planned route of the sample vehicle. Using the multiple road level models, based on the initial road query feature and the road coding features of the multiple road levels respectively, obtaining the comprehensive road feature includes: Based on the initial road query feature, constructing the query feature of the first road level in the multiple road levels; Based on the road coding feature of the first road level, constructing the key feature and value feature of the first road level; Using the road level model corresponding to the first road level, based on the cross-attention mechanism, performing feature processing on the query feature, key feature and value feature of the first road level to obtain the semantic feature of the first road level; and Using at least one road level model corresponding to the remaining at least one road level, based on the semantic feature of the first road level and the road coding features of the remaining at least one road level respectively, obtaining the comprehensive road feature.

12. The method according to claim 11, wherein, Using at least one road level model corresponding to the remaining at least one road level, based on the semantic feature of the first road level and the road coding features of the remaining at least one road level respectively, obtaining the comprehensive road feature includes: Based on the road coding features of the remaining at least one road level respectively, constructing the key features and value features of the at least one road level respectively; For the remaining at least one road level, sequentially constructing the query feature of the current road level based on the semantic feature of the previous road level, and using the road level model corresponding to the current road level to perform feature processing on the query feature, key feature and value feature of the current road level based on the cross-attention mechanism until the semantic feature of the last road level is obtained; and Based on the semantic feature of the last road level, determining the comprehensive road feature.

13. The method according to claim 12, wherein The initial road query feature is sequentially subjected to feature processing based on the cross-attention mechanism with the road coding features of the multiple road levels respectively from low to high in terms of the degree of association with the sample planned route.

14. The method according to any one of claims 11-13, wherein, Using the road coding model to encode the position information of the sample road nodes of the multiple road levels respectively, obtaining the road coding features of the multiple road levels respectively includes: For each road level in the multiple road levels, based on the road connection relationship of the road level, determining at least one sample path of the road level, where each sample path includes one or more sample road nodes that can be continuously traveled by the vehicle; and Using the road coding model to encode the position information of the one or more sample road nodes included in each of the at least one sample path, obtaining the path coding features of each of the at least one sample path, where the road coding feature of the road level corresponding to the at least one path includes the path coding features of each of the at least one sample path.

15. The method according to claim 14, wherein The road coding model includes a road point coding sub-model, a road point aggregation sub-model and a road point fusion sub-model. Among them, encoding the location information of one or more sample road nodes included in each of the at least one sample path by using the road encoding model to obtain the path encoding feature of each of the at least one sample path includes: For each of the at least one sample path, sampling one or more sample road nodes included in the sample path based on a preset number of samplings to obtain a plurality of sample road points; Encoding the coordinates of the plurality of sample road points by using the road point encoding sub-model to obtain the sample road point encoding feature of the sample path; Aggregating the sample road point encoding features of each of the at least one sample path by using the road point aggregation sub-model to obtain a sample road point aggregation feature; and Fusing the sample road point encoding features of each of the at least one sample path with the sample road point aggregation feature respectively by using the road point fusion sub-model to obtain the path encoding feature of each of the at least one sample path.

16. The method according to claim 15, wherein, Fusing the sample road point encoding features of each of the at least one sample path with the sample road point aggregation feature respectively by using the road point fusion sub-model to obtain the path encoding feature of each of the at least one sample path includes: Fusing the sample road point encoding features of each of the at least one sample path with the sample road point aggregation feature respectively along the hidden dimension by using the road point fusion sub-model to obtain the path encoding feature of each of the at least one sample path.

17. A control signal acquisition device for autonomous driving, comprising: A first road level determination unit configured to determine a plurality of road levels and road nodes of each of the plurality of road levels, the plurality of road levels having different degrees of association with the planned route of the vehicle, and the road nodes representing a section of road located within the same lane; A first road encoding unit configured to encode the location information of the road nodes of each of the plurality of road levels to obtain the road encoding feature of each of the plurality of road levels; A first feature extraction unit configured to obtain a road comprehensive feature based on a preset road query feature and the road encoding features of each of the plurality of road levels; And A first control signal generation unit configured to generate an autonomous driving control signal based on the road comprehensive feature, wherein, the first feature extraction unit includes: A first feature construction sub-unit configured to construct a query feature of a first road level among the plurality of road levels based on the road query feature; A second feature construction sub-unit configured to construct a key feature and a value feature of the first road level based on the road encoding feature of the first road level; A first cross-attention sub-unit configured to perform feature processing on the query feature, key feature and value feature of the first road level based on a cross-attention mechanism to obtain the semantic feature of the first road level; and A first feature extraction sub-unit configured to obtain the road comprehensive feature based on the semantic feature of the first road level and the road encoding features of the remaining at least one road level.

18. The device according to claim 17, wherein The first feature extraction subunit includes: A third feature construction subunit, configured to construct key features and value features for each of the at least one remaining road level based on the road coding features of each of the at least one remaining road level; A second cross-attention subunit, configured to, for the at least one remaining road level, sequentially construct query features for the current road level based on the semantic features of the previous road level, and perform feature processing on the query features, key features, and value features of the current road level based on the cross-attention mechanism until the semantic features of the last road level are obtained; and A second feature extraction subunit, configured to determine the comprehensive road feature based on the semantic features of the last road level.

19. The apparatus according to claim 18, wherein, The road query features are sequentially subjected to feature processing based on the cross-attention mechanism with the road coding features of each of the multiple road levels in ascending order of the degree of association with the planned route.

20. The apparatus according to any one of claims 17 - 19, wherein, The multiple road levels include: A planning level, where the planning level includes road nodes covered by the planned route; An alternative level, where the alternative level includes road nodes covered by alternative routes having the same end point as the planned route; and An other level, where the other level includes other road nodes near the vehicle, where the degree of association of the planning level, the alternative level, and the other level with the planned route decreases in sequence.

21. The device according to any one of claims 17-19, wherein The first road coding unit includes: A first path determination subunit, configured to, for each road level among the multiple road levels, determine at least one path for the road level based on the road connectivity relationship of the road level, where each of the paths includes one or more road nodes that can be continuously traveled by the vehicle; and A first path coding subunit, configured to encode the position information of the one or more road nodes included in each of the at least one path to obtain path coding features for each of the at least one path, where the road coding features of the road level corresponding to the at least one path include the path coding features for each of the at least one path.

22. The device according to claim 21, wherein The first path coding subunit includes: A first road point sampling subunit, configured to, for each of the at least one path, sample the one or more road nodes included in the path based on a preset number of samples to obtain multiple road points; A first road point coding subunit, configured to encode the coordinates of the multiple road points to obtain road point coding features for the path; A first feature aggregation subunit, configured to aggregate the road point coding features for each of the at least one path to obtain road point aggregation features; and A first feature fusion subunit, configured to fuse the road point coding features for each of the at least one path with the road point aggregation features respectively to obtain the path coding features for each of the at least one path.

23. The apparatus according to claim 22, wherein, The first feature fusion subunit is configured to: Concatenate the road point coding features of each of the at least one path with the road point aggregation feature along the hidden dimension to obtain the path coding feature of each of the at least one path.

24. The apparatus according to claim 21, wherein The first path determination subunit is configured to: For each road level among the multiple road levels, determine at least one path of the road level based on the road connection relationship of the road nodes included in the road level and located within the field of view of the vehicle.

25. The apparatus according to claim 22, wherein, The coordinates of the multiple road points are based on a coordinate system with the vehicle's location as the origin and the vehicle head direction as the first direction.

26. The apparatus according to any one of claims 17-19, wherein, The first control signal generation unit is configured to: Input the navigation information, environmental perception information, and the road comprehensive feature of the vehicle into the autonomous driving model to obtain the autonomous driving control signal output by the autonomous driving model, where the autonomous driving control signal is used to control the vehicle to perform autonomous driving.

27. A training device for an autonomous driving model, the autonomous driving model includes a road encoding model, a plurality of road level models corresponding to multiple road levels, and a control signal generation model, wherein, The training device includes: A second road level determination unit, configured to determine the sample road nodes of each of the multiple road levels, where the sample road nodes represent a section of road within the same lane; A second road coding unit, configured to use the road coding model to encode the position information of the sample road nodes of each of the multiple road levels to obtain the road coding features of each of the multiple road levels; A second feature extraction unit, configured to use the multiple road level models to obtain a road comprehensive feature based on the initial road query feature and the road coding features of each of the multiple road levels; A second control signal generation unit, configured to use the control signal generation model to generate an initial autonomous driving control signal based on the road comprehensive feature; and A training unit, configured to train the autonomous driving model based on the initial autonomous driving control signal and the true control signal to obtain a target autonomous driving model, where the multiple road levels have different degrees of association with the sample planning route of the sample vehicle, and the second feature extraction unit includes: A fourth feature construction subunit, configured to construct a query feature of the first road level among the multiple road levels based on the initial road query feature; A fifth feature construction subunit, configured to construct a key feature and a value feature of the first road level based on the road coding feature of the first road level; A third cross-attention subunit, configured to use the road level model corresponding to the first road level to perform feature processing on the query feature, key feature, and value feature of the first road level based on the cross-attention mechanism to obtain the semantic feature of the first road level; and A third feature extraction subunit, configured to use the at least one road level model corresponding to the remaining at least one road level to obtain the road comprehensive feature based on the semantic feature of the first road level and the road coding features of the remaining at least one road level.

28. The apparatus according to claim 27, wherein, The third feature extraction subunit includes: The sixth feature construction subunit is configured to construct the key features and value features of each of the at least one road level based on the road coding features of each of the remaining at least one road level; The fourth cross-attention subunit is configured to, for the remaining at least one road level, sequentially construct the query features of the current road level based on the semantic features of the previous road level, and use the road level model corresponding to the current road level to perform feature processing on the query features, key features, and value features of the current road level based on the cross-attention mechanism until the semantic features of the last road level are obtained; and The fourth feature extraction subunit is configured to determine the comprehensive road feature based on the semantic features of the last road level.

29. The apparatus according to claim 28, wherein The initial road query features are sequentially subjected to feature processing based on the cross-attention mechanism with the road coding features of each of the multiple road levels from low to high in terms of the degree of association with the sample planned route.

30. The device according to any one of claims 27 - 29, wherein, The second road coding unit includes: The second path determination subunit is configured to, for each road level among the multiple road levels, determine at least one sample path of the road level based on the road connectivity relationship of the road level, where each of the sample paths includes one or more sample road nodes that can be continuously traveled by a vehicle on the road level; and The second path coding subunit is configured to use the road coding model to encode the position information of the one or more sample road nodes included in each of the at least one sample path to obtain the path coding features of each of the at least one sample path, where the road coding features of the road level corresponding to the at least one path include the path coding features of each of the at least one sample path.

31. The apparatus according to claim 30, wherein, The road coding model includes a road point coding sub-model, a road point aggregation sub-model, and a road point fusion sub-model. Wherein, the second path coding subunit includes: The second road point sampling subunit is configured to, for each of the at least one sample path, sample the one or more sample road nodes included in the sample path based on a preset number of samplings to obtain a plurality of sample road points; The second road point coding subunit is configured to use the road point coding sub-model to encode the coordinates of the plurality of sample road points to obtain the sample road point coding features of the sample path; The second feature aggregation subunit is configured to use the road point aggregation sub-model to aggregate the sample road point coding features of each of the at least one sample path to obtain the sample road point aggregation features; and The second feature fusion subunit is configured to use the road point fusion sub-model to fuse the sample road point coding features of each of the at least one sample path with the sample road point aggregation features respectively to obtain the path coding features of each of the at least one sample path.

32. The apparatus according to claim 31, wherein, The second feature fusion subunit is configured to: Using the road point fusion sub-model, the sample road point encoding features of each of the at least one sample path are respectively concatenated with the sample road point aggregation feature along the hidden dimension to obtain the path encoding feature of each of the at least one sample path.

33. An electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-16.

34. An autonomous vehicle, comprising the electronic device according to claim 33.

35. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-16.

36. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1-16.

Citation Information

Patent Citations

  • Intelligent management method for bus lines

    CN105427582A

  • Methods and systems for autonomous vehicle navigation

    CN111351492A