Conditional Turning Trajectory Prediction Network for Urban Intersections

By processing the proxy tensor and map tensor of the target vehicle, and using machine learning models to predict its driving intention at intersections, the problem of inaccurate driving trajectory prediction by autonomous driving systems at urban intersections is solved, thereby improving the safety and efficiency of autonomous driving.

CN122497614APending Publication Date: 2026-07-31QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-12-13
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing autonomous driving systems struggle to accurately predict the driving intentions of other vehicles at urban intersections, leading to inaccurate driving trajectory predictions and impacting the safety and efficiency of autonomous driving.

Method used

Machine learning models are used to process the proxy tensor and map tensor of the target vehicle to predict its turning and flow classification intentions at intersections, and driving maneuvers are performed based on the prediction results.

Benefits of technology

It improves the accuracy of driving trajectory prediction at urban intersections, enhancing the safety and efficiency of autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122497614A_ABST
    Figure CN122497614A_ABST
Patent Text Reader

Abstract

Techniques for predicting driving trajectories are disclosed. In one or more aspects, a self-driving vehicle applies a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and performs driving maneuvers based on the predicted driving intention.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This patent application claims the benefits of U.S. Provisional Application No. 63 / 617,960, filed January 5, 2024, entitled “MULTI-TASK INTENTIONPREDICTION NETWORK FOR URBAN INTERSECTIONS,” and U.S. Non-Provisional Application No. 18 / 887,487, filed September 17, 2024, entitled “CONDITIONALTURN TRAJECTORY PREDICTION NETWORK FOR URBAN INTERSECTIONS,” both of which are assigned to the assignee of this application and are expressly incorporated herein by reference in their entirety. Background Technology 1. Technical Field

[0003] The various aspects disclosed herein relate to semi-autonomous or autonomous driving technologies.

[0004] 2. Relevant Technical Descriptions

[0005] Modern motor vehicles increasingly incorporate semi-autonomous or autonomous driving features, such as technologies that help drivers avoid drifting into adjacent lanes or making unsafe lane changes (e.g., Lane Departure Warning (LDW)), technologies that warn drivers of other vehicles behind them when a vehicle is reversing, or technologies that automatically brake if the vehicle in front of it suddenly stops or slows down (e.g., Forward Collision Warning (FCW)), and so on. The continuous evolution of automotive technology aims to provide even greater safety benefits and ultimately deliver automated driving systems (ADS) capable of taking control of the entire driving task without user intervention.

[0006] There are six defined levels for achieving full automation. At Level 0, the human driver performs all driving. At Level 1, the Advanced Driver Assistance System (ADAS) on the vehicle may sometimes assist the human driver in steering or braking / acceleration, but not simultaneously. At Level 2, the ADAS on the vehicle can, in some situations, effectively control both steering and braking / acceleration simultaneously. The human driver must maintain full attention and perform the remaining driving tasks at all times. At Level 3, the ADAS on the vehicle can perform all aspects of driving tasks in certain situations. In these situations, the human driver must be prepared to relinquish control when requested by the ADAS. In all other situations, the human driver performs the driving tasks. At Level 4, the ADAS on the vehicle can perform all driving tasks and monitor the driving environment, essentially performing all driving in some situations. In these situations, human attention is not required. At Level 5, the ADAS on the vehicle can perform all driving in all situations. The human occupant is merely a passenger and is never involved in driving. Summary of the Invention

[0007] The following is a simplified summary of the invention relating to one or more aspects disclosed herein. Therefore, this summary should not be considered an exhaustive overview relating to all conceived aspects, nor should it be considered to identify key or decisive elements relating to all conceived aspects or to depict the scope associated with any particular aspect. Thus, the sole purpose of this summary is to present, in a simplified form, certain concepts relating to one or more aspects involving the mechanisms disclosed herein, prior to the detailed description presented below.

[0008] In one aspect, a method for predicting driving trajectories performed by a self-driving vehicle includes: applying a machine learning model to one or more agent tensors and one or more map tensors associated with the target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and performing driving maneuvers based on the predicted driving intention.

[0009] In one aspect, a self-driving vehicle includes: one or more memories; one or more transceivers; and one or more processors communicatively coupled to the one or more memories and the one or more transceivers, the one or more processors being configured individually or in combination to: apply a machine learning model to one or more proxy tensors and one or more map tensors associated with the target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and perform driving maneuvers based on the predicted driving intention.

[0010] In one aspect, a self-driving vehicle includes: components for applying a machine learning model to one or more agent tensors and one or more map tensors associated with the target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and components for performing driving maneuvers based on the predicted driving intention.

[0011] In one aspect, a non-transitory computer-readable medium stores computer-executable instructions that, when executed by a self-driving vehicle, enable the self-driving vehicle to: apply a machine learning model to one or more agent tensors and one or more map tensors associated with the target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and perform driving maneuvers based on the predicted driving intention.

[0012] Based on the accompanying drawings and detailed description, other objects and advantages associated with the aspects disclosed herein will be apparent to those skilled in the art. Attached Figure Description

[0013] The accompanying drawings are provided to help describe various aspects of this disclosure, and are provided for illustrative purposes only and not to limit the aspects.

[0014] Figure 1 Example wireless communication systems according to various aspects of this disclosure are illustrated.

[0015] Figure 2A and Figure 2B Example wireless network architectures according to one or more aspects of this disclosure are illustrated.

[0016] Figure 3A It is a top view of a vehicle employing an integrated radar camera sensor behind the windshield, according to one or more aspects of this disclosure.

[0017] Figure 3B An example onboard computer (OBC) architecture according to one or more aspects of this disclosure is illustrated.

[0018] Figure 4 This is a diagram illustrating an example driving strategy pipeline according to one or more aspects of this disclosure.

[0019] Figure 5 Example neural networks according to various aspects of this disclosure are illustrated.

[0020] Figures 6A to 6C An example encoder-decoder machine learning model architecture for predicting turn and flow classification labels for agents of interest, based on various aspects of this disclosure, is illustrated.

[0021] Figure 7 This is an example of the various aspects of this disclosure. Figures 6A to 6C The example encoder-decoder machine learning model architecture is illustrated in the diagram with example enhancements.

[0022] Figure 8 Example intersection scenarios based on various aspects of this disclosure are illustrated.

[0023] Figure 9 An example method for predicting driving trajectories based on various aspects of this disclosure is illustrated. Detailed Implementation

[0024] Various aspects of this disclosure are provided in the following description and accompanying drawings of various examples provided for illustrative purposes. Alternative aspects may be devised without departing from the scope of this disclosure. Additionally, well-known elements of this disclosure will not be described in detail or will be omitted so as not to obscure the relevant details of this disclosure.

[0025] Various aspects are involved in autonomous driving as a whole. Some aspects involve more specifically driving trajectory prediction. In some examples, a self-driving vehicle applies a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle (e.g., at a road intersection). The predicted driving intention may include turn classification, flow classification, and optionally, a corresponding driving trajectory. The self-driving vehicle performs driving maneuvers based on the predicted driving intention (turn classification, flow classification, and corresponding driving trajectory) of the target vehicle.

[0026] Specific aspects of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages. In some examples, by applying a machine learning model to one or more agent tensors and one or more map tensors to obtain a predicted driving intention of a target vehicle, the described techniques can be used to improve trajectory prediction of the target agent / vehicle and thus improve autonomous driving performance.

[0027] The terms “exemplary” and / or “example” are used herein to mean “serving as an example, instance, or illustration.” Any aspect described herein as “exemplary” and / or “example” is not necessarily to be construed as superior to or better than other aspects. Similarly, the term “aspects of this disclosure” does not require that all aspects of this disclosure include the features, advantages, or modes of operation discussed.

[0028] Those skilled in the art will understand that any of a variety of different techniques and methods can be used to represent the information and signals described below. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the following description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof, depending in part on the specific application, in part on the desired design, in part on the corresponding technology, and so on.

[0029] Furthermore, many aspects are described according to a sequence of actions to be performed by elements of, for example, a computing device. It will be appreciated that the various actions described herein can be performed by specific circuitry (e.g., an application-specific integrated circuit (ASIC)), by program instructions executed by one or more processors, or by a combination of both. Additionally, the sequence of actions described herein can be considered to be entirely embodied in any form of non-transitory computer-readable storage medium storing a corresponding set of computer instructions that, when executed, will cause or command the associated processor of the device to perform the functionality described herein. Therefore, various aspects of this disclosure can be embodied in a variety of different forms, all of which are contemplated within the scope of the claimed subject matter. Furthermore, for each aspect described herein, any corresponding form of any such aspect may be described herein as, for example, "logic configured to perform the described actions."

[0030] As used herein, the terms “user equipment” (UE), “vehicle UE” (V-UE), “pedestrian UE” (P-UE), and “base station” are not intended to be specific to or otherwise limited to any particular radio access technology (RAT) unless otherwise stated. In general, a UE can be any wireless communication device used by a user to communicate over a wireless communication network (e.g., vehicle onboard computer, vehicle navigation device, mobile phone, router, tablet computer, laptop computer, asset location device, wearable device (e.g., smartwatch, glasses, augmented reality (AR) / virtual reality (VR) headset, etc.), vehicle (e.g., car, motorcycle, bicycle, etc.), Internet of Things (IoT) device, etc.). A UE can be mobile or can (e.g., at certain times) be stationary and can communicate with a radio access network (RAN). As used herein, the term “UE” can be interchangeably referred to as “mobile device,” “access terminal” or “AT,” “client device,” “wireless device,” “subscriber equipment,” “subscriber terminal,” “subscriber station,” “user terminal” or UT,” “mobile terminal,” “mobile station,” or variations thereof.

[0031] V-UE is a type of UE and can be any in-vehicle wireless communication device, such as a navigation system, alarm system, head-up display (HUD), onboard computer, in-vehicle infotainment system, automated driving system (ADS), advanced driver assistance system (ADAS), etc. Alternatively, V-UE can be a portable wireless communication device (e.g., cellular phone, tablet computer, etc.) carried by the driver or occupant of a vehicle. The term "V-UE" can refer to the in-vehicle wireless communication device or the vehicle itself, depending on the context. P-UE is a type of UE and can be a portable wireless communication device carried by a pedestrian (i.e., a user who is not driving or riding in a vehicle). Generally, the UE can communicate with the core network via the RAN, and through the core network, the UE can connect to external networks such as the Internet and other UEs. Of course, other mechanisms for connecting the UE to the core network and / or the Internet are also possible, such as through wired access networks, wireless local area network (WLAN) networks (e.g., based on IEEE 802.11, etc.).

[0032] A base station can communicate with a UE by operating under one of several RATs based on the network in which it is deployed, and may alternatively be referred to as an Access Point (AP), Network Node, Node B, Evolved Node B (eNB), Next Generation eNB (ng-eNB), New Radio (NR) Node B (also known as gNB or gNodeB), etc. The base station is primarily used to support the UE's radio access, including supporting the UE's data, voice, and / or signaling connections. In some systems, the base station may only provide edge node signaling functions, while in others, it may provide additional control and / or network management functions. The communication link through which the UE can transmit signals to the base station is called an uplink (UL) channel (e.g., reverse traffic channel, reverse control channel, access channel, etc.). The communication link through which the base station can transmit signals to the UE is called a downlink (DL) or forward link channel (e.g., paging channel, control channel, broadcast channel, forward traffic channel, etc.). As used herein, the term Traffic Channel (TCH) may refer to either the UL / reverse or DL / forward traffic channel.

[0033] The term "base station" can refer to a single physical transmit / receive point (TRP) or multiple physical TRPs that may or may not be co-located. For example, when the term "base station" refers to a single physical TRP, the physical TRP can be the antenna of a base station corresponding to a cell (or several cell sectors) of the base station. When the term "base station" refers to multiple co-located physical TRPs, the physical TRP can be the antenna array of the base station (e.g., as in a multiple-input multiple-output (MIMO) system or where the base station employs beamforming). When the term "base station" refers to multiple non-co-located physical TRPs, the physical TRP can be a distributed antenna system (DAS) (a network of spatially separated antennas connected via a transmission medium to a common source) or a remote radio headend (RRH) (a remote base station connected to a serving base station). Alternatively, a non-co-located physical TRP can be the serving base station from which the UE receives measurement reports and a neighboring base station where the UE is measuring its reference radio frequency (RF) signal. Because, as used herein, a TRP is the point by which a base station transmits and receives radio signals, references to transmitting from or receiving at a base station should be understood to refer to a specific TRP of the base station.

[0034] In some specific implementations supporting UE positioning, the base station may not support the UE's radio access (e.g., it may not support the UE's data, voice, and / or signaling connections). Instead, it may send a reference RF signal to the UE for measurement by the UE, and / or receive and measure signals sent by the UE. Such a base station may be referred to as a positioning beacon (e.g., in the case of sending RF signals to the UE) and / or as a location measurement unit (e.g., in the case of receiving and measuring RF signals from the UE).

[0035] An “RF signal” refers to an electromagnetic wave of a given frequency that transmits information across the space between a transmitter and a receiver. As used herein, a transmitter may send a single “RF signal” or multiple “RF signals” to a receiver. However, due to the propagation characteristics of RF signals through multipath channels, a receiver may receive multiple “RF signals” corresponding to each transmitted RF signal. The same transmitted RF signal on different paths between the transmitter and receiver may be referred to as a “multipath” RF signal. As used herein, an RF signal may also be referred to as a “wireless signal” or simply a “signal” where the context clearly indicates that the term “signal” refers to a wireless signal or an RF signal.

[0036] Figure 1An example wireless communication system 100 according to various aspects of this disclosure is illustrated. The wireless communication system 100 (which may also be referred to as a wireless wide area network (WWAN)) may include various base stations 102 (labeled "BS") and various UEs 104. Base station 102 may include macro cell base stations (high-power cellular base stations) and / or small cell base stations (low-power cellular base stations). In one aspect, macro cell base station 102 may include eNB and / or ng-eNB (wherein wireless communication system 100 corresponds to an LTE network) or gNB (wherein wireless communication system 100 corresponds to an NR network) or a combination of both, and small cell base stations may include femtocells, picocells, microcells, etc.

[0037] Base station 102 can collectively form a RAN and interface with core network 170 (e.g., evolved packet core (EPC) or 5G core (5GC)) via backhaul link 122, and interface with one or more location servers 172 (e.g., location management function (LMF) or secure user plane positioning (SUPL) positioning platform (SLP)) via core network 170. Location server 172 can be part of core network 170 or can be external to core network 170. Location server 172 can be integrated with base station 102. UE 104 can communicate with location server 172 directly or indirectly. For example, UE 104 can communicate with location server 172 via base station 102 currently serving UE 104. UE 104 can also communicate with location server 172 via another path, such as via application server (not shown), via another network, such as via wireless local area network (WLAN) access point (AP) (e.g., AP 150 described below), etc. For signaling purposes, communication between UE 104 and location server 172 can be represented as an indirect connection (e.g., via core network 170, etc.) or a direct connection (e.g., as shown via direct connection 128), wherein intermediate nodes (if present) are omitted from the signaling diagram for clarity.

[0038] In addition to other functions, base station 102 may perform functions associated with one or more of the following: transmitting user data, radio channel encryption and decryption, integrity protection, header compression, mobility control functions (e.g., handover, dual connectivity), inter-cell interference coordination, connection establishment and release, load balancing, distribution of non-access stratum (NAS) messages, NAS node selection, synchronization, RAN sharing, multimedia broadcast multicast service (MBMS), subscriber and equipment tracking, RAN information management (RIM), paging, location, and delivery of warning messages. Base stations 102 may communicate with each other directly or indirectly (e.g., via EPC / 5GC) on backhaul link 134, which may be wired or wireless.

[0039] Base station 102 can wirelessly communicate with UE 104. Each base station in base station 102 can provide communication coverage for a corresponding geographic coverage area 110. In one aspect, one or more cells can be supported by base station 102 in each geographic coverage area 110. A “cell” is a logical communication entity used to communicate with a base station (e.g., via a frequency resource, which is referred to as a carrier frequency, component carrier, carrier, frequency band, etc.) and can be associated with an identifier (e.g., Physical Cell Identifier (PCI), Enhanced Cell Identifier (ECI), Virtual Cell Identifier (VCI), Cell Global Identifier (CGI), etc.) used to distinguish cells operating via the same or different carrier frequencies. In some cases, different cells can be configured according to different protocol types that can provide access for different types of UEs (e.g., Machine Type Communication (MTC), Narrowband IoT (NB-IoT), Enhanced Mobile Broadband (eMBB), or other protocol types). Because a cell is supported by a specific base station, the term “cell” can refer to one or both of the logical communication entity and the base station that supports it, depending on the context. In some cases, the term "cell" can also refer to the geographic coverage area of ​​a base station (e.g., a sector), as long as the carrier frequency can be detected and used for communication within a portion of the geographic coverage area 110.

[0040] While the geographic coverage areas 110 of adjacent macro cell base stations 102 may partially overlap (e.g., in handover areas), some areas within geographic coverage areas 110 may substantially overlap with larger geographic coverage areas 110. For example, a small cell base station 102' (labeled "SC" for "small cell") may have a geographic coverage area 110' that substantially overlaps with the geographic coverage areas 110 of one or more macro cell base stations 102. A network that includes both small cell base stations and macro cell base stations can be referred to as a heterogeneous network. A heterogeneous network may also include a home eNB (HeNB) that can provide service to a restricted group referred to as a Closed Subscriber Group (CSG).

[0041] The communication link 120 between base station 102 and UE 104 may include uplink (also known as reverse link) transmission from UE 104 to base station 102 and / or downlink (DL) (also known as forward link) transmission from base station 102 to UE 104. The communication link 120 may use MIMO antenna techniques, including spatial multiplexing, beamforming, and / or transmit diversity. The communication link 120 may use one or more carrier frequencies. Carrier allocation may be asymmetric for the downlink and uplink (e.g., more or fewer carriers may be allocated to the downlink compared to the uplink).

[0042] The wireless communication system 100 may also include a WLAN access point (AP) 150 that communicates with a wireless local area network (WLAN) station (STA) 152 via a communication link 154 in unlicensed spectrum (e.g., 5 GHz). When communicating in unlicensed spectrum, the WLAN STA 152 and / or WLAN AP 150 may perform a free channel assessment (CCA) or listen-before-talk (LBT) process before communication to determine whether the channel is available.

[0043] Small cell base station 102' can operate in licensed and / or unlicensed spectrum. When operating in unlicensed spectrum, small cell base station 102' can employ LTE or NR technology and use the same 5GHz unlicensed spectrum as WLAN AP 150. Small cell base station 102' employing LTE / 5G in unlicensed spectrum can improve the coverage and / or increase the capacity of the access network. NR in unlicensed spectrum may be referred to as NR-U. LTE in unlicensed spectrum may be referred to as LTE-U, Licensed Assisted Access (LAA), or MULTEFIRE. ® .

[0044] The wireless communication system 100 may also include an mmW base station 180, which can operate in millimeter-wave (mmW) frequencies and / or near-mmW frequencies to communicate with the UE 182. Extremely high frequency (EHF) is a portion of the electromagnetic spectrum that contains radio frequency (RF). EHF has a range of 30 GHz to 300 GHz, with wavelengths between 1 mm and 10 mm. Radio waves in this band are referred to as millimeter waves. Near-mmW extends down to frequencies of 3 GHz with wavelengths of 100 mm. Ultra-high frequency (SHF) bands extend between 3 GHz and 30 GHz, and are also referred to as centimeter waves. Communication using mmW / near-mmW radio bands has high path loss and relatively short range. The mmW base station 180 and the UE 182 can utilize beamforming (transmit and / or receive) on the mmW communication link 184 to compensate for the extremely high path loss and short range. Furthermore, it should be understood that in alternative configurations, one or more base stations 102 may also use mmW or near-mmW and beamforming for transmission. Therefore, it should be understood that the foregoing examples are merely illustrative and should not be construed as limiting the various aspects disclosed herein.

[0045] Transmit beamforming is a technique used to focus RF signals in a specific direction. Traditionally, when a network node (e.g., a base station) broadcasts an RF signal, it broadcasts the signal in all directions (omnidirectionally). Using transmit beamforming, the network node determines where a given target device (e.g., a UE) is located (relative to the transmitting network node) and projects a stronger downlink RF signal in that specific direction, thus providing the receiving device with a faster and stronger RF signal (in terms of data rate). To change the directivity of the RF signal during transmission, the network node can control the phase and relative amplitude of the RF signal at each of one or more transmitters broadcasting the RF signal. For example, the network node can use an array of antennas (called a "phased array" or "antenna array") that forms an RF beam that can be "manipulated" to be pointed in different directions without actually moving the antennas. Specifically, RF currents from the transmitters are fed to individual antennas with the correct phase relationship, such that radio waves from the individual antennas add up in the desired direction to increase radiation, while canceling out in the undesired direction to suppress radiation.

[0046] Transmit beams can be quasi-co-located, meaning they appear to the receiver (e.g., the UE) as having the same parameters regardless of whether the network node's own transmit antennas are physically co-located. In NR, there are four types of quasi-co-located (QCL) relationships. Specifically, a given type of QCL relationship means that certain parameters of a second reference RF signal on a second beam can be derived based on information about the source reference RF signal on the source beam. Therefore, if the source reference RF signal is QCL type A, the receiver can use the source reference RF signal to estimate the Doppler shift, Doppler spread, average delay, and delay spread of the second reference RF signal transmitted on the same channel. If the source reference RF signal is QCL type B, the receiver can use the source reference RF signal to estimate the Doppler shift and Doppler spread of the second reference RF signal transmitted on the same channel. If the source reference RF signal is QCL type C, the receiver can use the source reference RF signal to estimate the Doppler shift and average delay of the second reference RF signal transmitted on the same channel. If the source reference RF signal is of type QCL D, the receiver can use the source reference RF signal to estimate the spatial reception parameters of a second reference RF signal transmitted on the same channel.

[0047] In receive beamforming, a receiver uses a receive beam to amplify an RF signal detected on a given channel. For example, the receiver may increase the gain setting of an antenna array in a particular direction and / or adjust the phase setting of the antenna array in a particular direction to amplify the RF signal received from that direction (e.g., increase its gain level). Therefore, when a receiver is described as performing beamforming in a certain direction, it means that the beam gain in that direction is high relative to the beam gain along other directions, or that the beam gain in that direction is the highest compared to the beam gain of all other receive beams available to the receiver in that direction. This results in a stronger received signal strength (e.g., reference signal received power (RSRP), reference signal received quality (RSRQ), signal-to-interference-plus-noise ratio (SINR), etc.) of the RF signal received from that direction.

[0048] The transmit and receive beams can be spatially correlated. Spatial correlation means that parameters for a second beam (e.g., transmit or receive beam) for a second reference signal can be derived based on information about a first beam (e.g., receive or transmit beam) for a first reference signal. For example, a UE can use a specific receive beam to receive a reference downlink reference signal (e.g., a synchronization signal block (SSB)) from a base station. The UE can then form a transmit beam for transmitting an uplink reference signal (e.g., a sounding reference signal (SRS)) to that base station based on the parameters of the receive beam.

[0049] It is important to note that, depending on the entity forming the "downlink" beam, the beam can be either a transmit beam or a receive beam. For example, if the base station is forming a downlink beam to transmit a reference signal to the UE, the downlink beam is a transmit beam. However, if the UE is forming a downlink beam, the downlink beam is a receive beam for receiving the downlink reference signal. Similarly, depending on the entity forming the "uplink" beam, the beam can be either a transmit beam or a receive beam. For example, if the base station is forming an uplink beam, the uplink beam is an uplink receive beam, while if the UE is forming an uplink beam, the uplink beam is an uplink transmit beam.

[0050] The electromagnetic spectrum is typically subdivided into various categories, bands, channels, etc., based on frequency / wavelength. In 5G NR, two initial operating bands have been designated as frequency ranges FR1 (410MHz to 7.125GHz) and FR2 (24.25GHz to 52.6GHz). It should be understood that although a portion of FR1 is greater than 6GHz, in various documents and articles, FR1 is often (interchangeably) referred to as the "sub-6GHz" band. A similar naming issue sometimes occurs with FR2, which is often (interchangeably) referred to as the "millimeter wave" band in documents and articles, although this differs from the designation used by the International Telecommunication Union.® Extremely high frequency (EHF) bands (30 GHz to 300 GHz) are designated as “millimeter wave” bands.

[0051] The frequencies between FR1 and FR2 are generally referred to as mid-band frequencies. Recent 5G NR studies have identified the operating bands used for these mid-band frequencies as the frequency range designation FR3 (7.125 GHz to 24.25 GHz). Bands falling within FR3 can inherit FR1 and / or FR2 characteristics, thus effectively extending the features of FR1 and / or FR2 to mid-band frequencies. Furthermore, higher frequency bands are currently being explored to extend 5G NR operation beyond 52.6 GHz. For example, three higher operating frequency bands have been identified as the frequency range designations FR4a or FR4-1 (52.6 GHz to 71 GHz), FR4 (52.6 GHz to 114.25 GHz), and FR5 (114.25 GHz to 300 GHz). Each of these higher frequency bands falls within the EHF band.

[0052] In light of the foregoing, unless otherwise specifically stated, it should be understood that, as used herein, the term "below 6 GHz" and the like can broadly refer to frequencies less than 6 GHz, within FR1, or including intermediate frequency band frequencies. Furthermore, unless otherwise specifically stated, it should be understood that, as used herein, the term "millimeter wave" and the like can broadly refer to frequencies that can include intermediate frequency band frequencies, within FR2, FR4, FR4-a or FR4-1 and / or FR5, or within the EHF band.

[0053] In multi-carrier systems such as 5G, one of the carrier frequencies is referred to as the "primary carrier," "anchor carrier," "primary serving cell," or "PCell," and the remaining carrier frequencies are referred to as "secondary carriers," "secondary serving cells," or "SCell." In carrier aggregation, the anchor carrier is the carrier operating on the primary frequency (e.g., FR1) used by UE 104 / 182 and the cell, where UE 104 / 182 performs an initial Radio Resource Control (RRC) connection establishment procedure or initiates an RRC connection re-establishment procedure. The primary carrier carries all common and UE-specific control channels and can be a carrier on a licensed frequency (however, this is not always the case). The secondary carrier is a carrier operating on a second frequency (e.g., FR2) that can be configured and used to provide additional radio resources once an RRC connection is established between UE 104 and the anchor carrier. In some cases, the secondary carrier can be a carrier on an unlicensed frequency. Secondary carriers may contain only the necessary signaling information and signals. For example, since the primary uplink and primary downlink carriers are typically UE-specific, the UE-specific signaling information and signals may not be present in the secondary carrier. This means that different UEs 104 / 182 within a cell can have different downlink primary carriers. The same applies to the uplink primary carrier. The network can change the primary carrier of any UE 104 / 182 at any time. This is done, for example, to balance the load on different carriers. Since a "serving cell" (whether PCell or SCell) corresponds to the carrier frequency / component carrier through which a base station communicates, the terms "cell," "serving cell," "component carrier," and "carrier frequency" can be used interchangeably.

[0054] For example, still refer to Figure 1 One of the frequencies used by macro cell base station 102 can be an anchor carrier (or "PCell"), and the other frequencies used by macro cell base station 102 and / or mmW base station 180 can be secondary carriers ("SCell"). Simultaneous transmission and / or reception on multiple carriers allows UE 104 / 182 to significantly increase its data transmission and / or reception rates. For example, compared to the data rate obtained by a single 20MHz carrier, two aggregated 20MHz carriers in a multi-carrier system would theoretically result in a doubling of the data rate (i.e., 40MHz).

[0055] exist Figure 1 In the example, the UE shown (for simplicity, in) Figure 1Any UE (shown as a single UE 104) can receive signal 124 from one or more Earth-orbiting spacecraft (SV) 112 (e.g., satellites). In one aspect, SV 112 may be part of a satellite positioning system that allows UE 104 to use as an independent source of location information. Satellite positioning systems typically include a system of transmitters (e.g., SV 112) positioned such that a receiver (e.g., UE 104) can determine its location on or above the Earth based at least in part on positioning signals (e.g., signal 124) received from the transmitters. Such transmitters typically transmit signals marked with a set number of repeating pseudo-random noise (PN) codes. While typically located in SV 112, transmitters may sometimes be located at ground-based control stations, base stations 102, and / or other UEs 104. UE 104 may include one or more dedicated receivers specifically designed to receive signal 124 in order to derive geographic location information from SV 112.

[0056] In a satellite positioning system, the use of signal 124 can be enhanced by various satellite-based augmentation systems (SBAS), which may be associated with or otherwise made capable of being used with one or more global and / or regional navigation satellite systems. For example, SBAS may include augmentation systems that provide integrity information, differential correction, etc., such as Wide Area Augmentation System (WAAS), European Geostationary Navigation Overlap Service (EGNOS), Multifunctional Satellite Augmentation System (MSAS), GPS-assisted geographic augmentation navigation, or GPS and geographic augmentation navigation system (GAGAN). Therefore, as used herein, a satellite positioning system may include any combination of one or more global and / or regional navigation satellites associated with such one or more satellite positioning systems.

[0057] On one hand, SV 112 may additionally or alternatively be part of one or more non-terrestrial networks (NTNs). In an NTN, SV 112 connects to an earth station (also referred to as a ground station, NTN gateway, or gateway), which in turn connects to elements in the 5G network, such as the modified base station 102 (without a ground antenna) or network nodes in a 5GC. This element then provides access to other elements in the 5G network and ultimately to entities outside the 5G network, such as internet web servers and other user equipment. Thus, as a replacement or supplement to communication signals from the ground base station 102, UE 104 can receive communication signals (e.g., signal 124) from SV 112.

[0058] Leveraging the increased data rates and reduced latency of NR (Radio Frequency I / O), vehicle-to-everything (V2X) communication technology is being implemented to support Intelligent Transportation Systems (ITS) applications, such as wireless communication between vehicles (V2V), between vehicles and roadside infrastructure (V2I), and between vehicles and pedestrians (V2P). The goal is to enable vehicles to sense their surroundings and communicate that information to other vehicles, infrastructure, and personal mobile devices. This type of vehicle communication will achieve safety, mobility, and environmental improvements that current technologies cannot provide. Once fully realized, this technology is expected to reduce collisions involving undamaged vehicles by 80%.

[0059] Still referencing Figure 1 The wireless communication system 100 may include multiple V-UEs 160, which can communicate with base station 102 on communication link 120 using a Uu interface (i.e., the air interface between the UE and the base station). V-UEs 160 can also communicate directly with each other on wireless sidelink 162, with roadside unit (RSU) 164 (roadside access point) on wireless sidelink 166, or with sidelink-capable UE 104 on wireless sidelink 168 using a PC5 interface (i.e., the air interface between UEs with sidelink capability). A wireless sidelink (or simply "sidelink") is an adaptation of core cellular network (e.g., LTE, NR) standards that allows direct communication between two or more UEs without requiring communication through a base station. Sidelink communication can be unicast or multicast and can be used for device-to-device (D2D) media sharing, V2V communication, V2X communication (e.g., cellular V2X (cV2X) communication, enhanced V2X (eV2X) communication, emergency rescue applications, etc. One or more V-UEs in a group of V-UEs 160 utilizing sidelink communication may be within the geographic coverage area 110 of base station 102. Other V-UEs 160 in such a group may be outside the geographic coverage area 110 of base station 102, or may be unable to receive transmissions from base station 102 for other reasons. In some cases, the groups of V-UEs 160 communicating via sidelink communication may utilize a one-to-many (1:M) system, where each V-UE 160 transmits to every other V-UE 160 in the group. In some cases, base station 102 facilitates the scheduling of resources for sidelink communication. In other cases, sidelink communication is performed between V-UEs 160 without involving base station 102.

[0060] On one hand, sidelinks 162, 166, and 168 can operate via a wireless communication medium of interest, which can be shared with other vehicles and / or infrastructure access points and other wireless communications between other RATs. “Medium” can include one or more time, frequency, and / or space communication resources (e.g., covering one or more channels across one or more carriers) associated with wireless communication between one or more transmitter / receiver pairs.

[0061] On one hand, sidelinks 162, 166, and 168 can be cV2X links. First-generation cV2X has been standardized in LTE, and the next generation is expected to be defined in NR. cV2X is a cellular technology that also enables device-to-device communication. In the United States and Europe, cV2X is expected to operate in licensed ITS bands below 6 GHz. Other bands may be allocated in other countries. Thus, as a specific example, the medium of interest utilized by sidelinks 162, 166, and 168 may correspond to at least a portion of licensed ITS bands below 6 GHz. However, this disclosure is not limited to this band or cellular technology.

[0062] On one hand, sidelinks 162, 166, and 168 can be Dedicated Short-Range Communications (DSRC) links. DSRC is a one-way or two-way short-to-medium-range wireless communication protocol that uses the Vehicle Environment Wireless Access (WAVE) protocol (also known as IEEE 802.11p) for V2V, V2I, and V2P communications. IEEE 802.11p is an approved modification of the IEEE 802.11 standard and operates in the licensed ITS band of 5.9 GHz (5.85 GHz to 5.925 GHz) in the United States. In Europe, IEEE 802.11p operates in the ITS G5A band (5.875 GHz to 5.905 MHz). Other bands may be allocated in other countries. The V2V communications briefly described above occur on a secure channel, which in the United States is typically a 10 MHz channel dedicated to security purposes. The remainder of the DSRC band (total bandwidth of 75MHz) is intended for other services of interest to drivers, such as road rules, toll collection, parking automation, etc. Therefore, as a specific example, the media of interest utilized by side links 162, 166, and 168 may correspond to at least a portion of the licensed ITS band at 5.9GHz.

[0063] Alternatively, the medium of interest may correspond to at least a portion of unlicensed frequency bands shared among various RATs. While different licensed frequency bands have been reserved for certain communication systems (e.g., by government entities such as the U.S. Federal Communications Commission (FCC), these systems (particularly those employing small cell access points) have recently expanded their operations to unlicensed National Information Infrastructure (U-NII) bands used by wireless local area network (WLAN) technologies, most notably the IEEE 802.11x WLAN technology commonly referred to as "Wi-Fi"). Example systems of this type include various variants of CDMA, TDMA, FDMA, orthogonal FDMA (OFDMA), single-carrier FDMA (SC-FDMA), and so on.

[0064] Communication between V-UEs 160 is referred to as V2V communication, communication between V-UE 160 and one or more RSUs 164 is referred to as V2I communication, and communication between V-UE 160 and one or more UEs 104 (where these UEs 104 are P-UEs) is referred to as V2P communication. V2V communication between V-UEs 160 may include information such as the location, speed, acceleration, heading, and other vehicle data of these V-UEs 160. V2I information received at a V-UE 160 from the one or more RSUs 164 may include, for example, road rules, parking automation information, etc. V2P communication between V-UE 160 and UE 104 may include information such as the location, speed, acceleration, and heading of V-UE 160, and the location, speed (e.g., in the case where UE 104 is carried by a cyclist), and heading of UE 104.

[0065] It should be noted that, although Figure 1 Only two UEs in the UE list are exemplified as V-UEs (V-UE 160), but any UE in the exemplified UEs (e.g., UE 104, 152, 182, 190) can be V-UEs. Furthermore, although only these V-UEs 160 and a single UE 104 have been exemplified as connected via a sidelink, Figure 1Any of the illustrated UEs, whether V-UE, P-UE, etc., can perform sidelink communication. Furthermore, although only UE 182 is described as capable of beamforming, any of the illustrated UEs (including V-UE 160) can perform beamforming. When V-UE 160 is capable of beamforming, it can beamform towards each other (i.e., towards other V-UEs 160), towards RSU 164, towards other UEs (e.g., UEs 104, 152, 182, 190), etc. Therefore, in some cases, V-UE 160 can utilize beamforming on sidelinks 162, 166, and 168.

[0066] The wireless communication system 100 may also include one or more UEs (such as UE 190) indirectly connected to one or more communication networks via one or more device-to-device (D2D) peer-to-peer (P2P) links. Figure 1 In one example, UE 190 has a D2D P2P link 192 with one of UEs 104 connected to one of the base stations 102 (e.g., UE 190 can indirectly obtain cellular connectivity through this D2D P2P link), and a D2D P2P link 194 with a WLANSTA 152 connected to a WLAN AP 150 (UE 190 can indirectly obtain WLAN-based Internet connectivity through this D2D P2P link). In one example, D2D P2P links 192 and 194 can utilize any known D2D RAT (such as LTE Direct (LTE-D), Wi-Fi Direct). ® ,Bluetooth ® (etc.) to support this. As another example, D2D P2P links 192 and 194 can be side links, as described above with reference to side links 162, 166 and 168.

[0067] Figure 2AAn example wireless network architecture 200 is illustrated. For instance, the 5GC 210 (also referred to as the Next Generation Core (NGC)) can be functionally viewed as control plane (C-plane) functions 214 (e.g., UE registration, authentication, network access, gateway selection, etc.) and user plane (U-plane) functions 212 (e.g., UE gateway functions, access to data networks, IP routing, etc.), which work together to form the core network. The user plane interface (NG-U) 213 and the control plane interface (NG-C) 215 connect the gNB 222 to the 5GC 210, specifically to user plane functions 212 and control plane functions 214, respectively. In an additional configuration, the ng-eNB 224 can also connect to the 5GC 210 via the NG-C 215 to the control plane function 214 and the NG-U 213 to the user plane function 212. Furthermore, the ng-eNB 224 can communicate directly with the gNB 222 via a backhaul connection 223. In some configurations, the next-generation RAN (NG-RAN) 220 may have one or more gNBs 222, while other configurations include one or more of both ng-eNBs 224 and gNBs 222. Either or both of the gNBs 222 or ng-eNBs 224 can communicate with one or more UEs 204 (e.g., any of the UEs described herein).

[0068] Another optional aspect may include a location server 230, which can communicate with the 5GC 210 to provide location assistance to the UE 204. The location server 230 may be implemented as multiple separate servers (e.g., physically separate servers, different software modules on a single server, different software modules distributed across multiple physical servers, etc.), or alternatively, each may correspond to a single server. The location server 230 may be configured to support one or more location services for the UE 204, which may be connected to the location server 230 via the core network, the 5GC 210, and / or via the Internet (not illustrated). Furthermore, the location server 230 may be integrated into a component of the core network, or alternatively, may be located outside the core network (e.g., a third-party server, such as an original equipment manufacturer (OEM) server or a service server).

[0069] Figure 2B Another example wireless network architecture 240.5GC 260 is illustrated (which can be used with...). Figure 2AThe 5GC 210 (corresponding to 5GC 210) can be functionally considered as a control plane function provided by the Access and Mobility Management Function (AMF) 264 and a user plane function provided by the User Plane Function (UPF) 262, which work together to form the core network (i.e., 5GC 260). The functions of AMF 264 include: registration management, connection management, reachability management, mobility management, lawful interception, transmission of session management (SM) messages between one or more UEs 204 (e.g., any of the UEs described herein) and the Session Management Function (SMF) 266, a transparent proxy service for routing SM messages, access authentication and access authorization, transmission of short message service (SMS) messages between UE 204 and the Short Message Service Function (SMSF) (not shown), and Secure Anchoring Functionality (SEAF). AMF 264 also interacts with the Authentication Server Function (AUSF) (not shown) and UE 204 and receives an intermediate key established as a result of the UE 204's authentication process. In the case of UMTS (Universal Mobile Telecommunications System) Subscriber Identity Module (USIM) authentication, AMF 264 retrieves security material from the AMF. AMF 264 also includes Security Context Management (SCM). The SCM receives a key from the SEAF and uses this key to derive an access network-specific key. AMF 264 functionality also includes location service management for regulatory services, transmission of location service messages between UE 204 and Location Management Function (LMF) 270 (which acts as location server 230), transmission of location service messages between NG-RAN 220 and LMF 270, Evolved Packet System (EPS) bearer identifier allocation for EPS interoperability, and UE 204 mobility event notification. Furthermore, AMF 264 also supports non-3GPP... ® (Third Generation Partner Program) Access network functionality.

[0070] The functions of UPF 262 include: acting as an anchor point for intra-RAT / inter-RAT mobility (where applicable), acting as an external Protocol Data Unit (PDU) session point interconnecting to a data network (not shown), providing packet routing and forwarding, packet inspection, user plane policy rule enforcement (e.g., strobing, redirection, traffic steering), lawful eavesdropping (user plane collection), traffic usage reporting, quality of service (QoS) processing for the user plane (e.g., uplink / downlink rate enforcement, reflective QoS marking in the downlink), uplink traffic verification (Service Data Flow (SDF) to QoS flow mapping), transport-level packet marking in the uplink and downlink, downlink packet buffering and downlink data notification triggering, and delivering and forwarding one or more "end markers" to the source RAN node. UPF 262 can also support the delivery of location service messages between UE 204 and location servers (such as SLP 272) on the user plane.

[0071] The functions of SMF 266 include session management, UE Internet Protocol (IP) address allocation and management, selection and control of user plane functions, service orientation configuration at UPF 262 for routing services to the correct destination, partial control of policy enforcement and QoS, and downlink data notification. The interface through which SMF 266 communicates with AMF 264 is called the N11 interface.

[0072] Another optional aspect may include an LMF 270, which can communicate with the 5GC 260 to provide location assistance to the UE 204. The LMF 270 can be implemented as multiple separate servers (e.g., physically separate servers, different software modules on a single server, different software modules distributed across multiple physical servers, etc.), or alternatively, each can correspond to a single server. The LMF 270 can be configured to support one or more location services for the UE 204, which can connect to the LMF 270 via the core network, the 5GC 260, and / or via the Internet (not illustrated). SLP 272 can support similar functions to LMF 270, but while LMF 270 can communicate with AMF 264, NG-RAN 220, and UE 204 on the control plane (e.g., using interfaces and protocols designed to transmit signaling messages rather than voice or data), SLP 272 can communicate with UE 204 and external clients (e.g., third-party server 274) on the user plane (e.g., using protocols designed to carry voice and / or data, such as Transmit Control Protocol (TCP) and / or IP).

[0073] Another optional aspect may include a third-party server 274, which can communicate with LMF 270, SLP 272, 5GC 260 (e.g., via AMF 264 and / or UPF 262), NG-RAN 220, and / or UE 204 to obtain location information (e.g., location estimation) of UE 204. Therefore, in some cases, the third-party server 274 may be referred to as a Location Services (LCS) client or an external client. The third-party server 274 may be implemented as multiple separate servers (e.g., physically separate servers, different software modules on a single server, different software modules distributed across multiple physical servers, etc.), or alternatively, each may correspond to a single server.

[0074] User plane interface 263 and control plane interface 265 connect 5GC 260, and specifically connect UPF 262 and AMF 264 to one or more gNB 222 and / or ng-eNB 224 in NG-RAN 220. The interface between gNB 222 and / or ng-eNB 224 and AMF 264 is referred to as the "N2" interface, while the interface between gNB 222 and / or ng-eNB 224 and UPF 262 is referred to as the "N3" interface. The gNB 222 and / or ng-eNB 224 of NG-RAN 220 can communicate directly with each other via backhaul connection 223, referred to as the "Xn-C" interface. One or more of gNB 222 and / or ng-eNB 224 can communicate with one or more UEs 204 via a radio interface referred to as the "Uu" interface.

[0075] The functionality of the gNB 222 is divided among the gNB Central Unit (gNB-CU) 226, one or more gNB Distributed Units (gNB-DU) 228, and one or more gNB Radio Units (gNB-RU) 229. The gNB-CU 226 is a logical node that includes base station functions other than those specifically allocated to the gNB-DU 228, including user data delivery, mobility control, radio access network sharing, location, session management, etc. More specifically, the gNB-CU 226 typically hosts the Radio Resource Control (RRC), Serving Data Adaptation Protocol (SDAP), and Packet Data Convergence Protocol (PDCP) protocols of the gNB 222. The gNB-DU 228 is a logical node that typically hosts the Radio Link Control (RLC) and Media Access Control (MAC) layers of the gNB 222. Its operation is controlled by the gNB-CU 226. One gNB-DU 228 can support one or more cells, and a cell is supported by only one gNB-DU 228. The interface 232 between gNB-CU 226 and one or more gNB-DU 228 is referred to as the "F1" interface. The physical (PHY) layer functionality of gNB 222 is typically managed by one or more independent gNB-RU 229s, which perform functions such as power amplification and signal transmission / reception. The interface between gNB-DU 228 and gNB-RU 229 is referred to as the "Fx" interface. Therefore, UE 204 communicates with gNB-CU 226 via the RRC, SDAP, and PDCP layers, with gNB-DU 228 via the RLC and MAC layers, and with gNB-RU 229 via the PHY layer.

[0076] Modern motor vehicles increasingly incorporate technologies that help drivers avoid drifting into adjacent lanes or making unsafe lane changes (e.g., Lane Departure Warning (LDW)), technologies that warn drivers of other vehicles behind them when a vehicle is reversing, or technologies that automatically brake if the vehicle in front of it suddenly stops or slows down (e.g., Forward Collision Warning (FCW)), and so on. The continued evolution of automotive technology aims to provide even greater safety benefits and ultimately deliver automated driving systems (ADS) capable of taking control of the entire driving task without user intervention.

[0077] There are six defined levels for achieving full automation. At Level 0, the human driver performs all driving. At Level 1, the Advanced Driver Assistance System (ADAS) on the vehicle may sometimes assist the human driver in steering or braking / acceleration, but not simultaneously. At Level 2, the ADAS on the vehicle can, in some situations, effectively control both steering and braking / acceleration simultaneously. The human driver must maintain full attention and perform the remaining driving tasks at all times. At Level 3, the ADAS on the vehicle can perform all aspects of driving tasks in certain situations. In these situations, the human driver must be prepared to relinquish control when requested by the ADAS. In all other situations, the human driver performs the driving tasks. At Level 4, the ADAS on the vehicle can perform all driving tasks and monitor the driving environment, essentially performing all driving in some situations. In these situations, human attention is not required. At Level 5, the ADAS on the vehicle can perform all driving in all situations. The human occupant is merely a passenger and is never involved in driving.

[0078] Autonomous and semi-autonomous driving safety technologies use a combination of hardware (sensors, cameras, and radar) and software to help vehicles identify certain safety risks so that they can warn the driver to take action (in the case of ADAS) or act autonomously (in the case of ADS) to avoid a collision. Vehicles equipped with ADAS or ADS include one or more camera sensors mounted on the vehicle that capture images of the scene in front of the vehicle, and possibly behind and to the sides. Radar systems can also be used to detect objects along the road and possibly behind and to the sides of the vehicle. Radar systems use RF waves to determine the range, direction, speed, and / or height of objects along the road. More specifically, a transmitter sends pulses of RF waves that bounce off any object in its path. The pulses reflected from the object return a small fraction of the energy of the RF waves to a receiver, which is typically located at the same location as the transmitter. Cameras and radars are typically oriented to capture their respective versions of the same scene.

[0079] Processors within vehicles (such as digital signal processors (DSPs)) analyze captured camera images and radar frames, attempting to identify objects within the captured scene. These objects can be other vehicles, pedestrians, road signs, objects on the road, etc. Radar systems provide reasonably accurate measurements of object distance and velocity under various weather conditions. However, radar systems typically have insufficient resolution to identify the features of detected objects. Camera sensors, on the other hand, usually provide sufficient resolution to identify object features. Cues about object shape and appearance extracted from captured images can provide sufficient characteristics for classifying different objects. Given the complementary properties of two sensors, data from both sensors can be combined (called "fusion") within a single system for improved performance.

[0080] To further enhance ADAS and ADS systems, especially at Level 3 and higher, autonomous and semi-autonomous vehicles can utilize high-definition (HD) map datasets. These datasets contain significantly more detailed information and true-ground accuracy than found in current conventional resources. Such HD maps provide accuracy within an absolute range of 7cm to 10cm, a highly detailed catalog of all road-related fixed physical assets, such as road lanes, road edges, shoulders, dividers, traffic signals, signs, paint markings, poles, and other data that aids in the safe navigation of autonomous / semi-autonomous vehicles on roads and at intersections. HD maps also provide electronic horizon prediction awareness, enabling autonomous / semi-autonomous vehicles to know what lies ahead.

[0081] It should be noted that autonomous or semi-autonomous vehicles can be, but do not have to be, V-UEs. Similarly, V-UEs can be, but do not have to be, autonomous or semi-autonomous vehicles. Autonomous or semi-autonomous vehicles are those equipped with ADAS or ADS. V-UEs are vehicles with cellular connectivity to 5G or other cellular networks. Autonomous or semi-autonomous vehicles that use or are able to use cellular technologies for positioning and / or navigation are V-UEs.

[0082] Now for reference Figure 3AAn example is illustrated of a V2X-enabled vehicle 300 (referred to as a "self-driving vehicle" or "primary vehicle"), which includes a radar camera sensor module 320 located in an interior compartment behind a windshield 362 of the V2X-enabled vehicle 300. The radar camera sensor module 320 includes a radar assembly configured to transmit radar signals through the windshield 362 within a horizontal coverage area 365 (shown by dashed lines) and to receive reflected radar signals reflected from any object within the horizontal coverage area 365. The radar camera sensor module 320 also includes a camera assembly for capturing images based on light waves seen and captured through the windshield 362 within the horizontal coverage area 360 (shown by dashed lines).

[0083] Although Figure 3A The example shown illustrates a co-located component where the radar and camera components share a housing; however, it should be understood that they can be individually housed in different locations within a V2X-enabled vehicle 300. For example, the camera could be... Figure 3A The location is shown, and the radar component can be located in the grille or front bumper of a vehicle 300 with V2X capability. Additionally, although... Figure 3A An example is shown of a radar camera sensor module 320 located behind the windshield 362, but it could alternatively be located in the top sensor array or elsewhere. Furthermore, although... Figure 3A Only a single radar camera sensor module 320 is illustrated, but it should be understood that a vehicle 300 with V2X capability may have multiple radar camera sensor modules 320 pointing in different directions (side, front, rear, etc.). The various radar camera sensor modules 320 may be under the vehicle's "skin" (e.g., behind the windshield 362, door panels, bumpers, grille, etc.) or within a top-mounted sensor array.

[0084] The radar camera sensor module 320 can detect one or more objects relative to the V2X-enabled vehicle 300 (or detect no objects). Figure 3A In the example, two objects, vehicles 370 and 380, are detectable by the radar camera sensor module 320 within horizontal coverage areas 360 and 365. The radar camera sensor module 320 can estimate parameters (attributes) of the detected objects, such as location, range, orientation, speed, size, and classification (e.g., vehicle, pedestrian, road sign, etc.). The radar camera sensor module 320 can be used as an onboard device on a V2X-enabled vehicle 300 for automotive safety applications such as adaptive cruise control (ACC), forward collision warning (FCW), collision mitigation or avoidance via automatic braking, and low-light braking (LDW).

[0085] Co-location of cameras and radar allows these components to share electronics and signal processing, and particularly enables early radar-camera data fusion. For example, radar and camera can be integrated onto a single board. Joint radar-camera alignment techniques can be used to align both the radar and camera. However, co-location of radar and camera is not required for implementing the techniques described herein.

[0086] Figure 3B An onboard computer (OBC) 380 of a V2X-enabled vehicle 300 is illustrated according to various aspects of this disclosure. In one aspect, the OBC 380 may be part of an ADAS or ADS. The OBC 380 may also be a V-UE of the V2X-enabled vehicle 300. The OBC 380 includes a non-transitory computer-readable storage medium (i.e., memory 304) and one or more processors 306 communicating with the memory 304 via a data bus 308. The memory 304 includes one or more storage modules storing computer-readable instructions executable by the one or more processors 306 to perform the functions of the OBC 380 described herein. For example, the combination of one or more processors 306 and memory 304 can implement various operations described herein.

[0087] One or more radar camera sensor modules 320 are coupled to OBC 380 (for simplicity, in Figure 3B Only one radar camera sensor module is shown in the diagram. In some aspects, the radar camera sensor module 320 includes at least one camera 312, at least one radar 314, and at least one optional light detection and ranging (LiDAR) sensor 316. The OBC 380 also includes one or more system interfaces 310 that connect one or more processors 306 to the radar camera sensor module 320 via a data bus 308, and optionally to other vehicle subsystems (not shown).

[0088] In at least some cases, the OBC 380 also includes one or more Wireless Wide Area Network (WWAN) transceivers 330 configured to communicate via one or more wireless communication networks (not shown) such as NR networks, LTE networks, and / or Global System for Mobile Communications (GSM) networks. The one or more WWAN transceivers 330 may be connected to one or more antennas (not shown) for communication with other network nodes (such as other V-UEs, pedestrian UEs, infrastructure access points, roadside units (RSUs), base stations (e.g., eNBs, gNBs), etc.) via at least one designated RAT (e.g., NR, LTE, GSM, etc.) through a wireless communication medium of interest (e.g., certain time / frequency resources in a specific spectrum). The one or more WWAN transceivers 330 may be configured in various ways to transmit and encode signals (e.g., messages, indications, information, etc.) according to the designated RAT and conversely, to receive and decode signals (e.g., messages, indications, information, pilots, etc.).

[0089] In at least some cases, the OBC 380 also includes one or more short-range wireless transceivers 340 (e.g., Wi-Fi transceivers, Bluetooth transceivers). ® Transceivers, etc.). One or more short-range wireless transceivers 340 may be connected to one or more antennas (not shown) for communicating with other network nodes (such as other V-UEs, pedestrian UEs, infrastructure access points, RSUs, etc.) via at least one designated RAT (e.g., cV2X), IEEE 802.11p (also known as Wireless Access for Vehicle Environments (WAVE)), Dedicated Short Range Communications (DSRC), etc.) through a wireless communication medium of interest. One or more short-range wireless transceivers 340 may be configured in various ways to transmit and encode signals (e.g., messages, indications, information, etc.) according to a designated RAT and conversely to receive and decode signals (e.g., messages, indications, information, pilots, etc.).

[0090] As used herein, a “transceiver” may include transmitter circuitry, receiver circuitry, or a combination thereof, but it is not necessary to provide both transmit and receive functionality in all designs. For example, in some designs, low-functionality receiver circuitry may be used to reduce costs when full communication is not necessary (e.g., simply providing a low-level sniffing receiver chip or similar circuitry).

[0091] In at least some cases, the OBC 380 also includes a Global Navigation Satellite System (GNSS) receiver 350. The GNSS receiver 350 may be connected to one or more antennas (not shown) for receiving satellite signals. The GNSS receiver 350 may include any suitable hardware and / or software for receiving and processing GNSS signals. The GNSS receiver 350 requests information and operation from other systems as appropriate and performs calculations necessary to determine the location of the vehicle 300 using measurements obtained through any suitable GNSS algorithm.

[0092] On one hand, the OBC 380 can utilize one or more WWAN transceivers 330 and / or one or more short-range wireless transceivers 340 to download one or more maps 302, which can then be stored in memory 304 and used for vehicle navigation. Map 302 can be one or more high-definition (HD) maps providing an accuracy of 7cm to 10cm in absolute range, a highly detailed catalog of all fixed physical assets associated with the road, such as road lanes, road edges, shoulders, dividers, traffic signals, signs, painted markings, poles, and other data that aids in the safe navigation of the V2X-enabled vehicle 300 on roads and at intersections. Map 302 can also provide electronic horizon prediction awareness, enabling the V2X-enabled vehicle 300 to know what lies ahead.

[0093] The V2X-enabled vehicle 300 may include one or more sensors 322, which may be coupled to one or more processors 306 via one or more system interfaces 310. The one or more sensors 322 may provide components for sensing or detecting information such as speed, heading (e.g., compass heading), headlight status, fuel consumption, etc., related to the state and / or environment of the V2X-enabled vehicle 300. By way of example, the one or more sensors 322 may include an odometer, speedometer, tachometer, accelerometer (e.g., microelectromechanical systems (MEMS) device), gyroscope, geomagnetic sensor (e.g., compass), altimeter (e.g., barometric altimeter), etc. Although shown as being located outside the OBC 380, some of these sensors 322 may be located on the OBC 380, and some may be located elsewhere within the V2X-enabled vehicle 300.

[0094] The OBC 380 may also include a driving strategy component 318. The driving strategy component 318 may be hardware circuitry that is part of or coupled to one or more processors 306, which, when executed, causes the OBC 380 to perform the functionality described herein. In other aspects, the driving strategy component 318 may be external to one or more processors 306 (e.g., part of a positioning processing system, integrated with another processing system, etc.). Alternatively, the driving strategy component 318 may be one or more memory modules stored in memory 304, which, when executed by one or more processors 306 (or a positioning processing system, another processing system, etc.), cause the OBC 380 to perform the functionality described herein. As a specific example, the driving strategy component 318 may include multiple positioning engines, a positioning engine aggregator, a sensor fusion module, etc. Figure 3B Possible locations for the driving strategy component 318 are illustrated. The driving strategy component may be part of, for example, memory 304, one or more processors 306, or any combination thereof, or may be a standalone component.

[0095] On the one hand, camera 312 can capture the observation area of ​​camera 312 at a certain periodic rate (such as in...). Figure 3A Image frames of the scene within the horizontal coverage area (illustrated as 360°) (also referred to herein as camera frames). Similarly, radar 314 can capture the observation area of ​​camera 314 (such as in...) at a certain periodic rate. Figure 3A The radar frames represent a scene within a horizontal coverage area (365°). Camera 312 and radar 314 capture their respective frames at the same or different periodic rates. Each camera and radar frame can be timestamped. Therefore, in cases where the periodic rates differ, the timestamps can be used to simultaneously or nearly simultaneously select the captured camera and radar frames for further processing (e.g., fusion).

[0096] For convenience, OBC 380 is... Figure 3B The example shown herein includes various components that can be configured according to the various examples described herein. However, it should be understood that the illustrated components may have different functionalities in different designs. In particular, Figure 3B Various components are optional in alternative configurations, and various aspects include configurations that may vary due to design choices, cost, equipment usage, or other considerations. For the sake of brevity, examples of the various alternative configurations are not provided herein, but will be readily understood by those skilled in the art.

[0097] Figure 3B The components can be implemented in various ways. In some specific implementations, Figure 3BThe components may be implemented in one or more circuits, such as, for example, one or more processors and / or one or more ASICs (which may include one or more processors). Here, each circuit may use and / or combine at least one memory component for storing information or executable code used by the circuit to provide that functionality. For example, some or all of the functionality represented by blocks 302 to 350 may be implemented by the processor and memory components of the OBC 380 (e.g., by executing appropriate code and / or by appropriate configuration of the processor components). For simplicity, various operations, actions, and / or functions are described herein as being performed “by the UE,” “by the OBC,” or “by the vehicle.” However, it should be understood that such operations, actions, and / or functions may actually be performed by specific components or combinations of components of the OBC 380 (such as one or more processors 306, one or more transceivers 330 and 340, memory 304, driving strategy component 318, etc.).

[0098] In autonomous or semi-autonomous driving scenarios, the autonomous vehicle needs to make various driving decisions, such as when to change lanes (e.g., to avoid obstacles, move to the exit lane, etc.), where to merge into traffic, and whether to overtake another vehicle. These types of decisions are called "driving policies" and can be executed by the OBC 380 (e.g., one or more processors 306, driving policy components 318, memory 304, etc.) based on information from the radar camera sensor module 320 and / or sensor 322.

[0099] Driving strategy involves trajectory prediction and route planning functionality. Trajectory prediction follows a data-driven approach that combines flashing light status information and trajectory history of other vehicles (referred to as "agents") around the self-driving vehicle with map geometry (e.g., from map 302). Graph-based neural networks learn multi-agent interactions, while the weighted multimodal distribution of trajectories represents the uncertainty of agent intent and motion. Randomized predicted trajectories are used for tree search and a dynamically programmed optimizer for risk minimizing self-manipulation. Note that tree search is only one approach; trajectory prediction can also be performed using graph-based neural networks, where the weights of the neural network can be updated to adjust which trajectories / paths remain feasible.

[0100] Route planning attempts to understand the probabilistic evolution of the world by exploring a belief space. Self-actions are defined by generating possible trajectories, and surrogate actions are defined by predicting inputs. Route planning efficiently prunes the search space (e.g., a search tree) and evaluates the risk and reward of candidate trajectories. The output is a coarse reference trajectory along with corresponding beliefs about the world and related semantics.

[0101] It should be noted that a driving trajectory is not necessarily a single driving maneuver (such as lane changing, braking, merging, etc.), but more precisely, it is a driving path that may be taken in the next few seconds to minutes. Therefore, a driving trajectory may include one or more planned driving maneuvers over a period of time from the next few seconds to the next few minutes.

[0102] Figure 4 This is a diagram 400 illustrating an example driving strategy pipeline according to various aspects of this disclosure. For example... Figure 4 As shown, at a high level, sensing and perception information (e.g., from camera 312, radar 314, lidar sensor 316, sensor 322) is fed into a Real World Model (RWM) block, which outputs map data (e.g., from map 302), object detection results (e.g., object detection results for both stationary and moving objects), trajectory predictions for detected moving objects, and the positions of vehicles to lane-level planner blocks, global trajectory search blocks, and local trajectory optimization blocks.

[0103] The lane-level planner block receives at least map data and driving objectives (e.g., lane change, merging, overtaking, etc.) from the RWM block and outputs the desired route plan [r] and advanced lane instructions to the global trajectory search block. The global trajectory search block generates a set of coarse reference trajectories (denoted as t_r) and a set of search and semantic parameters s_r based on the desired route plan [r] and information from the RWM block. The global trajectory search block outputs t_r and s_r to the local trajectory optimization block, which optimizes the set of coarse reference trajectories t_r and local reaction trajectories based on information from the RWM block to determine the set of optimized trajectories [t_o].

[0104] The arbitration block (e.g., within a local trajectory optimization block) selects the minimum-cost candidate trajectory t_c^ from the set of optimized trajectories [t_o] received from the local trajectory optimization block. The arbitration block will select the minimum cost candidate trajectory t_c^ The output is fed into a safety verification block (e.g., within a local trajectory optimization block), which verifies the minimum cost candidate trajectory t_c^. The safety of the candidate trajectory t_c is considered, and if it is safe, the candidate trajectory t_c is selected. As the ultimate "favored" trajectory t^ Output to the lateral control block, and set the trajectory t^ The speed output is sent to the longitudinal control block. Based on these inputs, the lateral and longitudinal control blocks output steering, throttle, and braking control signals to the corresponding vehicle systems.

[0105] Machine learning can be used to generate models that can facilitate various aspects associated with data processing. One specific application of machine learning involves the generation of models for predicting driving trajectories such as stopping, turning, and changing lanes.

[0106] Machine learning models are generally categorized as supervised or unsupervised. Supervised models can be further subdivided into regression models or classification models. Supervised learning involves learning a function that maps inputs to outputs based on example input-output pairs. For example, given a training dataset with two variables, age (input) and height (output), a supervised learning model can be generated to predict a person's height based on their age. In regression models, the output is continuous. An example of a regression model is linear regression, which simply attempts to find a line that best fits the data. Extensions of linear regression include multiple linear regression (e.g., finding a best-fitting plane) and multinomial regression (e.g., finding a best-fitting curve).

[0107] Another example of a machine learning model is the decision tree model. In a decision tree model, the tree structure is defined as having multiple nodes. Decisions are made to move from the root node at the top of the decision tree to a leaf node at the bottom (i.e., a node that has no other children). Generally, a higher number of nodes in a decision tree model is associated with higher decision accuracy.

[0108] Another example of a machine learning model is the decision forest. Random forests are an ensemble learning technique built on top of decision trees. Random forests involve creating multiple decision trees using a bootstrap dataset of the original data and randomly selecting a subset of variables at each step of the decision trees. The model then selects the pattern of all predictions from each decision tree. By relying on a "majority decision" model, the risk of errors from individual trees is reduced.

[0109] Another example of a machine learning model is a neural network (NN). A neural network is essentially a network of mathematical equations. It takes one or more input variables and produces one or more output variables by passing them through the network of equations. In other words, a neural network receives a vector of inputs and returns a vector of outputs.

[0110] Figure 5An example neural network 500 according to various aspects of this disclosure is illustrated. The neural network 500 includes an input layer "i" that receives "n" (or more) inputs (illustrated as "input 1", "input 2", and "input n"), one or more hidden layers (illustrated as hidden layers "h1", "h2", and "h3") for processing the inputs from the input layer, and an output layer "o" that provides "m" (or more) outputs (labeled as "output 1" and "output m"). The number of inputs "n", hidden layers "h", and outputs "m" may be the same or different. In some designs, hidden layers "h" may include linear functions and / or activation functions, with each node of a successive hidden layer (illustrated as a circle) processing the linear function and / or activation function from the node of the previous hidden layer.

[0111] In classification models, the output is discrete. An example of a classification model is logistic regression. Logistic regression is similar to linear regression, but it's used to model the probabilities of a finite number of outcomes (usually two). Essentially, it's a logistic equation created in a way that ensures the output value can only be between "0" and "1". Another example of a classification model is a support vector machine (SVM). For example, given data from two classes, an SVM will find a hyperplane, or boundary, that maximizes the margin between the two classes. Many hyperplanes can separate the two classes, but only one hyperplane maximizes the margin or distance between them. Another example of a classification model is Naive Bayes, based on Bayes' theorem. Other examples of classification models include decision trees, random forests, and neural networks, which are similar to the examples described above, except that the output is discrete rather than continuous.

[0112] Unlike supervised learning, unsupervised learning is used to derive inferences and find patterns from input data without referring to labeled results. Two examples of unsupervised learning models include clustering and dimensionality reduction.

[0113] Clustering is an unsupervised technique involving the grouping or clustering of data points. Clustering is commonly used for customer segmentation, fraud detection, and document classification. Common clustering techniques include k-means clustering, hierarchical clustering, mean-shift clustering, and density-based clustering. Dimensionality reduction is the process of reducing the number of random variables under consideration by obtaining a set of main variables. More simply, dimensionality reduction is the process of reducing the dimension of a feature set (or, even more simply, reducing the number of features). Most dimensionality reduction techniques can be categorized as feature elimination or feature extraction. An example of dimensionality reduction is called Principal Component Analysis (PCA). In its simplest sense, PCA involves projecting higher-dimensional data (e.g., three-dimensional) onto a smaller space (e.g., two-dimensional). This produces lower-dimensional (e.g., two-dimensional instead of three-dimensional) data while preserving all the original variables in the model.

[0114] Regardless of the machine learning model used, at a high level, the machine learning module (e.g., implemented by the processing system) can be configured to iteratively analyze the training input data (e.g., the presence of other vehicles, detection of road boundaries, etc.) and correlate the training input data with the output dataset (e.g., a set of possible or highly probable candidate trajectories of other vehicles and / or the self-vehicle), so that the same output dataset can be determined later when similar input data (e.g., from other target UEs at the same or similar locations) is provided.

[0115] When a self-driving vehicle (e.g., a vehicle 300 with V2X capabilities) needs to autonomously navigate through intersections, predicting the future trajectories of road users around the self-driving vehicle is crucial for determining a smooth and comfortable self-driving trajectory and preventing accidents or undesirable driving behaviors. The task of trajectory prediction is particularly significant when the self-driving vehicle needs to perform unprotected left and right turns (e.g., turns without stopping oncoming traffic, such as at stop signs or traffic lights), and for road users who violate traffic rules (e.g., turning from a non-turning lane, failing to stop at a stop sign, etc.). For tasks like these, the planned trajectory is largely conditioned on the predicted trajectories of interacting road users. More formally, the goal is to predict the possible future states of other road vehicles (referred to as "agents") observed around the self-driving vehicle, given its recent history and the map context around each involved / involved vehicle / agent. These trajectory predictions are performed by behavior planning and / or motion planning modules (see, for example...). Figure 4 These modules ultimately plan safe (e.g., collision-free) and comfortable trajectories for autonomous vehicles to execute.

[0116] This disclosure provides a technique for predicting the intentions (e.g., future trajectories) of road vehicles (agents) at urban intersections (and determining their probabilities). More specifically, a data-driven system collects and annotates, at specific time instances (referred to as “interest timestamps”), the turning classification labels (left turn, right turn, straight, U-turn) and flow classification labels (free flow, starting, decelerating, stopping) of each vehicle / agent of interest (or “interest agent”, “target agent”, or “target vehicle”, etc., which is the agent currently considered or processed by the predictive model) at and around the intersection. These turning and flow classifications are used as baseline ground truth labels for a multi-task (i.e., turning and flow) classification problem. Furthermore, along with the turning and flow classifications, the future trajectory of the target agent (e.g., the x and y coordinates of the target agent’s location sampled at 10 Hz for the next four seconds) is recorded. This is used as the baseline ground truth trajectory for regression.

[0117] Along with the ground truth labels for turns and flow, agent features of interest, neighboring agent features, and map features are recorded as input features to a machine learning model to learn and predict labels. Agent features may include one-second vehicle history sampled at, for example, 0.1-second intervals, and features may include (for both the target agent and neighboring agents) such as agent location (x... t y t Previous timestamp location (x) t-1 y t-1 ), speed (vx) t vy t ), angular velocity (ang) vel Parameters include acceleration (acc), flash status, time offset (t), and mask (the mask value indicates whether the agent exists in the scene at the interest timestamp). For map features (e.g., from map 302), the nearest lane center and lane boundaries (which may be referred to as "road geometry") are extracted and represented as a set of polylines (i.e., a set of points where line segments connect consecutive points), where each polyline contains at most a threshold number of points (e.g., 20 points). Each map polyline may include parameters such as location (x, y), direction (dir), and acceleration (acc). x ,dir y ), point type (e.g., lane center type or lane boundary type), and previous location (prev) x , prev y ) characteristics.

[0118] A deep learning network is provided that uses these proxy features and map features as tensors and learns to predict corresponding ground truth turns and flow classification labels, as well as corresponding future trajectories. The proposed network can have, for example, an encoder-decoder architecture. The encoder side can consist of one or more factorized attention modules. The first stage uses PointNet-like layers to encode information along the spatial (for map features) and temporal (for proxy features) dimensions. Then, the subsequent second stage uses one or more attention modules to encode information across polylines by using the target proxy polyline as the query and all polylines as keys and values. The decoder has a shared fully connected layer, followed by separate fully connected layers for each classification output to predict the label.

[0119] The proposed network also includes multiple regression heads, one for each turn (left turn, right turn, U-turn, straight ahead). For classification, a series of fully connected layers (called "LayerNorm" and "Rectified Linear Unit" (ReLU)) are used, followed by a final fully connected layer to output the predicted probability of lane change classification (left lane change, lane keeping, right lane change). For each regression head corresponding to the classification, a linear layer and batch norm are used with skip connections, followed by ReLU and a final linear layer to output the regression. Each output regression head can be composed of 2... It consists of N coordinate points, where "2" represents the x and y coordinates, and N represents the number of output sample points to represent the learned three-second (e.g.) future trajectory.

[0120] During the training phase, the baseline ground truth turn classifications (left turn, right turn, U-turn, straight ahead) along with the baseline ground truth trajectories are used to train one of the classification heads and the corresponding regression heads, thus conditioned the trajectory heads on the correct turn classification data. This joint backbone with separate heads for each trajectory category makes the deep learning algorithm highly efficient for predicting turns at urban intersections. The turn and flow probabilities, along with the generated trajectories for each turn category, are used to produce the final output and are consumed by the planning module (see, for example...). Figure 4 ).

[0121] The problem setting for predicting turns and flows at intersections helps to break down the challenging problem to fit the constraints of dataset size, computational budget, and training resources, while still allowing the network to generate good predictions for challenging intersection scenarios. The multi-task setting facilitates learning two different tasks using the same network while allowing a shared network backbone, which enables the network to learn common features, thus facilitating faster learning. The exposed network and features are designed to achieve high accuracy (e.g., precision, recall, F1 score) for both classification tasks (turns and flows). Furthermore, the network's inference latency is optimized to enable its deployment in real-time systems on roads.

[0122] Figures 6A to 6C An example encoder-decoder machine learning model architecture for predicting turning and flow classification labels for an agent (vehicle) of interest, according to various aspects of this disclosure, is illustrated. Specifically, Figure 6A Figure 610 illustrates the input tensor of an example encoder-decoder machine learning model architecture. Figure 6B This is a diagram 630 illustrating the encoder side of an example encoder-decoder machine learning model architecture, and... Figure 6C This is a diagram 650 illustrating the decoder side of an example encoder-decoder machine learning model architecture.

[0123] like Figure 6A and Figure 6B As shown, the model takes a P x N x F surrogate tensor and a P x N x F map tensor as input. For the surrogate tensor, P represents the polyline point of the target surrogate's trajectory in the last second (e.g.) (note that the polyline can represent road geometry and / or surrogate trajectory), and F represents the feature set (x) of each polyline point at time t (the timestamp of interest). t y t x t-1 y t-1 vx t vy t ang vel , acc, flash, t, mask), and N represents the number of polylines. For the map tensor, P represents the 20 polyline points of the lane associated with the target agent (e.g.), and F represents the feature set (x, y, dir) of each polyline point. x dir y Point type, prev x 、prev y (and mask), and N represents the number of polylines.

[0124] In the first stage of the encoder side of the encoder-decoder machine learning model, such as Figure 6B As shown, one or more PointNet layers of the proxy multi-segment encoder module (in) Figure 6B In the example, there are three) that convert the P x ​​N x F proxy tensor into a 1 x N x F' tensor. Similarly, one or more PointNet layers of the map polyline encoder module (in...) Figure 6B (In the example, there are three) convert the P x ​​N x F map tensor into a 1 x N x F' tensor.

[0125] It's important to note that PointNet is a neural network that directly consumes point clouds and provides a unified architecture for applications such as object classification, partial segmentation, and scene semantic parsing. The PointNet network learns a set of optimization functions / standards that select interesting or informative points in the point cloud and encode the reasons for their selection. The network's final fully connected layers aggregate these learned optimal values ​​into a global descriptor for the entire shape (shape classification) or use them for prediction of each point's label (shape segmentation).

[0126] In the second stage on the encoder side of the encoder-decoder machine learning model, a 1 x N x F' tensor is fed to one or more factorized self-attention modules. These modules encode information across polylines by using the target surrogate polyline as the query (denoted as "Q") and all polylines as keys (denoted as "K") and values ​​(denoted as "V"). The output of the one or more factorized self-attention modules is a 1 x 1 x F'' tensor.

[0127] Self-attention is a fundamental concept in Natural Language Processing (NLP) and deep learning, particularly prominent in transformer-based models. It enables models to weigh the importance of different parts of an input sequence when making predictions or capturing dependencies between words. Its role is to impart contextual intelligence, allowing the model to discern the importance of individual elements within a sequence and dynamically adjust their impact on the final output. This arrangement is especially important when the meaning of input elements is based on the meaning of other input elements (e.g., in language processing tasks).

[0128] In self-attention, a query (Q) is an element that seeks information (e.g., a target surrogate polyline). For each element in the input sequence, a query vector is computed. These queries represent elements within the sequence that should receive attention. A key (K) helps identify and locate important elements in the sequence. Similar to the queries, a key vector is computed for each element of interest. Values ​​(V) carry information. Again, for each element, a value vector is computed. These vectors maintain what needs to be considered when determining the importance of elements within the sequence.

[0129] For each element in the input sequence (e.g., predicting a polyline), a query vector, a key vector, and a value vector are computed. These vectors form the basis of the attention mechanism's operations. Then, an attention score is computed for each pair of elements in the sequence. The attention score between the query and the key quantifies their compatibility or relevance. Finally, the attention scores are used as weights to perform a weighted aggregation of the value vectors. This aggregation results in a self-attention output, representing an enhanced and context-informed representation of the input sequence (here, a 1 x 1 F'' tensor).

[0130] On the decoder side, such as Figure 6CAs illustrated, a shared fully connected multilayer perceptron (MLP) layer is followed by a separate fully connected MLP layer for each classification output, which predicts the turning and flow classification labels for the agent of interest. The encoder-decoder machine learning model can be applied to multiple agents of interest (e.g., all vehicles sufficiently close to the autonomous vehicle at an intersection to influence its driving decisions) to obtain a predicted classification label for each agent of interest. A rule-based trajectory generator can then use the predicted classification labels to generate candidate trajectories for consumption by the driving planning module (see example...). Figure 4 ).

[0131] Figure 7 This is an example of the various aspects of this disclosure. Figures 6A to 6C The example encoder-decoder machine learning model architecture illustrated in Figure 700 is an example of an enhanced diagram. Figure 7 As shown, on the decoder side of the disclosed encoder-decoder machine learning model architecture, an MLP layer is shared between the classification head and the regression head. The classification head includes a series of MLP layers that provide turning classification output and an optional streaming classification head that can also provide streaming classification output. Along with the classification head, there are four regression heads, each of which includes a skipped MLP followed by another MLP.

[0132] Each regression head corresponds to one of the four turning categories (i.e., left turn, right turn, straight, and U-turn) and outputs the trajectory for that turn. The classification head provides the probability of each turning trajectory, and the regression head predicts the future trajectory. The predicted trajectory can be output as a series of xy coordinates. In some cases, the trajectory can be predicted up to three seconds into the future.

[0133] Therefore, although Figures 6A to 6C The encoder-decoder machine learning model architecture illustrated in the example simply outputs the probability that the target agent follows a specific turn classification on the decoder side, but... Figure 7 The decoder-side output illustrated in the example also includes predicted trajectories for turn classification. Having a separate regression head facilitates training a dedicated trajectory generator for each turn intent. Compared to rule-based methods, this provides more realistic trajectories and is better able to model complex scenarios such as stop-and-go traffic, cuts, merges, splits, and interactions with other agents. Figure 7 The decoder side illustrated in the example is also better able to handle input noise from object fusion.

[0134] Figure 8This is illustration 800 illustrating an example intersection scenario according to various aspects of this disclosure. In illustration 800, a self-driving vehicle (whose trajectory is represented by a series of squares) follows two agents of interest (whose current trajectories are represented by a series of small circles). A series of future predicted trajectory points (represented by a series of larger circles) are determined based on one or more P x ​​N x F agent tensors (e.g., determined at least in part based on sensor information such as from camera 312, radar 314, and / or lidar sensor 316) and one or more P x ​​N x F map tensors (e.g., from map 302) applied to the disclosed encoder-decoder machine learning model to the agents of interest.

[0135] Figure 9 An example method 900 for predicting driving trajectories according to various aspects of this disclosure is illustrated. In one aspect, method 900 can be performed by a self-driving vehicle (e.g., any of the vehicles described herein).

[0136] At point 910, the self-driving vehicle will use machine learning models (e.g., ... Figures 6A to 6C (As illustrated) Applied to one or more agent tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both. In one aspect, operation 910 may be performed by one or more WWAN transceivers 330, one or more short-range wireless transceivers 340, one or more processors 306, memory 304, and / or driving strategy component 318, any or all of which may be considered as components for performing the operation.

[0137] At 920, the autonomous vehicle performs driving maneuvers based on a prediction of driving intentions. In one aspect, operation 920 may be performed by one or more WWAN transceivers 330, one or more short-range wireless transceivers 340, one or more processors 306, memory 304, and / or driving strategy components 318, any one or all of which may be considered as components for performing the operation.

[0138] It should be understood that the technical advantage of Method 900 is that it improves trajectory prediction for the target agent / vehicle, and thus improves autonomous driving performance. More specifically, Method 900 can be used to simultaneously predict the agent's future lateral (turning) and longitudinal (traffic flow) intentions at intersections. It learns to do this using a common map and agent feature representations. It also learns from a joint backbone, thus helping the task learn common representations better and faster. Method 900 has low inference latency, allowing for the generation of well-reasonable future trajectories with high accuracy by employing the learned intention predictions, thereby improving autonomous driving performance.

[0139] As can be seen in the detailed description above, different features are grouped together in the examples. This manner of disclosure should not be construed as an intention to have more features than those explicitly mentioned in each clause. Rather, the various aspects of this disclosure may include fewer features than those in the individual example clauses disclosed. Therefore, the following clauses should be regarded accordingly as incorporated into the description, where each clause may serve as a separate example. Although each dependent clause may refer in the clause to a specific combination with one of the other clauses, the aspect of that dependent clause is not limited to that specific combination. It should be understood that other example clauses may also include combinations of aspects of a dependent clause with the subject matter of any other dependent or independent clause, or combinations of any feature with other dependent and independent clauses. The various aspects disclosed herein explicitly include these combinations unless explicitly stated or readily inferred that a particular combination is not intended for use (e.g., contradictory aspects, such as defining an element as both an electrical insulator and an electrical conductor). Furthermore, it is contemplated that aspects of a clause may be included in any other independent clause, even if that clause does not directly depend on the independent clause.

[0140] Specific implementation examples are described in the following numbered clauses:

[0141] Clause 1. A method for predicting a driving trajectory performed by a self-driving vehicle, the method comprising: applying a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and performing driving maneuvers based on the predicted driving intention.

[0142] Clause 2. The method according to Clause 1, wherein the one or more surrogate tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the trajectories of the target vehicle and one or more adjacent vehicles of the target vehicle in a recent time period, a plurality of points of each of the plurality of polylines, and a plurality of features of each of the plurality of polylines.

[0143] Clause 3. The method according to Clause 2, wherein the plurality of features includes: the x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the previous x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the previous y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the x-axis velocity value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, and the target vehicle and The y-axis velocity value of the one or more adjacent vehicles at each of the plurality of points, the angular velocity value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the acceleration value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the flash status of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the time offset between the current timestamp of the ego vehicle and the timestamps of each of the plurality of points, the mask value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, or any combination thereof.

[0144] Clause 4. The method described in Clause 3, wherein the length of the most recent time period is one second.

[0145] Clause 5. The method according to any one of Clauses 1 to 4, wherein the one or more map tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the target vehicle and one or more adjacent vehicles traveling along their lanes, lane boundaries or both; a plurality of points of each of the plurality of polylines; and a plurality of features of each of the plurality of polylines.

[0146] Clause 6. The method according to Clause 5, wherein said plurality of features include: the x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the x-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the point type of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the mask value of the target vehicle at each of the plurality of points; or any combination thereof.

[0147] Clause 7. The method according to any one of Clauses 5 to 6, wherein the number of said plurality of points is 20 points.

[0148] Clause 8. The method according to any one of Clauses 5 to 7, wherein the point type along which the target vehicle and the one or more adjacent vehicles travel, the lane, the lane boundary, or both, includes: lane center point type, lane boundary point type, or a combination thereof.

[0149] Clause 9. The method according to any one of Clauses 1 to 8, wherein the machine learning model is an encoder-decoder machine learning model.

[0150] Clause 10. The method according to Clause 9, wherein the encoder side of the encoder-decoder machine learning model includes a first stage and a second stage.

[0151] Clause 11. The method according to Clause 10, wherein the first stage comprises: a proxy polyline module applied to the one or more proxy tensors, wherein the proxy polyline module converts the one or more proxy tensors into one or more two-dimensional proxy tensors; and a map polyline module applied to the one or more map tensors, wherein the map polyline module converts the one or more map tensors into one or more two-dimensional map tensors.

[0152] Clause 12. The method according to Clause 11, wherein the first stage further comprises: a self-attention module applied to the one or more two-dimensional agent tensors and the one or more two-dimensional map tensors to obtain a one-dimensional vector representing the predicted driving intention of the target vehicle.

[0153] Clause 13. The method according to any one of Clauses 10 to 12, wherein the second phase comprises: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for the turn classification and the flow classification.

[0154] Clause 14. The method according to any one of Clauses 10 to 13, wherein: the predicted driving intention further includes a turning trajectory, and the second phase includes: a shared fully connected multilayer perception machine (MLP) layer, and one or more separate fully connected MLP layers for each of the plurality of turning trajectories.

[0155] Clause 15. The method according to any one of Clauses 1 to 14, wherein the turn classification represents the probability that the target vehicle will perform one of a plurality of turn categories.

[0156] Clause 16. The method described in Clause 15, wherein the plurality of turning categories includes: left turn, right turn, straight and U-turn.

[0157] Clause 17. The method according to any one of Clauses 1 to 16, wherein the flow classification represents the probability that the target vehicle will perform one of a plurality of flow categories.

[0158] Clause 18. The method described in Clause 17, wherein the plurality of flow categories includes: free flow, start, deceleration, and stop.

[0159] Clause 19. The method according to any one of Clauses 1 to 18, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

[0160] Clause 20. The method according to any one of Clauses 1 to 19, wherein the driving maneuver includes: lane change before or after the road intersection, left turn at the road intersection, right turn at the road intersection, U-turn at the road intersection, driving straight through the road intersection, merging into the lane of the target vehicle, or a hard braking event.

[0161] Clause 21. A self-driving vehicle comprising: one or more memories; one or more transceivers; and one or more processors communicatively coupled to the one or more memories and the one or more transceivers, the one or more processors being individually or in combination configured to: apply a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and perform driving maneuvers based on the predicted driving intention.

[0162] Clause 22. The self-driving vehicle as described in Clause 21, wherein the one or more proxy tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the trajectories of the target vehicle and one or more adjacent vehicles of the target vehicle in a recent time period, a plurality of points of each of the plurality of polylines, and a plurality of features of each of the plurality of polylines.

[0163] Clause 23. The self-propelled vehicle as described in Clause 22, wherein the plurality of features includes: the x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the x-axis velocity value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the target vehicle... The y-axis velocity values ​​of the vehicle and the one or more adjacent vehicles at each of the plurality of points, the angular velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the acceleration values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the flash status of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the time offset between the current timestamp of the self vehicle and the timestamps of each of the plurality of points, the mask values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, or any combination thereof.

[0164] Clause 24. Self-driving vehicle as described in Clause 23, wherein the length of the most recent time period is one second.

[0165] Clause 25. A self-contained vehicle according to any one of Clauses 21 to 24, wherein the one or more map tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the target vehicle and one or more adjacent vehicles traveling along their lanes, lane boundaries or both; a plurality of points of each of the plurality of polylines; and a plurality of features of each of the plurality of polylines.

[0166] Clause 26. The self-driving vehicle as described in Clause 25, wherein said plurality of features include: the x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the x-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the point type of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the mask value of the target vehicle at each of the plurality of points; or any combination thereof.

[0167] Clause 27. A self-driving vehicle pursuant to any one of Clauses 25 to 26, wherein the number of said plurality of points is 20 points.

[0168] Clause 28. The self-driving vehicle according to any one of Clauses 25 to 27, wherein the point type along which the target vehicle and the one or more adjacent vehicles travel, the lane, lane boundary, or both, includes: center point type, boundary point type, or a combination thereof.

[0169] Clause 29. A self-driving vehicle pursuant to any one of Clauses 21 to 28, wherein the machine learning model is an encoder-decoder machine learning model.

[0170] Clause 30. The self-driving vehicle as described in Clause 29, wherein the encoder side of the encoder-decoder machine learning model includes a first stage and a second stage.

[0171] Clause 31. The self-driving vehicle as described in Clause 30, wherein the first stage comprises: a proxy polyline module applied to the one or more proxy tensors, wherein the proxy polyline module converts the one or more proxy tensors into one or more two-dimensional proxy tensors; and a map polyline module applied to the one or more map tensors, wherein the map polyline module converts the one or more map tensors into one or more two-dimensional map tensors.

[0172] Clause 32. The self-driving vehicle as described in Clause 31, wherein the first stage further comprises: a self-attention module applied to the one or more two-dimensional agent tensors and the one or more two-dimensional map tensors to obtain a one-dimensional vector representing the predicted driving intention of the target vehicle.

[0173] Clause 33. A self-contained vehicle according to any one of Clauses 30 to 32, wherein the second phase comprises: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for the turn classification and the flow classification.

[0174] Clause 34. A self-driving vehicle according to any one of Clauses 30 to 33, wherein: the predicted driving intention further includes a turning trajectory, and the second phase includes: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for each of the plurality of turning trajectories.

[0175] Clause 35. A self-contained vehicle pursuant to any one of Clauses 21 to 34, wherein the turning classification represents the probability that the target vehicle will perform one of a plurality of turning categories.

[0176] Clause 36. Self-propelled vehicles as described in Clause 35, wherein the multiple turning categories include: left turn, right turn, straight ahead, and U-turn.

[0177] Clause 37. A self-contained vehicle according to any one of Clauses 21 to 36, wherein the flow classification represents the probability that the target vehicle will execute one of a plurality of flow categories.

[0178] Clause 38. Self-propelled vehicles as described in Clause 37, wherein the multiple flow categories include: free flow, starting, deceleration, and stopping.

[0179] Clause 39. A self-driving vehicle pursuant to any one of Clauses 21 to 38, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

[0180] Clause 40. A self-driving vehicle according to any one of Clauses 21 to 39, wherein the driving maneuver includes: lane change before or after the road intersection, left turn at the road intersection, right turn at the road intersection, U-turn at the road intersection, driving straight through the road intersection, merging into the lane in which the target vehicle is driving, or a hard braking event.

[0181] Clause 41. A self-driving vehicle, the self-driving vehicle comprising: components for applying a machine learning model to one or more proxy tensors and one or more map tensors associated with the target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and components for performing driving maneuvers based on the predicted driving intention.

[0182] Clause 42. The self-driving vehicle as described in Clause 41, wherein the one or more proxy tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the trajectories of the target vehicle and one or more adjacent vehicles of the target vehicle in a recent time period, a plurality of points of each of the plurality of polylines, and a plurality of features of each of the plurality of polylines.

[0183] Clause 43. The self-propelled vehicle as described in Clause 42, wherein the plurality of features includes: the x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the x-axis velocity value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points; the target vehicle... The y-axis velocity values ​​of the vehicle and the one or more adjacent vehicles at each of the plurality of points, the angular velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the acceleration values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the flash status of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the time offset between the current timestamp of the self vehicle and the timestamps of each of the plurality of points, the mask values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, or any combination thereof.

[0184] Clause 44. Self-propelled vehicle as described in Clause 43, wherein the length of the most recent time period is one second.

[0185] Clause 45. A self-contained vehicle according to any one of Clauses 41 to 44, wherein the one or more map tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the lanes, lane boundaries or both along which the target vehicle and one or more adjacent vehicles travel; a plurality of points of each of the plurality of polylines; and a plurality of features of each of the plurality of polylines.

[0186] Clause 46. The self-driving vehicle as described in Clause 45, wherein said plurality of features include: the x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the x-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the point type of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the mask value of the target vehicle at each of the plurality of points; or any combination thereof.

[0187] Clause 47. A self-driving vehicle pursuant to any one of Clauses 45 to 46, wherein the number of said plurality of points is 20 points.

[0188] Clause 48. The self-driving vehicle according to any one of Clauses 45 to 47, wherein the point type along which the target vehicle and the one or more adjacent vehicles travel, the lane, lane boundary, or both, includes: components for center point type, components for boundary point type, or combinations thereof.

[0189] Clause 49. A self-driving vehicle pursuant to any one of Clauses 41 to 48, wherein the machine learning model is an encoder-decoder machine learning model.

[0190] Clause 50. The self-driving vehicle as described in Clause 49, wherein the encoder side of the encoder-decoder machine learning model includes a first stage and a second stage.

[0191] Clause 51. The self-driving vehicle as described in Clause 50, wherein the first stage comprises: a proxy polyline module applied to the one or more proxy tensors, wherein the proxy polyline module converts the one or more proxy tensors into one or more two-dimensional proxy tensors; and a map polyline module applied to the one or more map tensors, wherein the map polyline module converts the one or more map tensors into one or more two-dimensional map tensors.

[0192] Clause 52. The self-driving vehicle as described in Clause 51, wherein the first stage further comprises: a self-attention module applied to the one or more two-dimensional agent tensors and the one or more two-dimensional map tensors to obtain a one-dimensional vector representing the predicted driving intention of the target vehicle.

[0193] Clause 53. A self-contained vehicle according to any one of Clauses 50 to 52, wherein the second phase comprises: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for the turn classification and the flow classification.

[0194] Clause 54. The self-driving vehicle according to any one of Clauses 50 to 53, wherein: the predicted driving intention further includes a turning trajectory, and the second phase includes: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for each of the plurality of turning trajectories.

[0195] Clause 55. A self-contained vehicle pursuant to any one of Clauses 41 to 54, wherein the turning classification represents the probability that the target vehicle will perform one of a plurality of turning categories.

[0196] Clause 56. Self-propelled vehicles as described in Clause 55, wherein the multiple turning categories include: left turn, right turn, straight ahead, and U-turn.

[0197] Clause 57. A self-contained vehicle pursuant to any one of Clauses 41 to 56, wherein the flow classification represents the probability that the target vehicle will execute one of a plurality of flow categories.

[0198] Clause 58. Self-propelled vehicles as described in Clause 57, wherein the multiple flow categories include: free flow, starting, deceleration, and stopping.

[0199] Clause 59. A self-driving vehicle pursuant to any one of Clauses 41 to 58, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

[0200] Clause 60. A self-driving vehicle pursuant to any one of Clauses 41 to 59, wherein the driving maneuver includes: lane change before or after the road intersection, left turn at the road intersection, right turn at the road intersection, U-turn at the road intersection, driving straight through the road intersection, merging into the lane in which the target vehicle is driving, or a hard braking event.

[0201] Clause 61. A non-transitory computer-readable medium storing computer-executable instructions that, when executed by a self-driving vehicle, cause the self-driving vehicle to: apply a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and perform driving maneuvers based on the predicted driving intention.

[0202] Clause 62. The non-transitory computer-readable medium according to Clause 61, wherein the one or more surrogate tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the trajectories of the target vehicle and one or more adjacent vehicles of the target vehicle in a recent time period, a plurality of points of each of the plurality of polylines, and a plurality of features of each of the plurality of polylines.

[0203] Clause 63. The non-transitory computer-readable medium according to Clause 62, wherein said plurality of features include: the x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the previous x-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the previous y-coordinate of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the x-axis velocity value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the target The y-axis velocity values ​​of the vehicle and the one or more adjacent vehicles at each of the plurality of points, the angular velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the acceleration values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the flash status of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, the time offset between the current timestamp of the self vehicle and the timestamps of each of the plurality of points, the mask values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, or any combination thereof.

[0204] Clause 64. The non-transitory computer-readable medium as described in Clause 63, wherein the length of the most recent time period is one second.

[0205] Clause 65. A non-transitory computer-readable medium according to any one of Clauses 61 to 64, wherein the one or more map tensors are three-dimensional tensors, the three-dimensional tensors representing: a plurality of polylines representing the lanes, lane boundaries or both along which the target vehicle and one or more adjacent vehicles travel; a plurality of points of each of the plurality of polylines; and a plurality of features of each of the plurality of polylines.

[0206] Clause 66. The non-transitory computer-readable medium as described in Clause 65, wherein said plurality of features include: the x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the x-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the y-direction of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous x-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the previous y-coordinate of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the point type of the target vehicle and the one or more adjacent vehicles along the lane, lane boundary, or both they are traveling in; the mask value of the target vehicle at each of the plurality of points; or any combination thereof.

[0207] Clause 67. A non-transitory computer-readable medium pursuant to any one of Clauses 65 to 66, wherein the number of said plurality of points is 20 points.

[0208] Clause 68. A non-transitory computer-readable medium according to any one of Clauses 65 to 67, wherein the point type along which the target vehicle and the one or more adjacent vehicles travel, the lane, the lane boundary, or both, comprises: center point type, boundary point type, or a combination thereof.

[0209] Clause 69. A non-transitory computer-readable medium pursuant to any one of Clauses 61 to 68, wherein the machine learning model is an encoder-decoder machine learning model.

[0210] Clause 70. The non-transitory computer-readable medium as described in Clause 69, wherein the encoder side of the encoder-decoder machine learning model includes a first stage and a second stage.

[0211] Clause 71. The non-transitory computer-readable medium according to Clause 70, wherein the first stage comprises: a proxy polyline module applied to the one or more proxy tensors, wherein the proxy polyline module converts the one or more proxy tensors into one or more two-dimensional proxy tensors; and a map polyline module applied to the one or more map tensors, wherein the map polyline module converts the one or more map tensors into one or more two-dimensional map tensors.

[0212] Clause 72. The non-transitory computer-readable medium according to Clause 71, wherein the first stage further comprises: a self-attention module applied to the one or more two-dimensional agent tensors and the one or more two-dimensional map tensors to obtain a one-dimensional vector representing the predicted driving intention of the target vehicle.

[0213] Clause 73. A non-transitory computer-readable medium according to any one of Clauses 70 to 72, wherein the second phase comprises: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for the turn classification and the flow classification.

[0214] Clause 74. A non-transitory computer-readable medium according to any one of Clauses 70 to 73, wherein: the predicted driving intention further includes a turning trajectory, and the second phase includes: a shared fully connected multilayer perceptron (MLP) layer, and one or more separate fully connected MLP layers for each of the plurality of turning trajectories.

[0215] Clause 75. A non-transitory computer-readable medium according to any one of Clauses 61 to 74, wherein the turn classification represents the probability that the target vehicle will perform one of a plurality of turn categories.

[0216] Clause 76. The non-transitory computer-readable medium as described in Clause 75, wherein the plurality of turning categories includes: left turn, right turn, straight ahead, and U-turn.

[0217] Clause 77. A nontransitory computer-readable medium pursuant to any one of Clauses 61 to 76, wherein the flow classification represents the probability that the target vehicle will perform one of a plurality of flow categories.

[0218] Clause 78. The non-transitory computer-readable medium as described in Clause 77, wherein the plurality of stream categories includes: free stream, start-up, deceleration, and stop.

[0219] Clause 79. A non-transitory computer-readable medium pursuant to any one of Clauses 61 to 78, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

[0220] Clause 80. A non-transitory computer-readable medium pursuant to any one of Clauses 61 to 79, wherein the driving maneuver includes: lane change before or after the road intersection, left turn at the road intersection, right turn at the road intersection, U-turn at the road intersection, driving straight through the road intersection, merging into the lane of the target vehicle, or a hard braking event.

[0221] Those skilled in the art will understand that information and signals can be represented using any of a variety of different techniques and skills. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, light fields or optical particles, or any combination thereof.

[0222] Furthermore, those skilled in the art will understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in connection with the aspects disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure.

[0223] The various exemplary logic blocks, modules, and circuits described in conjunction with the aspects disclosed herein may be implemented or performed using a general-purpose processor, a digital signal processor (DSP), an ASIC, a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor may be a microprocessor, but in alternative embodiments, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration.

[0224] The methods, sequences, and / or algorithms described in conjunction with the aspects disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium known in the art. Example storage media are coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal (e.g., a UE). Alternatively, the processor and storage medium may reside as discrete components in the user terminal.

[0225] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on or transmitted via a computer-readable medium. A computer-readable medium includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. A storage medium may be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and is accessible to a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included within the definition of a medium. As used herein, disks and optical discs include: compact optical discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs. Disks typically reproduce data magnetically, while optical discs reproduce data optically using lasers. Combinations of these should also be included within the scope of computer-readable media.

[0226] While the foregoing disclosure illustrates exemplary aspects of this disclosure, it should be noted that various changes and modifications may be made herein without departing from the scope of this disclosure as defined by the appended claims. For example, the functions, steps, and / or actions of the method claims according to aspects of this disclosure described herein need not be performed in any particular order. Furthermore, no component, function, action, or instruction described or claimed herein should be construed as critical or essential unless explicitly stated otherwise. Additionally, as used herein, the terms “set,” “group,” etc., are intended to include one or more of the stated elements. Furthermore, as used herein, the terms “having,” “comprising,” “including,” etc., do not exclude the presence of one or more additional elements (e.g., an element “having” A may also have B). Furthermore, the phrase “based on” is intended to mean “at least partially based on” unless otherwise explicitly stated. Furthermore, as used herein, the term “or” is intended to be open-ended when used in a series and is interchangeable with “and / or” unless otherwise explicitly stated (e.g., if used in conjunction with “any” or “only one”), or these alternatives are mutually exclusive (e.g., “one or more” should not be interpreted as “one and more”). Additionally, although components, functions, actions, and instructions may be described or claimed in the singular, plural forms may also be considered unless explicitly stated as limited to the singular. Thus, as used herein, the articles “a,” “an,” “the,” and “described” are intended to include one or more of the stated elements. Additionally, as used herein, the terms “at least one” and “one or more” include “one” component, function, action, or instruction that performs or is capable of performing the described or claimed functionality, and also include “two or more” components, functions, actions, or instructions that perform or are capable of performing the described or claimed functionality in combination.

Claims

1. A self-driving vehicle, said self-driving vehicle comprising: One or more memory units; One or more transceivers; and One or more processors, communicatively coupled to one or more memories and one or more transceivers, wherein the one or more processors are configured individually or in combination to: A machine learning model is applied to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and The driving maneuvers are performed based on the predicted driving intentions.

2. The self-driving vehicle according to claim 1, wherein the one or more proxy tensors are three-dimensional tensors, the three-dimensional tensors representing: Multiple polylines representing the trajectories of the target vehicle and one or more adjacent vehicles within a recent time period. Multiple points of each of the plurality of polylines, and Each of the plurality of polylines has multiple features.

3. The self-driving vehicle according to claim 2, wherein the plurality of features includes: The x-coordinates of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The y-coordinates of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The previous x-coordinates of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The previous y-coordinates of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The x-axis velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The y-axis velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The angular velocity values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The acceleration values ​​of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The flashing status of the target vehicle and the one or more adjacent vehicles at each of the plurality of points. The time offset between the current timestamp of the self-driving vehicle and the timestamp of each of the plurality of points. The mask value of the target vehicle and the one or more adjacent vehicles at each of the plurality of points, or Any combination of them.

4. The self-driving vehicle according to claim 3, wherein the length of the most recent time period is one second.

5. The self-driving vehicle according to claim 1, wherein the one or more map tensors are three-dimensional tensors, and the three-dimensional tensors represent: This refers to multiple segments of lines along which the target vehicle and one or more adjacent vehicles travel, including lanes, lane boundaries, or both. Multiple points of each of the plurality of polylines, and Each of the plurality of polylines has multiple features.

6. The self-driving vehicle according to claim 5, wherein the plurality of features includes: The target vehicle and the one or more adjacent vehicles travel along the lane, lane boundary, or both of their x-coordinates. The target vehicle and the one or more adjacent vehicles travel along the y-coordinate of the lane, lane boundary, or both. The target vehicle and the one or more adjacent vehicles travel along the lane, lane boundary, or both in the x-direction. The target vehicle and the one or more adjacent vehicles travel along the y-direction of the lane, lane boundary, or both. The target vehicle and the one or more adjacent vehicles travel along the lane, lane boundary, or both of their previous x-coordinates. The target vehicle and the one or more adjacent vehicles travel along the lane, lane boundary, or both of their previous y-coordinates. The target vehicle and the one or more adjacent vehicles travel along the point type of the lane, lane boundary, or both. The mask value of the target vehicle at each of the plurality of points, or Any combination of them.

7. The self-driving vehicle according to claim 5, wherein the number of the plurality of points is 20.

8. The self-driving vehicle of claim 5, wherein the point type along which the target vehicle and the one or more adjacent vehicles travel, the lane, the lane boundary, or both, comprises: Center point type, Boundary point type, or Their combination.

9. The self-driving vehicle according to claim 1, wherein the machine learning model is an encoder-decoder machine learning model.

10. The self-driving vehicle of claim 9, wherein the encoder side of the encoder-decoder machine learning model includes a first stage and a second stage.

11. The self-driving vehicle of claim 10, wherein the first stage comprises: A proxy polyline module applied to the one or more proxy tensors, wherein the proxy polyline module converts the one or more proxy tensors into one or more two-dimensional proxy tensors, and A map polyline module applied to the one or more map tensors, wherein the map polyline module converts the one or more map tensors into one or more two-dimensional map tensors.

12. The self-driving vehicle of claim 11, wherein the first stage further comprises: A self-attention module is applied to the one or more two-dimensional agent tensors and the one or more two-dimensional map tensors to obtain a one-dimensional vector representing the predicted driving intention of the target vehicle.

13. The self-driving vehicle of claim 10, wherein the second stage comprises: A shared, fully connected multilayer perceptron (MLP) layer, and One or more separate fully connected MLP layers are used for the turn classification and the flow classification.

14. The self-driving vehicle according to claim 10, wherein: The predicted driving intention also includes the turning trajectory, and The second phase includes: A shared, fully connected multilayer perceptron (MLP) layer, and One or more separate fully connected MLP layers for each of the multiple turning trajectories.

15. The self-driving vehicle of claim 1, wherein the turn classification represents the probability that the target vehicle will perform one of a plurality of turn categories.

16. The self-driving vehicle of claim 15, wherein the plurality of turning categories includes: Turn left. Turn right. Go straight, and U-shaped turn.

17. The self-driving vehicle of claim 1, wherein the flow classification represents the probability that the target vehicle will execute one of a plurality of flow categories.

18. The self-driving vehicle of claim 17, wherein the plurality of flow categories include: Free flow, Starting point Slow down, and stop.

19. The self-driving vehicle of claim 1, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

20. The self-driving vehicle of claim 1, wherein the driving controls include: Lane changes before or after the road intersection, The left turn at the intersection of the roads The right turn at the road intersection The U-shaped turn at the road intersection Drive straight through the intersection of the roads. Merging into the lane in which the target vehicle is driving, or Hard braking event.

21. A method for predicting driving trajectories performed by a self-driving vehicle, the method comprising: A machine learning model is applied to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and The driving maneuvers are performed based on the predicted driving intentions.

22. The method of claim 21, wherein the one or more proxy tensors are three-dimensional tensors, the three-dimensional tensors representing: Multiple polylines representing the trajectories of the target vehicle and one or more adjacent vehicles within a recent time period. Multiple points of each of the plurality of polylines, and Each of the plurality of polylines has multiple features.

23. The method of claim 21, wherein the one or more map tensors are three-dimensional tensors, the three-dimensional tensors representing: This refers to multiple segments of lines along which the target vehicle and one or more adjacent vehicles travel, including lanes, lane boundaries, or both. Multiple points of each of the plurality of polylines, and Each of the plurality of polylines has multiple features.

24. The method of claim 21, wherein the machine learning model is an encoder-decoder machine learning model.

25. The method of claim 21, wherein the turn classification represents the probability that the target vehicle will perform one of a plurality of turn categories.

26. The method of claim 21, wherein the flow classification represents the probability that the target vehicle will execute one of a plurality of flow categories.

27. The method of claim 21, wherein the predicted driving intention includes a turning trajectory associated with the turning classification.

28. The method of claim 21, wherein the driving operation comprises: Lane changes before or after the road intersection, The left turn at the intersection of the roads The right turn at the road intersection The U-shaped turn at the road intersection Drive straight through the intersection of the roads. Merging into the lane in which the target vehicle is driving, or Hard braking event.

29. A self-driving vehicle, said self-driving vehicle comprising: A component for applying a machine learning model to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both. and Components for performing driving maneuvers based on the predicted driving intentions.

30. A non-transitory computer-readable medium storing computer-executable instructions, which, when executed by a self-driving vehicle, cause the self-driving vehicle to: A machine learning model is applied to one or more proxy tensors and one or more map tensors associated with a target vehicle to obtain a predicted driving intention of the target vehicle at a road intersection, wherein the predicted driving intention includes turn classification, flow classification, or both; and The driving maneuvers are performed based on the predicted driving intentions.