Coupling detection and prediction transformer

The combined detection and prediction method using a transformer model improves the accuracy and generalization of object detection and prediction in autonomous systems, addressing the limitations of conventional methods by integrating sensor and map coding for enhanced navigation.

JP2026509270APending Publication Date: 2026-03-17WAABI CANADA INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing autonomous systems face challenges in accurately and timely detecting the position of objects and predicting their path movement in real-world environments, with conventional methods separating detection and prediction as distinct tasks, leading to impaired accuracy and generalization capabilities.

Method used

A computer implementation method that combines detection and prediction using a transformer model to iteratively modify a spatiotemporal object structure with sensor and map coding, allowing simultaneous decoding of location and path information through a decoder model.

Benefits of technology

Enhances the accuracy and generalization of object detection and prediction in autonomous systems, enabling timely and precise navigation in dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509270000001_ABST
    Figure 2026509270000001_ABST
Patent Text Reader

Abstract

The coupled detection and prediction transformer includes the ability to perform operations that include obtaining the sensor data encoding of sensor data acquired from a sensor, initializing a spatiotemporal object structure, and obtaining the map encoding of the geographical region corresponding to the sensor data. The operations further include, by the transformer model, iteratively modifying the spatiotemporal object structure using the sensor data encoding and map encoding to generate a modified spatiotemporal object structure, and by the decoder model, decoding the spatiotemporal object structure to obtain location and path information for at least one object within the geographical region, and outputting the location and path information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application is a non - provisional application of U.S. Patent Application No. 63 / 450,634, filed on March 7, 2023, which is hereby incorporated by reference in its entirety, and claims the benefit thereof.

Background Art

[0002] Background An autonomous system is an automated means of transportation that does not require a human pilot or driver to move in and respond to a real - world environment. Rather, an autonomous system includes a virtual driver, which is the decision - making part of the autonomous system. Specifically, the virtual driver controls the operation of the autonomous system. The virtual driver is an artificial intelligence system that learns how to interact in the real world and then executes the interactions when in the real world.

[0003] Part of the interaction in the real world is to detect the position of objects and predict the path movement of objects within the environment. To navigate the real - world environment autonomously and more safely, the prediction not only has to be accurate and generalized across many scenarios, but also has to be made in a timely manner so that the autonomous system can react appropriately. That is, existing technologies may have impaired accuracy and generalization capabilities compared to the computational requirements.

[0004] Conventional systems separate detection and prediction as distinct sequential tasks. Objects are first detected using sensor data, and then the path of each detected object is predicted.

Summary of the Invention

[0005] Summary In general, in one aspect, one or more embodiments relate to a computer implementation method that includes obtaining sensor data coding for sensor data acquired from a sensor, initializing a spatiotemporal object structure, and obtaining map coding for a geographical area corresponding to the sensor data. The method further includes iteratively modifying the spatiotemporal object structure using the sensor data coding and map coding by a transformer model to generate a modified spatiotemporal object structure. The method further includes decoding the spatiotemporal object structure by a decoder model to obtain location and path information for at least one object within the geographical area. The method further includes outputting the location and path information.

[0006] In general, one or more embodiments relate to a system comprising at least one computer processor and memory for storing computer-readable program code for causing the computer processor to perform an operation. The operation includes obtaining sensor data coding of sensor data acquired from a sensor, initializing a spatiotemporal object structure, and obtaining map coding of a geographical area corresponding to the sensor data. The operation further includes, by a transformer model, iteratively modifying the spatiotemporal object structure using the sensor data coding and map coding to generate a modified spatiotemporal object structure. The operation further includes, by a decoder model, decoding the spatiotemporal object structure to obtain location and path information of at least one object within the geographical area. The operation further includes outputting the location and path information.

[0007] In general, in one embodiment, one or more embodiments relate to a non-temporary computer-readable medium containing computer-readable program code for causing a computer system to perform an operation. The operation includes obtaining sensor data coding of sensor data acquired from a sensor, initializing a spatiotemporal object structure, and obtaining map coding of a geographical area corresponding to the sensor data. The operation further includes, by a transformer model, iteratively modifying the spatiotemporal object structure using the sensor data coding and map coding to generate a modified spatiotemporal object structure. The operation further includes, by a decoder model, decoding the spatiotemporal object structure to obtain location and path information of at least one object within the geographical area. The operation further includes outputting the location and path information.

[0008] Other aspects of the present invention will become apparent from the following description and the appended claims. [Brief explanation of the drawing]

[0009] [Figure 1] This document describes an autonomous system having a virtual driver according to one or more embodiments. [Figure 2] This shows a virtual driver according to one or more embodiments. [Figure 3] An example of a spatiotemporal object structure according to one or more embodiments is shown. [Figure 4] This document illustrates exemplary embodiments of a transformer model in one or more embodiments. [Figure 5] A flowchart for coupling detection and prediction according to one or more embodiments is shown. [Figure 6] A flowchart is shown for iteratively modifying the spatiotemporal object structure according to one or more embodiments. [Figure 7] One or more embodiments of coupled detection and prediction using LiDAR are shown. [Figure 8A]This document shows a computing system according to one or more embodiments of the present invention. [Figure 8B] This document shows a computing system according to one or more embodiments of the present invention. [Modes for carrying out the invention]

[0010] Similar elements in various figures are indicated with the same reference number for consistency. Detailed explanation Generally, embodiments relate to the combined detection and prediction of objects within a geographical area from sensor data and map data. The geographical area includes objects and map elements. Each object and map element has a physical location within the geographical area that indicates the exact point or location where the object or map element is located. Map elements are stationary within the geographical area and may be reflected in the map of the geographical area. Objects are located within the geographical area and are designed to move within the geographical area. For example, an object may be a person, a cyclist, an autonomous system, a vehicle, a non-intelligent object (e.g., debris moving within the geographical area), or another object that may move.

[0011] Detection is the identification of the precise physical location of an object within a geographical area. Here, "precise" refers to being within an engineering threshold for the operation of the autonomous system and may include a bounding box. Prediction is the determination of a future trajectory that is unknown at the time of sensor data acquisition. Combined detection and prediction involve creating a single unified code that can be decoded simultaneously to perform detection and prediction. Sensor data coding for sensor data and map coding for map data are acquired. Sensor data coding is the coding of the current state of the geographical area acquired by one or more sensors of the autonomous system. Map coding is the coding of a map of the geographical area. One or more embodiments include a transformer model that iteratively modifies a spatiotemporal object structure using sensor data coding and map coding to create a modified spatiotemporal object structure with unified coding. A decoder model decodes the modified spatiotemporal object structure to acquire location and path information for at least one object within the geographical area.

[0012] Referring to the drawings, Figures 1 and 2 show illustrative diagrams of an autonomous system and a virtual driver. Referring to Figure 1, the autonomous system (116) is a mode of autonomous driving that can be used without requiring a human pilot or human driver to move and react to a real-world environment. The autonomous system (116) can be fully autonomous or semi-autonomous. As a mode of driving, the autonomous system (116) is housed in a housing configured to move through a real-world environment. Examples of autonomous systems include self-driving vehicles (e.g., self-driving trucks and cars), drones, airplanes, robots, and the like.

[0013] The autonomous system (116) includes a virtual driver (102), which is the decision-making part of the autonomous system (116). The virtual driver (102) is an artificial intelligence system that learns how to interact with the real world and interacts accordingly. The virtual driver (102) is software that runs on a processor and makes decisions that cause the autonomous system (116) to interact with the real world, including moving, signaling, and stopping, or maintaining its current state. Specifically, the virtual driver (102) is decision-making software that runs on hardware (not shown). The hardware may include a hardware processor, memory, or other storage devices, and one or more interfaces. The hardware processor is any hardware processing unit configured to process computer-readable program code and to perform actions described in the computer-readable program code.

[0014] The real-world environment is a part of the real world designed to be navigated by an autonomous system (116) once trained. Thus, the real-world environment, along with the agent, can include concrete and land, buildings, and other objects within a geographical area. The agent is another object in the real-world environment that can move through it. The agent may have an independent decision-making function. This independent decision-making function allows the agent to determine how to move through the environment and can be based on visual or tactile cues from the real-world environment. For example, the agent could include other autonomous and non-autonomous traffic systems (e.g., other vehicles, bicycles, robots), pedestrians, animals, etc.

[0015] In the real world, a geographical region is the actual area within the real world that surrounds an autonomous system. In other words, from the perspective of a virtual driver, the geographical region is the area in which the autonomous system moves.

[0016] The real-world environment changes as the autonomous system (116) moves through the real-world environment. For example, the geographical area may change, or agents may move locations, which may include the addition of new agents or the departure of existing agents.

[0017] To interact with the real-world environment, the autonomous system (116) includes various types of sensors (104), such as LiDAR sensors, among others, used to acquire measurements of the real-world environment, and a camera to capture images from the real-world environment. LiDAR (i.e., an acronym for “Light Detection and Ranging”) is a sensing technique that uses light in the form of pulsed lasers to measure the range (variable distance) to various objects in the environment. The autonomous system (116) may also include other types of sensors. Sensor (104) provides input to a virtual driver (102).

[0018] In addition to the sensor (104), the autonomous system (116) includes one or more actuators (108). The actuators are hardware and / or software configured to control one or more physical parts of the autonomous system based on control signals from a virtual driver (102). In one or more embodiments, the control signals specify an action for the autonomous system (e.g., turning on a turn signal, applying the brakes by a specified amount, pressing the accelerator by a specified amount, turning the steering wheel or tires by a specified amount). The actuators (108) are configured to perform the action. In one or more embodiments, the control signals may specify a new state for the autonomous system, and the actuators may be configured to perform the new state in order to bring the autonomous system to the new state. For example, the control signals may specify that the autonomous system should turn by a specific amount while accelerating at a predetermined acceleration, and the actuators determine the amount of wheel movement and acceleration to apply to the accelerator to achieve the specific amount of turning and acceleration.

[0019] Although not shown in FIG. 1, embodiments may be used in a virtual environment, such as when a virtual driver is being trained and / or tested in a simulated environment created by a simulator. In such scenarios, to the virtual driver, the simulated environment appears as the real-world environment.

[0020] FIG. 2 shows a virtual driver according to one or more embodiments. The virtual driver (102) is operably connected to a sensor (104). The sensor (104) can include various types of sensors as described above with reference to FIG. 1. The sensor may be a real sensor or a virtual sensor. A real sensor can be used, for example, on a physical autonomous system when the physical autonomous system is moving in the real world. A virtual sensor can be used to test or train the virtual driver (102). When the sensor is a virtual sensor, the interface between the virtual sensor and the virtual driver (102) matches the interface between a real sensor and the virtual driver in one or more embodiments.

[0021] The sensor (104) is configured to output sensor data (204) to the virtual driver (102). The sensor data (204) is data collected by sensors in the geographic area around the autonomous system.

[0022] One type of sensor (104) is a LiDAR sensor. The LiDAR sensor is configured to emit pulsed light waves into the surrounding environment. The pulsed light waves reflect off surrounding objects and return to the sensor. The sensor uses the time difference between the transmitted pulsed light wave and the pulsed light wave returning to the LiDAR sensor to calculate the distance the pulsed light wave has traveled. The LiDAR sensor can also record the direction and intensity of the returned light wave. The light wave is transmitted as part of a LiDAR sweep. A LiDAR sweep is a capture of the geographic area by the LiDAR sensor along a predetermined angular range around the LiDAR sensor.

[0023] A LiDAR sensor is configured to provide sensor data (204) in the form of a collection of LiDAR points. Each LiDAR point may have distance, direction, and intensity. The distance and direction may be relative to the LiDAR sensor. LiDAR points can be transformed into a three-dimensional (3D) space. For example, each point in the list may be transformed into a three-dimensional Cartesian coordinate system having a value that specifies the intensity. When transformed into 3D space, the LiDAR points are a LiDAR point cloud. That is, a LiDAR point cloud is a single scan of the surrounding region surrounding the LiDAR sensor. In other embodiments, a LiDAR point cloud may be a frame having multiple LiDAR sweeps of the sensing region, where each sweep in a single point cloud is associated with a specific point in time.

[0024] Another type of sensor is a camera. Different types of cameras may be used. For example, the camera may be a visible light camera, a full-spectrum camera, a thermal imaging camera, or another type of camera. In such scenarios, the sensor data includes a set of one or more images from the camera.

[0025] The sensor (104) is communicatively connected to a virtual driver (102). The virtual driver (102) includes a detection and prediction system (206). The detection and prediction system (206) is a machine learning framework configured to perform object binding detection and prediction of the future path of an object. In binding detection and prediction, the detection of the object's current location information may be performed concurrently with the detection of the object's path information.

[0026] Detection is the estimated position information of an object when sensor data is acquired. In other words, detection is the current position information. Object position information is information that identifies a specific location of an object. Specifically, position information is the area within the geographical region where the object is located. For example, position information may be the definition of a bounding box around the object, the object's center of gravity, the object's surface, or other position information around the object. Position information can be defined in relation to an autonomous system, a geographical region, or another reference frame. For example, position information may be the predicted position of the object's center of gravity as defined by a bird's-eye view (BEV) of the geographical region, as well as the predicted estimated height, width, and length of the object. Other mechanisms may be used to specify the position of an object within the geographical region. Position information may further include the object's orientation. Position information may be called the object's attitude at the defined point in time.

[0027] Path information defines the path an object takes through a future geographical area. "Future" refers to the time since the last capture of sensor data from which location and path information were acquired. Path information specifies the object's predicted future trajectory. Path information can be defined discretely as a sequence of location data in a set of future time steps. For example, path information may be a sequence of attitudes. Path information may also be defined continuously, such as a continuous trajectory defined by lines.

[0028] The sensor data encoder model (208) is a machine learning model configured to generate encodings of sensor data (204). In one or more embodiments, encoding is the extraction of higher-level features from the sensor data. Encoding may be vector embedding of the sensor data. In one or more embodiments, encoding is not an explicit encoding of identified objects. For example, encoding may encode features at the pixel or voxel level. The sensor data encoder model (208) may include a convolutional neural network, a graph neural network, a multilayer perceptron model, or other machine learning models. For example, in the case of LiDAR, the sensor data encoder model (208) may be configured to voxelize one or more LiDAR point clouds and generate encodings from the voxelized LiDAR point clouds.

[0029] The map encoder model (210) is a machine learning model configured to generate map encodings of a map of a geographical area. Specifically, the encoder model (210) is configured to encode map elements within a geographical area and the relationships between map elements. Map elements are physical parts of a geographical area that can be reflected in the map of the geographical area. For example, a map element may be a curb, a specific lane marker, a specific location between two lane markers, or a lane marker and a curb, lights, a stop sign, or a construction area, or any other physical object / location within the geographical area. Map elements may or may not be distinguishable in the real world. For example, if a map element is a specific real-world location between two lane markers, the specific location exists in the real world and has a physical location, but the specific location may not have any real-world signs or other signs at that location. In one or more embodiments, the map of the geographical area directly or indirectly specifies the stationary locations of map elements.

[0030] The map encoder model (210) is a software process configured to encode map elements of a geographical area as relative positions to each other. Specifically, the map encoder model (210) can be configured to calculate, for each map element, the relative position of the map element with respect to other map elements within the geographical area. Thus, for each map element, a set of relative positions of the map element with respect to other map elements can be defined. In one or more embodiments, the relative positions are defined on a pair-by-pair basis between pairs of map elements. The map encoder model (210) can be further configured to encode the relative positions into a feature set of the pair of map elements. Furthermore, a map element may have a feature set that defines the map element. For example, the feature set may encode size information, the type of geographical area in which the map element is located (e.g., urban, rural, etc.), the type of map element, and other features relating to the map element or properties surrounding the map element.

[0031] The output of the map encoder model (210) may be map element coding in a map graph. A map graph is a graph data structure having map element nodes connected by edges. Each map element node relates to an individual corresponding map element. An edge connecting two map element nodes may be associated with the relative position coding of the pair of corresponding map element nodes. A map element node may be associated with a feature set generated based on the general characteristics of the map element.

[0032] In one or more embodiments, a spatiotemporal object structure (e.g., initial spatiotemporal object structure (212), modified spatiotemporal object structure (214)) is a data structure that associates objects with the spatiotemporal features of corresponding objects. Specifically, each object has a corresponding spatiotemporal feature within the spatiotemporal object structure. The spatiotemporal feature may include a set of features for the object at each of several time steps, including the current time. The feature set is a high-level feature set having the object coding of the object at the time step. The feature set at the current time corresponds to object detection, and the feature set for future time steps corresponds to object prediction. The final feature set relates to the target location of the object. The target location is the location within the geographical area to which the object is about to move. For example, the target location may be 5 seconds in the future, but the predicted time step may be up to 2 seconds in the future. For example, the spatiotemporal object structure may be a three-dimensional vector structure in which the first dimension corresponds to the object and the second dimension corresponds to the time step. The value at the intersection of a specific object and a specific time step in a three-dimensional vector structure represents the feature set (i.e., the third dimension) of the specific object and the specific time step. An example of a spatiotemporal object structure is shown in Figure 3.

[0033] The initial spatiotemporal object structure (212) is any pre-modification spatiotemporal object structure. For example, the initial spatiotemporal object structure may have placeholders for corresponding sets of spatiotemporal features of a given number of objects. As a more specific example, a detection and prediction system (206) may be configured to detect up to 300 objects over 11 future time steps. In this example, the initial spatiotemporal object structure (212) may have separate placeholders for each combination of 300 objects and 12 time steps. The values ​​of the placeholders may generally be random, regardless of the sensor data or based on coarse detection. Coarse detection can optimize recall over precision. Coarse detection is the detection of object locations that have an error exceeding a threshold amount (e.g., more ghost objects are detected) or where the probability of being accurate is less than a threshold.

[0034] The modified spatiotemporal object structure (214) is the spatiotemporal object structure after modification by the transformer model (216). In each iteration of the transformer model (216), the spatiotemporal features are more accurate estimates of the spatial features at various time steps.

[0035] The transformer model (216) is a machine learning model configured to transform an initial spatiotemporal object structure (212) into a modified spatiotemporal object structure (214). Specifically, the transformer model (216) is configured to iteratively modify spatiotemporal features using sensor data coding and map coding. An example of the transformer model is shown in Figure 4.

[0036] Continuing to refer to Figure 1, the decoder model (218) is a machine learning model configured to perform detection and prediction from a modified spatiotemporal object structure. In one or more embodiments, the decoder model (218) includes a detection decoder (not shown) and a separate prediction decoder. The detection decoder may be an MLP that outputs the centroid, yaw, length, width, and confidence of existence of an object. The prediction decoder may be a transformer model, a recurrent neural network, etc. The detection decoder (218) may be configured to output a predicted object type. In one or more embodiments, the object type is a type of physical object. For example, the object type may be a vehicle, a person, an airplane, a ball, a person on a bicycle, a building, a truck, or other type of object. The decoder model (218) may have a set of defined classes, each class being for a specific object type. In one or more embodiments, the decoder model (218) may include a network layer configured to output the probability of each of at least a subset of the classes of each object.

[0037] Figure 3 shows an example of a spatiotemporal object structure (300) according to one or more embodiments. Specifically, the exemplary spatiotemporal object structure may be a three-dimensional vector structure stored in heap memory when performing coupling detection and prediction. As shown in Figure 3, the exemplary spatiotemporal object structure (300) includes a first dimension (302) of the object and a second position (304) of the time step. The time step is a future time step that coincides with the object's path; that is, the time step may not coincide with the iterations of the transformer model. More specifically, as the object moves through a geographical area over time, the continuous future time is discretized into a set of time steps. The granularity of the discretization may be based on the location or type of the autonomous system. For example, an autonomous system, such as a robot operating in or designed to operate in a sparsely populated area, may have time steps in the order of seconds. As another example, an autonomous system, such as a vehicle designed to operate in an urban area, may have time steps in the order of milliseconds.

[0038] For a specific object (e.g., Object 2), the spatiotemporal feature includes a first pair for the current time (e.g., Object 2, Time 0 (306)), where the current time is the time of sensor data capture. The spatiotemporal feature may further include multiple pairs for future time steps (e.g., Object 2, Time 1 (308) to Object 2, Time t (310)). The spatiotemporal feature may further include pairs of object and target positions (e.g., Object 2, Target (312)).

[0039] As shown in Figure 3, the exemplary spatiotemporal object structure (300) has pairs for each object and time step, including a feature set and an anchor. The feature set is the feature set described above with reference to Figure 3. The anchor is the predicted position of the object at that point in time. The anchor may also include the predicted orientation of the object. For example, in BEV, the anchor may be the x,y coordinates that predict the position of the object. The x,y coordinates may be defined with respect to the boundary or centroid of the geographical region. Other mechanisms for defining the position may be used without departing from the claims. Initially, the anchor may be a rough prediction of the object. As another example, the anchor may be randomized based on a grid or based on another initialization strategy. The feature set may generally be an initial set of learned features used across multiple geographical regions. In such a scenario, the feature set is not pre-trained. The transformer model is configured to modify the anchor and feature set over one or more iterations.

[0040] Figure 4 shows an exemplary embodiment of a transformer model (400) according to one or more embodiments. The transformer model includes an object-map attention layer (402), one or more sensor type attention layers (e.g., a sensor type X attention layer (404), a sensor type Y attention layer (406)), an object self-attention layer (408), an object time attention layer (410), and an anchor update layer (412). Each of these layers is described below.

[0041] The term "attention layer" corresponds to the standard definition used in the field of machine learning. An attention layer is a type of neural network layer used to selectively focus on different parts of data when processing data. Generally, attention works by measuring the cosine similarity between two vectors or embeddings. Cosine similarity determines the importance of elements, thereby allowing the model to selectively focus on important inputs that lead to informed predictions and decisions. For example, with each attention layer, a transformer model can assign greater weights to different parts of a map, sensor data, or spatiotemporal object structure, and the weights are determined by the attention mechanism.

[0042] The object-map attention layer (402) is configured to perform cross-attention of feature sets to map element coding. In one or more embodiments, the object-map attention layer is configured to process spatiotemporal features of individual objects at individual time steps associated with the map, without considering the set of objects and the set of time steps. In other words, in relation to an exemplary spatiotemporal object structure, the object-map attention layer may be configured to process each pair individually. In one or more embodiments, the object-map attention layer (402) is a k-nearest neighbor attention layer.

[0043] One or more sensor type attention layers (e.g., sensor type X attention layer (404), sensor type Y attention layer (406)) may each be specific to a sensor type. In one or more embodiments, the sensor type attention layer is configured to process the spatiotemporal features of individual objects at individual time steps related to sensor data, without considering maps, sets of objects, and sets of time steps. In other words, in relation to an exemplary spatiotemporal object structure, the object-sensor data attention layer may be configured to process each pair individually. In one or more embodiments, the object-map attention layer (402) is a deformable attention layer.

[0044] The object self-attention layer (408) is configured to process the spatiotemporal features of individual objects in relation to other objects in the same time step, without considering a set of maps, sensor data, or time steps. In other words, in relation to an exemplary spatiotemporal object structure, the object self-attention layer may be configured to process pairs in relation to other pairs in the same time step. Thus, the object self-attention layer (408) processes predicted objects in relation to other predicted objects. In one or more embodiments, the object self-attention compares pairs only in the same time step.

[0045] The object time attention layer (410) is configured to process the spatiotemporal features of individual objects in relation to the same object at other time steps, without considering maps, sensor data, or other objects. In other words, in relation to an exemplary spatiotemporal object structure, the object time attention layer may be configured to process pairs in relation to other pairs of the same predicted objects. In one or more embodiments, object self-attention compares pairs only for the same object.

[0046] The anchor update layer (412) is configured to modify the anchor for each pair. The anchor update layer (412) may be configured to predict an offset from the current anchor position of the pair based on the pair's feature set. In one or more embodiments, the anchor update model is a small multilayer perceptron (MLP) model having a set of neural network layers.

[0047] Figures 1, 2, 3, and 4 show the configuration of the components, but other configurations may be used without departing from the spirit of the present invention. For example, various components can be combined to create a single component. As another example, a function performed by a single component may be performed by two or more components.

[0048] Figures 5 and 6 show flowcharts according to one or more embodiments. While the various steps in these flowcharts are presented and described sequentially, at least some steps may be performed in a different order, combined, or omitted, and at least some steps may be performed in parallel. Furthermore, the steps may be performed actively or passively.

[0049] Figure 5 shows a flowchart for coupling detection and prediction according to one or more embodiments.

[0050] In block 502, sensor data about a geographical area is acquired from the sensor. For example, one or more cameras can take one or more pictures of a geographical area. Furthermore, a LiDAR sensor can perform one or more LiDAR sweeps. To perform a LiDAR sweep, light pulses are transmitted by the LiDAR transmitter of the LiDAR sensor. The light pulses can be reflected by objects in the environment, even if unknown, and the time and intensity of the reflected light pulses returning to the receiver are determined. The time is used to determine the distance to the object. The result is a LiDAR point in the LiDAR data. In the case of a LiDAR sweep, multiple light pulses are transmitted at various angles, and multiple reflected lights are received to acquire multiple points in the LiDAR point cloud. The LiDAR point cloud may be acquired as a list of LiDAR points. The list may be converted back into a LiDAR point cloud.

[0051] A simulator that uses a sensor simulation model of a sensor, rather than using an actual sensor, can generate simulated sensor data. Specifically, the simulator can generate a scene and render the scene. The scene is located within a virtual geographical area. A machine learning model, which is part of the simulator, can determine the intensity or color of LiDAR points or pixels based on the location of the virtual sensor in the scene. The result is a set of simulated sensor data that mimics the actual sensor data of a particular simulated scene. The simulated sensor data can be at virtually any resolution. For example, the simulated sensor data may match the resolution of an actual sensor.

[0052] In block 504, the spatiotemporal object structure is initialized. Memory is allocated for the spatiotemporal object structure. Different techniques can be used to initialize the set of spatial features within the spatiotemporal object structure.

[0053] For example, in the first technique, the values ​​of the feature set may be defined randomly. In such a scenario, the values ​​can be initialized using random selection. In another example, a learned set of values ​​can be used as a spatial feature set, and the learned set of values ​​is common across several geographical regions. The learned set of values ​​can be learned by training a detection-prediction system through a training process.

[0054] Furthermore, corresponding anchors associated with the feature set are initialized. In one example, initialization can be performed by dividing the geographical area into a two-dimensional bird's-eye view grid or a three-dimensional grid and assuming that stationary objects reside within each grid cell. In another example, a coarse detector model can process sensor data and generate coarse detection of objects within the geographical area. A coarse detector is a lightweight detector that may return false positives. Therefore, even objects with a probability of existence below a given threshold can still be identified by a coarse detector. For example, in the case of a LiDAR point cloud, a course detection model can be based on locations with a higher density of LiDAR points.

[0055] In block 506, map coding of the geographical region corresponding to the sensor data is obtained. Map coding is any coding of the map of the geographical region. Map coding can be generated by applying a neural network model to the map elements within the geographical region. The following is an example of a technique for generating map coding. Other techniques may be used without departing from the scope of the present invention.

[0056] In this example, the physical location of a map element is obtained from the map data. In some embodiments, one or more map elements are determined from the map data. For example, if the map data includes lane markers, a map element could be a defined geographical point between two lane markers. From the physical location of a map element, its relative location to other map elements is calculated. Map element coding is generated using the relative location. Map element coding encodes the features of the map element (e.g., the type of map element, speed limits in the map element, and other features) and encodes the relative location. Map coding is generated using map element coding. Each map element corresponds to a map element node in the map coding. Map element nodes are connected by edges based on adjacency between them. The edge between two map element nodes is associated with the relative location between the two corresponding map elements in order to generate map coding. A graph neural network can be applied to the map layer to further update the map layer based on adjacent map elements. Thus, map element coding in a map reflects not only its own features in the map but also information about other map elements.

[0057] In block 508, the sensor coding model generates sensor data coding for the sensor data. In the case of a camera, the sensor coding model may include a convolutional neural network and other networks that transform a series of camera images into a projected bird's-eye view of the region. As another example, in the case of a LiDAR sensor, a set of LiDAR points is acquired. Each LiDAR point has a position and intensity in a Cartesian coordinate system. In some embodiments, the LiDAR points can be processed through a first set of neural network layers to generate a first set of LiDAR point features. For example, for each LiDAR point, the position values ​​in three coordinate systems may be concatenated with the intensity. The result is an input feature set to a first set of neural network layers that generate an output feature set for each LiDAR point. The output feature set may be point-level based. For example, each LiDAR point may be associated with a first set of point features and its original position (e.g., in a Cartesian coordinate system).

[0058] Each LiDAR point can be individually voxelized by a voxelization process. The geographical region is divided into a grid (i.e., a voxel grid), so that a one-to-one mapping can exist between voxels in the voxel grid and non-overlapping subregions of the geographical region. During the voxelization process, the location of a LiDAR point within the geographical region coincides with the subregion of the geographical region containing that location. The voxelization process associates the LiDAR point with a voxel mapped to the subregion. Furthermore, the 3D space offset from the centroid of the voxel is identified. For voxels in the voxel grid, each voxel can be associated with multiple sets of features from different LiDAR points. The sets of features may be aggregated. For example, aggregation may involve averaging. The LiDAR voxels are encoded to obtain encoded voxels. In one or more embodiments, the voxel grid is processed via a convolutional neural network (CNN). The voxel grid being processed can have three dimensions of location and a fourth dimension of the voxel feature set. The CNN updates the voxel features to take into account the features of surrounding voxels. Thus, encoding updates the features of each voxel. In one or more embodiments, the size of the current voxel grid does not change when encoding LiDAR voxels. Each encoded voxel is associated with a corresponding sub-region of a geographical area. Furthermore, this process may be repeated for a window of LiDAR point clouds, thereby updating the encoding in the voxel grid for each subsequent point cloud (e.g., by aggregating the respective voxel encodings). Thus, sensor data encoding can reflect not only current sensor data but also historical sensor data.

[0059] In block 510, the transformer model iteratively modifies the spatiotemporal object structure using sensor data coding and map coding to generate a modified spatiotemporal object structure. The transformer model updates the spatiotemporal object structure with features from map coding, features from sensor data coding, features from other predicted objects in the region, and features from the same object over time. Additionally, anchors may be updated. Through iterative updates, the spatiotemporal object structure becomes a more accurate representation of the geographic region, including the objects within the geographic region. The iterative modification of objects to a time-step feature map is shown in Figure 6.

[0060] Continuing with Figure 5, in block 512, the detection decoder decodes the spatiotemporal object structure to obtain location information for at least one object within the geographical area. Objects that do not exist can be identified based on a low confidence score generated by the detection decoder model. In one or more embodiments, the detection decoder may operate independently for each object's feature set and initial timestep pair of anchors. The detection decoder may be an MLP model that predicts the anchor's current position and its offset from the feature set. For example, the feature set may be concatenated with the anchor and then processed by a neural network of the MLP. The offset is then used to update the anchor and determine the object's current position. Similarly, the detection decoder can generate predictions of the object's dimensions. For example, the detection decoder may output the object's length and width.

[0061] In block 514, the predictive decoder decodes the spatiotemporal object structure to obtain path information for at least one object within the geographical region. The predictive decoder can operate in the same or similar manner as the detection decoder, but for each future time step. In such a scenario, the feature set can be concatenated with an anchor and then processed by the predictive decoder's MLP, independently for each time step and each object.

[0062] Another example of processing by the predictive decoder is the use of maps, both target and optional. For each time step of an object, the target location (in the spatiotemporal object structure) may be concatenated with the feature set and anchors to obtain the input set. The predictive decoder can then process the input set to output object locations (i.e., future waypoints for the time step) into the future. In one or more embodiments, the combination of future waypoints of the object forms the object's path information.

[0063] In block 516, position information and path information are output. The virtual driver of the autonomous system can use the agent's trajectory and the virtual driver's own path to its destination to determine the current trajectory of the autonomous system that satisfies safety criteria (e.g., collision avoidance, ensuring sufficient stopping distance) and other criteria (e.g., shortest path, reduction of lane changes) and facilitates movement to the destination. The virtual driver can then output control signals to one or more actuators. In a real-world environment, control signals are used by actuators to cause the autonomous system to perform actions such as moving in a specific direction at a specific speed or acceleration, waiting, displaying turn signals, or performing other actions. In a simulated environment, control signals are captured by a simulator that simulates the actuators and the resulting behavior of the autonomous system. The simulator simulates the autonomous system and thereby trains the virtual driver. That is, the virtual driver's behavior can be evaluated using the output that simulates the autonomous system in the simulated environment.

[0064] Figure 6 shows a flowchart for iteratively modifying the spatiotemporal object structure according to one or more embodiments.

[0065] In block 602, the object-map attention layer modifies the spatiotemporal object structure. The object-map attention layer can process each pair within the spatiotemporal object structure individually. For a pair, the following actions can be performed: In one or more embodiments, coordinates specified by an anchor are used to identify map elements adjacent to the location specified by the coordinates. For example, four adjacent map elements can be used. Orientation within the anchor may be used to calculate the relative distance between the agent and the map. Cross-attention can then be performed while the identified adjacent map elements are map-encoded into the pair's feature set. The result of the object-map attention layer is to modify the pair's feature set based on the map of the geographical region. This process can be repeated for each pair of spatiotemporal features. Other types of attention layers can be used to update the feature set based on the map-encoded elements.

[0066] In block 604, the sensor type attention layer modifies the spatiotemporal object structure. In one or more embodiments, each sensor type has a corresponding attention layer. Thus, block 604 may be executed separately for each sensor type. For example, a first sensor type attention layer can process LiDAR data, and a second sensor type attention layer can process camera sensor data. Sensor type attention layers can process pairs in series with other sensor type attention layers. For example, a feature set may be modified for a first sensor type, and then for a second sensor type. To execute the sensor type attention layer, a set of offsets is determined using the anchors of the pair. In one or more embodiments, a predetermined number of offsets are used. Cross-attention can then be performed between the sensor data encoding at each of the offset positions and the feature set of the pair. The result of the sensor type attention layer is to modify the feature set of the pair based on the sensor data of the sensor type.

[0067] In block 606, the object self-attention layer modifies the spatiotemporal object structure. The object self-attention layer considers other currently predicted objects within the region at the same time step. In Figure 3, the object self-attention layer performs attention along the object dimension. The object self-attention layer has multiple different predicted objects that are paying attention to each other. For example, the number of different predicted objects may be four. Any other number of predicted objects can be used without departing from the claims. Furthermore, the objects that are paying attention to each other may be based on the adjacency of the objects by their respective anchors. The orientation within the anchors may be used to calculate the relative distance between different objects. The result of the object self-attention layer is to modify the paired feature set based on other predicted objects within the region.

[0068] In block 608, the object time attention layer modifies the spatiotemporal object structure. The object time attention layer updates the object based on the object's past and future time steps. In Figure 3, the object time attention layer performs attention along the time step dimension. The object time attention layer has a feature set of a pair of single predicted objects that are of interest to each other. One of the pair has the object's target location. Therefore, modification of the pair's feature set also incorporates the object's predicted target location. The result of the object time attention layer is to modify the pair's feature set based on other predicted objects in the region.

[0069] In block 610, the anchor update layer modifies the spatiotemporal object structure. In one or more embodiments, the anchor update layer processes a set of features coupled with a pair of anchors to generate an offset. This processing may be performed through a set of neural network layers of the MLP. The anchors are then modified according to the offset. In other words, the anchors are moved from their current position to a current position with an offset to create a new current position. The new current position becomes the anchor for the next iteration or decoder. The process of the anchor update layer may be performed individually for each pair of spatiotemporal feature structures.

[0070] In block 612, a decision is made as to whether to perform another iteration. In one or more embodiments, the process is repeated until a set number of iterations have been performed. In such embodiments, the decision is in block 612 and may be whether or not the number of iterations has been performed. If the number of iterations has not been performed, the operation can repeat block 602 for the next set of iterations. If the number of iterations has been performed, the process continues to Figure 5. The process may also be stopped before the set number of iterations is completed, for example, to satisfy latency constraints.

[0071] The iterative process of a transformer model involves iteratively updating the spatiotemporal object structure. The spatiotemporal object structure contains placeholders for objects in one or more embodiments. Therefore, even if predicted object locations exist, not all locations necessarily correspond to actual objects. Before the first iteration, the predicted objects may be inaccurate or nonexistent. Furthermore, the object's motion may be unknown. Throughout the update process, the spatiotemporal object structure is iteratively updated to reflect a more accurate set of features and anchors for the objects.

[0072] The training of a transformer model can be performed by comparing a set of objects predicted by a decoder (i.e., a detection or prediction decoder) for a given time step with the actual locations of the objects in a given map and given sensor data. The results are used to calculate the loss that is backpropagated through the detection and prediction system.

[0073] Figure 7 shows an example of coupled detection and prediction using LiDAR according to one or more embodiments. The following examples are for illustrative purposes only and are not intended to limit the scope of the present invention.

[0074]

number

[0075] An exemplary architecture may include an input encoder (e.g., a LiDAR encoder (704), a map encoder (708)), a spatiotemporal object feature set and anchors in a spatiotemporal object structure (710), a transformer model (712), and a decoder model (714). The input encoder is configured to abstract heterogeneous sensory inputs, such as LiDAR and high-definition (HD) maps, into feature sets and spatial anchors. The spatiotemporal object feature set and anchor set encodes the latent attributes of the object, as well as its current and future orientation. The transformer model (712) includes an attention mechanism applied to update the object feature set and anchors with information from map encoding and sensor data encoding. The decoder model (714) generates detection and prediction results (716) from the updated object feature set (718).

[0076]

number

[0077]

number

[0078]

number

[0079]

number

[0080]

number

[0081]

number

[0082]

number

[0083]

number

[0084]

number

[0085]

number

[0086]

number

[0087]

number

[0088]

number

[0089]

number

[0090]

number

[0091]

number

[0092]

number

[0093]

number

[0094] Embodiments can be implemented on computing systems specifically designed to achieve improved technical results. When implemented on a computing system, the features and elements of the Disclosure provide significant technical advancements over computing systems that do not implement the features and elements of the Disclosure. Any combination of mobile, desktop, server, router, switch, embedded device, or other types of hardware can be improved by including the features and elements described in the Disclosure. For example, as shown in Figure 8A, a computing system (800) may include one or more computer processors (802), non-persistent storage (804), persistent storage (806), communication interfaces (808) (e.g., Bluetooth interface, infrared interface, network interface, optical interface, etc.), and a number of other elements and functions that implement the features and elements of the Disclosure. The computer processor (802) may be an integrated circuit for processing instructions. The computer processor may be one or more cores or microcores of a processor. The computer processor (802) includes one or more processors. One or more processors can include a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), or a combination thereof.

[0095] The input device (810) may include a touchscreen, keyboard, mouse, microphone, touchpad, electronic pen, or any other type of input device. The input device (810) may receive input from a user in response to data and messages presented by the output device (812). The input may include text input, audio input, video input, etc., which may be processed and transmitted by the computing system (800) according to this disclosure. The communication interface (808) may include an integrated circuit for connecting the computing system (800) to a network (not shown) (e.g., a local area network (LAN), a wide area network (WAN) such as the Internet, a mobile network, or any other type of network) and / or another device such as another computing device.

[0096] Furthermore, the output device (812) may include a display device, a printer, an external storage device, or any other output device. One or more of the output devices may be the same as or different from the input devices. The input and output devices may be connected locally or remotely to the computer processor (802). Many different types of computing systems exist, and the aforementioned input and output devices may take other forms. The output device (812) can display data and messages sent and received by the computing system (800). The data and messages may include text, audio, video, etc., and may include the data and messages described above in other figures of this disclosure.

[0097] Software instructions in the form of computer-readable program code for performing embodiments may be stored, in whole or in part, temporarily or permanently on a non-temporary computer-readable medium such as a CD, DVD, storage device, diskette, tape, flash memory, physical memory, or any other computer-readable storage medium. Specifically, software instructions may correspond to computer-readable program code configured to perform one or more embodiments that, when executed by a processor, may include transmitting, receiving, presenting, and displaying data and messages described in other figures of this disclosure.

[0098] The computing system (800) in Figure 8A may be connected to a network or may be part of a network. For example, as shown in Figure 8B, the network (820) may include multiple nodes (e.g., node X (822), node Y (824)). Each node may correspond to a computing system such as the computing system shown in Figure 8A, or a combined group of nodes may correspond to the computing system shown in Figure 8A. As an example, the embodiment may be implemented on a node of a distributed system connected to other nodes. As another example, the embodiment may be implemented on a distributed computing system having multiple nodes, and each part may be located on a different node within the distributed computing system. Furthermore, one or more elements of the aforementioned computing system (800) may be located remotely and connected to other elements via a network.

[0099] Nodes in the network (820) (e.g., node X (822), node Y (824)) can be configured to provide services to a client device (826), including receiving requests and sending responses to the client device (826). For example, the nodes may be part of a cloud computing system. The client device (826) may be a computing system, such as the computing system shown in Figure 8A. Furthermore, the client device (826) may include and / or run one or more embodiments, in whole or in part.

[0100] The computing system in Figure 8A may include the ability to present raw and / or processed data, such as the results of comparisons and other processing. For example, presenting data can be achieved by various presentation methods. Specifically, data may be presented by being displayed in a user interface, transmitted to different computing systems, and stored. The user interface may include a graphical user interface (GUI) that displays information on a display device. The GUI may include various GUI widgets that organize what data is displayed and how the data is presented to the user. Furthermore, the GUI may present data directly to the user, for example, data presented as actual data values ​​via text, or data rendered into a visual representation of the data by a computing device, such as by visualizing a data model.

[0101] As used herein, the term "connected" has multiple meanings. A connection can be direct or indirect (e.g., through another component or network). A connection can be wired or wireless. A connection can be a temporary, permanent, or semi-permanent communication channel between two entities.

[0102] The various descriptions in the drawings may be combined and may include, or be contained within, features described in other drawings of this application. The various elements, systems, components, and processes shown in the drawings may be omitted, repeated, combined, and / or modified as shown in the drawings. Therefore, the scope of this disclosure should not be considered to be limited to any particular configuration shown in the drawings.

[0103] In this application, ordinal numbers (e.g., 1st, 2nd, 3rd, etc.) may be used as adjectives for elements (i.e., any noun in this application). The use of ordinal numbers does not imply or create a particular order of elements, nor does it limit any element to only a single element, unless expressly disclosed by the use of “before,” “after,” “single,” and other such terms. Rather, the use of ordinal numbers is for distinguishing elements. For example, a first element is different from a second element, and a first element may encompass two or more elements and may follow (or precede) a second element in the order of elements.

[0104] Furthermore, unless otherwise specified, "or" is "inclusive or" and therefore includes "and". Additionally, unless otherwise specified, items joined by "or" may include any combination of items with any number of each item.

[0105] The above description includes numerous specific details to provide a more complete understanding of the disclosure. However, it will be apparent to those skilled in the art that the art can be carried out without these specific details. In other examples, well-known features are not described in detail to avoid unnecessarily complicating the description. Furthermore, other embodiments not expressly described above can be devised that do not deviate from the claims disclosed herein. Thus, the scope should be limited only by the appended claims.

Claims

1. A method performed by a computer, To obtain the sensor data encoding of sensor data acquired from a sensor, Initializing the spatiotemporal object structure, To obtain the map encoding of the geographical area corresponding to the aforementioned sensor data, The transformer model iteratively modifies the spatiotemporal object structure using the sensor data coding and the map coding to generate the modified spatiotemporal object structure, The decoder model decodes the spatiotemporal object structure to obtain location and path information for at least one object within the geographical area, A method performed by a computer, which includes outputting the location information and the route information.

2. Further includes controlling the autonomous system equipped with the sensors according to the location information and route information, The method performed by a computer as described in claim 1.

3. The object-map attention layer modifies the spatiotemporal object structure, The sensor type attention layer modifies the spatiotemporal object structure, The object self-attention layer modifies the spatiotemporal object structure, The method performed by a computer according to claim 1, further comprising modifying the spatiotemporal object structure by an object time attention layer.

4. The spatiotemporal object structure includes, for at least one object, multiple pairs of multiple time steps, Each of the plurality of pairs includes the feature set and anchor of the object and the time step of the plurality of time steps, wherein the anchor includes the current predicted position of the object in the time step. The method performed by a computer as described in claim 1.

5. The anchor update layer further includes modifying the anchors of the plurality of pairs using the feature set of the pair, The method performed by a computer as described in claim 4.

6. The object-map attention layer allows The computer-operated method according to claim 4, further comprising modifying the spatiotemporal object structure, wherein the object-map attention layer performs cross-attention of the corresponding feature sets to map element coding according to the anchors of the plurality of pairs.

7. The acquisition of the aforementioned sensor data is To obtain multiple LiDAR sweeps in the aforementioned geographical area, The process involves voxelizing the aforementioned multiple LiDAR sweeps into multiple voxelized point clouds, A computer-based method according to claim 1, comprising running a convolutional neural network on the plurality of voxelized point clouds in order to obtain a plurality of feature maps.

8. The decoder model includes a detection decoder and a prediction decoder, and the method is The detection decoder decodes the spatiotemporal object structure to obtain the location information of at least one object within the geographical area, The method performed by a computer according to claim 1, further comprising decoding the spatiotemporal object structure using the predictive decoder to obtain the route information of the at least one object within the geographical area.

9. A computer-based method according to claim 1, further comprising decoding the spatiotemporal object structure to predict the type of at least one object.

10. Dividing the aforementioned geographical area into a grid, This further includes initializing the spatiotemporal object structure and assigning an object to each of the multiple grid cells of the grid, The method performed by a computer as described in claim 1.

11. A coarse detector model is used to process the geographic region and identify multiple predicted object locations within the geographic region. The further includes initializing the spatiotemporal object structure with the multiple predicted object positions, The method performed by a computer as described in claim 1.

12. To acquire sensor data from the sensor in the geographical area, The sensor encoding model further includes generating sensor data encoding for the sensor data, The method performed by a computer as described in claim 1.

13. It is a system, At least one computer processor, The system comprises a memory for storing computer-readable program code for causing at least one computer processor to perform an operation, and the operation is as follows: To obtain the sensor data encoding of the sensor data acquired from the sensor, Initializing the spatiotemporal object structure, To obtain the map encoding of the geographical area corresponding to the aforementioned sensor data, The transformer model iteratively modifies the spatiotemporal object structure using the sensor data coding and the map coding to generate the modified spatiotemporal object structure, The decoder model decodes the spatiotemporal object structure to obtain location and path information for at least one object within the geographical area, A system that includes outputting the aforementioned location information and the aforementioned route information.

14. The operation further includes controlling the autonomous system equipped with the sensors according to the location information and the route information. The system according to claim 13.

15. The aforementioned operation involves modifying the spatiotemporal object structure by an object-map attention layer, The sensor type attention layer modifies the spatiotemporal object structure, The object self-attention layer modifies the spatiotemporal object structure, The system according to claim 13, further comprising modifying the spatiotemporal object structure by an object time attention layer.

16. The spatiotemporal object structure includes, for at least one object, multiple pairs of multiple time steps, Each of the plurality of pairs includes the feature set and anchor of the object and the time step of the plurality of time steps, wherein the anchor includes the current predicted position of the object in the time step. The system according to claim 13.

17. The acquisition of the aforementioned sensor data is To obtain multiple LiDAR sweeps in the aforementioned geographical area, The process involves voxelizing the aforementioned multiple LiDAR sweeps into multiple voxelized point clouds, The system according to claim 13, comprising running a convolutional neural network on the plurality of voxelized point clouds in order to obtain a plurality of feature maps.

18. The decoder model includes a detection decoder and a prediction decoder, and the operation is as follows: The detection decoder decodes the spatiotemporal object structure to obtain the location information of at least one object within the geographical area, The system according to claim 13, further comprising decoding the spatiotemporal object structure using a predictive decoder to obtain the route information of the at least one object within the geographical area.

19. A non-temporary computer-readable medium comprising computer-readable program code for causing a computer system to perform an action, the action being: To obtain the sensor data encoding of sensor data acquired from a sensor, Initializing the spatiotemporal object structure, To obtain the map encoding of the geographical area corresponding to the aforementioned sensor data, The transformer model iteratively modifies the spatiotemporal object structure using the sensor data coding and the map coding to generate the modified spatiotemporal object structure, The decoder model decodes the spatiotemporal object structure to obtain location and path information for at least one object within the geographical area, A non-temporary computer-readable medium that includes outputting the aforementioned location information and the aforementioned route information.

20. The aforementioned operation involves modifying the spatiotemporal object structure by an object-map attention layer, The sensor type attention layer modifies the spatiotemporal object structure, The object self-attention layer modifies the spatiotemporal object structure, The non-temporary computer-readable medium according to claim 19, further comprising modifying the spatiotemporal object structure by an object time attention layer.