Vehicle lane position adjustment using fused scene data
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2026-08-13
Smart Images

Figure CN2025075686_13082026_PF_FP_ABST
Abstract
Description
VEHICLE LANE POSITION ADJUSTMENT USING FUSED SCENE DATAI. Field
[0001] The present disclosure is generally related to performing lane position adjustment of a vehicle. II. Description of Related Art
[0002] Vehicle automation systems seek to improve the safety and / or ease of use of a vehicle by, for example, relieving drivers of some or all of the responsibilities of navigating and controlling their vehicles. These vehicle autonomous systems use sensor data from various onboard sensors, offboard sensors, or both, to perceive their environment and make real-time decisions. Vehicle automation systems can support full autonomy (e.g., driverless) vehicle control or partial autonomy (e.g., driver supported) vehicle control. Examples of automation systems that support partial autonomy include driver assistance systems, such as adaptive cruise control systems, parking assistance systems, or lane following systems.
[0003] Many vehicle automation systems include perception systems that are intended to process sensor data to attempt to represent an environment around the vehicle. Perception data, from the perception system, can be provided to a planning system that can perform various tasks, such as predicting the future trajectories of other objects (such as vehicles and pedestrians) on the roadway. The vehicle automation systems can use such motion predictions to make decisions to control the vehicle. Each of these tasks (e.g., perception, planning, and control) is complex and can involve extensive data processing. As a result, tradeoffs arise between factors that tend to decrease demand for computing resources (e.g., processor time, memory, power, etc. ) , factors that tend to increase performance, factors that tend to increase cost, etc. For example, vehicle control performance (e.g., in terms of improved safety, legal compliance, low error rates, passenger comfort, etc. ) may be improved by the use of more sensor data; however, using more sensor data while generating results quickly enough for real-time vehicle control can significantly increase the cost of providing sufficient processing capacity.
[0004] Traffic path assist is one example of vehicle automation that assists with vehicle control by generating lane position signals based on the location of the vehicle. For example, the lane position signals generated by traffic path assistance can be used to keep the vehicle in the center of its lane. In some circumstances, traveling in the center of the lane may not be the most preferred situation. For example, it may be safer for the vehicle to move to the right of the center of the lane when a large truck is approaching in a lane to the left of the vehicle. However, due to the above-described tradeoffs, it can be difficult to implement a traffic path assistance that is powerful enough to determine optimum positioning of the vehicle within a lane to maximize safety across a large variety of possible traffic scenarios without also incurring a prohibitively large cost, latency, demand for computing resources, or combinations thereof.III. Summary
[0005] According to one implementation of the present disclosure, a device includes a memory configured to store scene data that represent a scene associated with a vehicle. The device also includes one or more processors configured to process at least a first portion of the scene data to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene. The one or more processors are configured to process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects. The one or more processors are configured to generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. The one or more processors are configured to, based on the fused scene data, determine a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects. The one or more processors are also configured to determine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0006] According to another implementation of the present disclosure, a method includes processing, at a device, at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene. The method includes processing, at the device, at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects. The method includes generating, at the device, fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. The method includes determining, at the device and based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects. The method also includes determining, at the device and based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0007] According to another implementation of the present disclosure, a non-transitory computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to process at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene. The instructions further cause the one or more processors to process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects. The instructions further cause the one or more processors to generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. The instructions further cause the one or more processors to determine, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects. The instructions further cause the one or more processors to determine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0008] Other aspects, advantages, and features of the present disclosure will become apparent after review of the entire application, including the following sections: Brief Description of the Drawings, Detailed Description, and the Claims.IV. Brief Description of the Drawings
[0009] FIG. 1 is a diagram that illustrates an example of a scene that includes multiple objects associated with a roadway and a particular illustrative aspect of a system operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0010] FIG. 2 is a diagram of an illustrative aspect of components and operations associated with the system of FIG. 1, in accordance with some examples of the present disclosure.
[0011] FIG. 3 is a block diagram of an illustrative aspect of components of the system of FIG. 1, in accordance with some examples of the present disclosure.
[0012] FIG. 4 is a diagram of an illustrative aspect of operations associated with the system of FIG. 1, in accordance with some examples of the present disclosure.
[0013] FIG. 5 is a diagram of an example of an integrated circuit operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0014] FIG. 6 is a diagram of a mobile device operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0015] FIG. 7 is a diagram of a headset operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0016] FIG. 8 is a diagram of a wearable electronic device operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0017] FIG. 9 is a diagram of a mixed reality or augmented reality glasses device operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0018] FIG. 10 is a diagram of a headset, such as a virtual reality, mixed reality, or augmented reality headset, operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0019] FIG. 11 is a diagram of an example of a vehicle operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.
[0020] FIG. 12 is a diagram of a particular implementation of a method of determining a lane position adjustment using fused scene data that may be performed by the system of FIG. 1, in accordance with some examples of the present disclosure.
[0021] FIG. 13 is a block diagram of a particular illustrative example of a device that is operable to determine a lane position adjustment using fused scene data, in accordance with some examples of the present disclosure.V.Detailed Description
[0022] The above-described problems associated with providing vehicle automation including traffic path assistance are solved by using fused scene data from various sources to reduce a computational burden associated with determining a lane position adjustment for a vehicle, as described herein. Some conventional traffic path assistance applications have a normal operational mode and a special operational mode. Such traffic path assistance applications use the normal operational mode unless a special condition (e.g., a large truck approaching in an adjacent lane) is detected, and use the special operational mode when a special condition is detected. When the normal operational mode is in use, the traffic path assistance application attempts to maintain the vehicle in alignment with the center of the lane in which the vehicle is positioned. When the special operational mode is in use, the traffic path assistance application attempts to maintain the vehicle in alignment with a path that is offset from the center of the lane. To illustrate, in order to reduce an amount of computational resources required by a traffic path assistance system, the system may be designed to detect other objects (e.g., vehicles and pedestrians) using a perception system of the vehicle and predict the future trajectories of other objects on the roadway, and then compare the object locations and predicted trajectories to a discrete set of traffic scenarios. If a matching traffic scenario is detected, the system can select a traffic path adjustment that has been pre-computed for that particular traffic scenario. However, such an approach may provide less than optimum traffic path adjustments when the object locations and predicted trajectories may warrant a traffic path adjustment but do not match one of the predetermined traffic scenarios.
[0023] In accordance with the disclosed aspects, traffic path assistance including dynamic lateral offset is provided using a fused scene data input to a machine learning model, such as an attention network and a multilayer perceptron network. The fused scene data combines data from various sources, including object data for each detected object that may be obtained via a perception system, and roadway data for portions of the roadway associated with each of the detected objects, fusing map data into the objects and reducing a data input size to the model.
[0024] According to an aspect, because the roadway data associated with each object is combined with the object data for that object, the roadway data and object data serve as a single input for the object, enabling an attention network to be used to calculate interactions of the objects with significantly reduced computational resources as compared to techniques in which the object data and the roadway data are provided as separate inputs to a model. An output of the attention network, such as a set of object-scene embeddings corresponding to the calculated interactions, can be processed by a multilayer perceptron to output a lateral offset value for a traffic assist path of an ego vehicle.
[0025] According to an aspect, one or more constraints may be applied to the lateral offset, such as a maximum lateral offset amount, a maximum rate of lateral offset, or both, for additional safety and / or comfort of occupants of vehicle.
[0026] A technical advantage provided by the disclosed techniques includes generating more accurate traffic path adjustments for a much larger variety of situations as compared to conventional traffic assistance systems that use comparisons to predetermined traffic scenarios, while also reducing the computational complexity, latency, and computing resources required to train and use a machine learning model to dynamically determine traffic path adjustments as compared to systems in which the various sources of object data and roadway data are separately input into a machine learning model.
[0027] Particular aspects of the present disclosure are described below with reference to the drawings. In the description, common features are designated by common reference numbers. As used herein, various terminology is used for the purpose of describing particular implementations only and is not intended to be limiting of implementations. For example, the singular forms “a, ” “an, ” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. Further, some features described herein are singular in some implementations and plural in other implementations. To illustrate, FIG. 1 depicts a device 102 including one or more processors ( “processor (s) ” 108 of FIG. 1) , which indicates that in some implementations the device 102 includes a single processor 108 and in other implementations the device 102 includes multiple processors 108. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural (as indicated by “ (s) ” ) unless aspects related to multiple of the features are being described.
[0028] In some drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically distinct, the same reference number is used for each, and the different instances are distinguished by addition of a letter to the reference number. When the features as a group or a type are referred to herein e.g., when no particular one of the features is being referenced, the reference number is used without a distinguishing letter. However, when one particular feature of multiple features of the same type is referred to herein, the reference number is used with the distinguishing letter. For example, referring to FIG. 1, multiple sets of fused scene data are illustrated and associated with reference numbers 168A, 168B, and 168C. When referring to a particular one of these sets of fused scene data, such as a first set of fused scene data 168A, the distinguishing letter “A” is used. However, when referring to any arbitrary one of these sets of fused scene data or to these sets of fused scene data as a group, the reference number 168 is used without a distinguishing letter.
[0029] As used herein, the terms “comprise, ” “comprises, ” and “comprising” may be used interchangeably with “include, ” “includes, ” or “including. ” Additionally, the term “wherein” may be used interchangeably with “where. ” As used herein, “exemplary” indicates an example, an implementation, and / or an aspect, and should not be construed as limiting or as indicating a preference or a preferred implementation. As used herein, an ordinal term (e.g., “first, ” “second, ” “third, ” etc. ) used to modify an element, such as a structure, a component, an operation, etc., does not by itself indicate any priority or order of the element with respect to another element, but rather merely distinguishes the element from another element having a same name (but for use of the ordinal term) . As used herein, the term “set” refers to one or more of a particular element, and the term “plurality” refers to multiple (e.g., two or more) of a particular element.
[0030] As used herein, “coupled” may include “communicatively coupled, ” “electrically coupled, ” or “physically coupled, ” and may also (or alternatively) include any combinations thereof. Two devices (or components) may be coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) directly or indirectly via one or more other devices, components, wires, buses, networks (e.g., a wired network, a wireless network, or a combination thereof) , etc. Two devices (or components) that are electrically coupled may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling, as illustrative, non-limiting examples. In some implementations, two devices (or components) that are communicatively coupled, such as in electrical communication, may send and receive signals (e.g., digital signals or analog signals) directly or indirectly, via one or more wires, buses, networks, etc. As used herein, “directly coupled” may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without intervening components.
[0031] In the present disclosure, terms such as “obtaining, ” “determining, ” “calculating, ” “estimating, ” “shifting, ” “adjusting, ” etc. may be used to describe how one or more operations are performed. It should be noted that such terms are not to be construed as limiting and other techniques may be utilized to perform similar operations. Additionally, as referred to herein, “obtaining, ” “generating, ” “calculating, ” “estimating, ” “using, ” “selecting, ” “accessing, ” and “determining” may be used interchangeably. For example, “obtaining, ” “generating, ” “calculating, ” “estimating, ” or “determining” a parameter (or a signal) may refer to actively generating, estimating, calculating, or determining the parameter (or the signal) or may refer to using, selecting, or accessing the parameter (or signal) that is already generated, such as by another component or device.
[0032] As used herein, the term “machine learning” should be understood to have any of its usual and customary meanings within the fields of computers science and data science, such meanings including, for example, processes or techniques by which one or more computers can learn to perform some operation or function without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in data and generate a result based on the analysis. For certain types of machine learning, the results that are generated include data that indicates an underlying structure or pattern of the data itself. Such techniques, for example, include so called “clustering” techniques, which identify clusters (e.g., groupings of data elements of the data) .
[0033] For certain types of machine learning, the results that are generated include a data model (also referred to as a “machine-learning model” or simply a “model” ) . Typically, a model is generated using a first data set to facilitate analysis of a second data set. For example, a first portion of a large body of data may be used to generate a model that can be used to analyze the remaining portion of the large body of data. As another example, a set of historical data can be used to generate a model that can be used to analyze future data.
[0034] Since a model can be used to evaluate a set of data that is distinct from the data used to generate the model, the model can be viewed as a type of software (e.g., instructions, parameters, or both) that is automatically generated by the computer (s) during the machine learning process. As such, the model can be portable (e.g., can be generated at a first computer, and subsequently moved to a second computer for further training, for use, or both) . Additionally, a model can be used in combination with one or more other models to perform a desired analysis. To illustrate, first data can be provided as input to a first model to generate first model output data, which can be provided (alone, with the first data, or with other data) as input to a second model to generate second model output data indicating a result of a desired analysis. Depending on the analysis and data involved, different combinations of models may be used to generate such results. In some examples, multiple models may provide model output that is input to a single model. In some examples, a single model provides model output to multiple models as input.
[0035] Examples of machine-learning models include, without limitation, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neuro-fuzzy inference systems, as well as combinations, ensembles and variants of these and other types of models. Variants of neural networks include, for example and without limitation, prototypical networks, autoencoders, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variants of decision trees include, for example and without limitation, random forests, boosted decision trees, etc.
[0036] Since machine-learning models are generated by computer (s) based on input data, machine-learning models can be discussed in terms of at least two distinct time windows –a creation / training phase and a runtime phase. During the creation / training phase, a model is created, trained, adapted, validated, or otherwise configured by the computer based on the input data (which in the creation / training phase, is generally referred to as “training data” ) . Note that the trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform particular operations, such as classification, prediction, encoding, or other data analysis or data synthesis operations. During the runtime phase (or “inference” phase) , the model is used to analyze input data to generate model output. The content of the model output depends on the type of model. For example, a model can be trained to perform classification tasks or regression tasks, as non-limiting examples. In some implementations, a model may be continuously, periodically, or occasionally updated, in which case training time and runtime may be interleaved or one version of the model can be used for inference while a copy is updated, after which the updated copy may be deployed for inference.
[0037] In some implementations, a previously generated model is trained (or re-trained) using a machine-learning technique. In this context, “training” refers to adapting the model or parameters of the model to a particular data set. Unless otherwise clear from the specific context, the term “training” as used herein includes “re-training” or refining a model for a specific data set. For example, training may include so called “transfer learning. ” In transfer learning a base model may be trained using a generic or typical data set, and the base model may be subsequently refined (e.g., re-trained or further trained) using a more specific data set.
[0038] A data set used during training is referred to as a “training data set” or simply “training data. ” The data set may be labeled or unlabeled. “Labeled data” refers to data that has been assigned a categorical label indicating a group or category with which the data is associated, and “unlabeled data” refers to data that is not labeled. Typically, “supervised machine-learning processes” use labeled data to train a machine-learning model, and “unsupervised machine-learning processes” use unlabeled data to train a machine-learning model; however, it should be understood that a label associated with data is itself merely another data element that can be used in any appropriate machine-learning process. To illustrate, many clustering operations can operate using unlabeled data; however, such a clustering operation can use labeled data by ignoring labels assigned to data or by treating the labels the same as other data elements.
[0039] Training a model based on a training data set generally involves changing parameters of the model with a goal of causing the output of the model to have particular characteristics based on data input to the model. To distinguish from model generation operations, model training may be referred to herein as optimization or optimization training. In this context, “optimization” refers to improving a metric, and does not mean finding an ideal (e.g., global maximum or global minimum) value of the metric. Examples of optimization trainers include, without limitation, backpropagation trainers, derivative free optimizers (DFOs) , and extreme learning machines (ELMs) . As one example of training a model, during supervised training of a neural network, an input data sample is associated with a label. When the input data sample is provided to the model, the model generates output data, which is compared to the label associated with the input data sample to generate an error value. Parameters of the model are modified in an attempt to reduce (e.g., optimize) the error value. As another example of training a model, during unsupervised training of an autoencoder, a data sample is provided as input to the autoencoder, and the autoencoder reduces the dimensionality of the data sample (which is a lossy operation) and attempts to reconstruct the data sample as output data. In this example, the output data is compared to the input data sample to generate a reconstruction loss, and parameters of the autoencoder are modified in an attempt to reduce (e.g., optimize) the reconstruction loss.
[0040] As used herein, a “vehicle” includes means of transportation of goods or people. For example, cars, buses, trains, airplanes, drones, trucks, boats, ships, submersibles, dirigibles, motorcycles, bicycles, etc. A driverless car is an example of a vehicle. In general, vehicle automation systems disclosed herein are configured to operate during use of a vehicle on a roadway. A “roadway” includes a physical area that is specifically configured to be traversed by a vehicle. In particular, a roadway, as the term is used herein, includes one or more lanes, where each “lane” refers to an area designated for movement of vehicles in a particular direction. Such lanes are often delineated by markings, but markings are not necessary where roadway characteristics or local laws or custom provide sufficient information to generally differentiate lanes.
[0041] Vehicle automation systems generally use sensor data from sensors onboard the vehicle, offboard the vehicle, or both. As used herein, a “sensor” includes a device (e.g., a hardware component) that generates a signal or other data that is representative of information about an environment around or near the sensor.
[0042] Sensors can be classified into two main categories: active sensors and passive sensors. An active sensor emits energy, such as electromagnetic waves or sound waves, into the environment and detects reflections (or "returns" ) of the emitted energy. Sensor data is generated based on the reflected signals. Examples of active sensors include, without limitation, radar, ultrasonic, sonar, and light detection and ranging (LiDAR) systems.
[0043] In contrast, a passive sensor detects ambient energy in the environment, such as light, radio waves, sound, or heat, that was not emitted by the sensor itself. Examples of passive sensors include, without limitation, cameras, microphones, and thermal sensors. Some sensors can operate in active or passive modes depending on the setting. For example, an infrared camera may passively detect thermal energy in a first mode and may emit infrared light for active sensing in a second mode.
[0044] As used herein “sensor data” refers to a signal or a representation of a signal output by a sensor. For example, some sensors are configured to output sensor data as a signal, such as a voltage signal, that represents information as continuous values, discontinuous values, or discrete values. In some cases, the signal is encoded to represent the sensor data. For example, a sensor can generate continuous states that are sample by digitization circuitry to generate a bits (or a bitstream) representing the sensor data. Irrespective the type of signal generated by a sensor, the signal can be sampled to generate bits which can be stored or processed, and such bits are also referred to herein as sensor data.
[0045] As used herein, a “scene” refers to conditions within an area around a vehicle and “scene data” refers to representation of the scene based on sensor data (possibly from more than one sensor) and optionally other data (e.g., map data) associated with or descriptive of the scene. A scene often includes one or more objects, such as other vehicles, pedestrians, roadway features, etc., which may be represented in or described by the scene data. One example of a type of scene data is a “bird's eye view (BEV) image, ” which represents the scene from above based on aggregation of map data and one or more images.
[0046] FIG. 1 illustrates an example of a scene 100 that includes multiple objects associated with a roadway 110 having multiple roadway features. In FIG. 1, the objects in the scene 100 include vehicles 140, 142, 144, 146, 148, and 150, and a pedestrian 152. The roadway features include lanes 112, 114, 116, 118, lane markings 130, 132, 134, 136, 138, and roadway boundary features (e.g., a sidewalk 120 and a wall 122) . The various objects and roadway features in FIG. 1 are merely one example. Different scenes can include more, fewer, or different objects; more, fewer, or different roadway features; or both.
[0047] One of the vehicles of the scene 100 is an ego vehicle (e.g., a vehicle that includes sensors for capturing aspects of the scene 100 and, for purposes of assisted driving or self-driving and navigation, uses a coordinate system centered on itself) . Although the ego vehicle can include any of the vehicles of FIG. 1, for purposes of illustration, vehicle 142 is designated the ego vehicle for the examples provided below. The ego vehicle 142 includes or is associated with a traffic path assist system 160.
[0048] The traffic path assist system 160 includes a device 102 that is coupled to (or includes) one or more sensors 104. The device 102 includes a memory 106 coupled to one or more processors 108. The memory 106 includes scene data 166 representing the scene 100 that includes the ego vehicle 142. The processor 108 includes a lane position adjuster 180 that is configured to obtain the scene data 166 and use the scene data 166 to determine a lane position adjustment 178 to control (or assist with control of) particular aspects of operation of the ego vehicle 142. The memory 106 may also include other data, such as parameters associated with one or more models used in the lane position adjuster 180, instructions that are executable by the processor 108 to perform operations described herein, other data, parameters, or instructions, etc., or any combination thereof. In a particular embodiment, the traffic path assist system 160 (e.g., the processor 108) is integrated in the ego vehicle 142 and can generate vehicle control signals that are configured to control positioning of the ego vehicle 142 within an associated lane (e.g., the lane 114) of the roadway 110.
[0049] The scene data 166 is based on sensor data 162 generated by sensors of the ego vehicle 142 (e.g., the sensor (s) 104) , such as cameras, lidar sensors, radar sensors, etc. Examples of the sensor (s) 104 are described in further detail with reference to FIG. 3. In some cases, the ego vehicle 142 can be configured to receive at least a portion of the sensor data 162, at least a portion of the scene data 166, or both, from one or more remote sensors, such as infrastructure sensors (e.g., a roadway camera) or sensors of another vehicle. To illustrate, at least a portion of the sensor data 162 can be received at the ego vehicle 142 via transmissions using a vehicle-to-vehicle (V2V) communication protocol or a vehicle-to-everything (V2X) communication protocol, as described further with reference to FIG. 3.
[0050] The scene data 166 can include raw sensor data 162 from the sensors 104, processed sensor data 162, predictions or embeddings based on the sensor data 162, other data derived from the sensor data 162, such as object data 190, roadway data 192, and localization data 194, or combinations thereof. As a specific example, sensor data 162 from one or more of the sensors 104 can be filtered, downsampled, upsampled, contrast modified, color modified, etc. As another example, some of the sensor data 162 can be combined to form a bird’s eye view representation of the scene 100. As yet another example, the processor 108 can execute machine-vision processes (e.g., edge detection, blob detection, object detection, etc. ) to annotate or modify the sensor data 162. In some examples, the processor 108 uses one or more scene models 164 (e.g., machine-learning (ML) -based perception model (s) , ML-based prediction model (s) , etc. ) to generate the object data 190, the roadway data 192, and the localization data 194 based at least partially on the sensor data 162.
[0051] The processor 108 is configured to process at least a first portion of the scene data 166 to determine object data 190, indicative of object characteristics (e.g., type, relative position, relative speed, etc. ) , for each object of a set of objects detected in the scene 100. As illustrated, the object data 190 includes N sets of object data 190A-190N (N is a positive integer) . For example, because the scene 100 includes seven objects (six vehicles 140, 142, 144, 146, 148, and 150, and one pedestrian 152) , the processor 108 generates seven sets of object data 190 (e.g., N=7) . The processor 108 is also configured to process at least a second portion of the scene data 166 to determine the roadway data 192 representing roadway characteristics (e.g., curvature, lane boundary conditions, etc. ) for respective portions of the roadway 110 associated with each of the objects, and to determine the localization data 194 representing map locations of the objects (e.g., each object’s position relative to the roadway 110, each object’s lane identification, etc. ) . Examples of the object data 190, the roadway data 192, and the localization data 194 are described in further detail with reference to FIG. 2.
[0052] The processor 108 is configured to generate fused scene data 168 for each object of the set of objects. To illustrate, a data fusion component 182 of the lane position adjuster 180 is configured to generate a set of fused scene data 168 for each object i (1 ≤i ≤ N) in the set of objects, such as fused scene data 168A for a first object (object 1) , fused scene data 168B for an ith object (2 ≤ i ≤ N-1) , and fused scene data 168C for an Nth object (object N) . The fused scene data 168B for a particular object i, such as a second object (e.g., i = 2) includes object data associated with the particular object i, roadway data associated with the particular object i, and localization data associated with the particular object i. In an illustrative, non-limiting example, the data fusion component 182 concatenates the object data, the roadway data, and the localization data for the object i into a single vector or other data structure to generate the fused scene data 168B.
[0053] The processor 108 is configured to determine, based on the fused scene data 168, a set of object-scene embeddings 174 that includes an object-scene embedding for each object of the set of objects. To illustrate, the fused scene data 168 is provided as input data 170 to a first network 172 and processed by the first network 172 to generate the object-scene embeddings 174. In an illustrative example, the first network 172 includes or corresponds to an attention network, as described in further detail with reference to FIG. 2. According to an aspect, the set of object-scene embeddings 174 includes one object-scene embedding for each object of the set of objects.
[0054] The processor 108 is configured to determine, based on the set of object-scene embeddings 174, a lane position adjustment 178 for a lateral position of the ego vehicle 142 within a lane of the roadway 110. To illustrate, the object-scene embeddings 174 are provided to a second network 176 and processed by the second network 176 to generate the lane position adjustment 178. In an illustrative example, the second network 176 includes or corresponds to multilayer perceptron (MLP) network, as described in further detail with reference to FIG. 2.
[0055] In an example, the lane position adjustment 178 can indicate lateral position of the ego vehicle 142 within the lane 114 associated with the ego vehicle 142. The lateral position can be indicated relative to a reference location associated with the lane 114, such as relative to a centerline of the lane 114 or relative to a boundary of the lane 114.
[0056] The lane position adjustment 178 is based on information about each of the objects represented by a corresponding object-scene embedding 174, which may include every object detected in the scene 100 or a subset of the objects detected in the scene 100. For example, an object-scene embedding 174 can be generated for each object that is within a threshold distance from the ego vehicle 142, where the threshold distance can be selected based on the speed of the ego vehicle 142. In this example, objects that are detected in the scene 100 but greater than the threshold distance from the ego vehicle 142 can be ignored for purposes of determining the lane position adjustment 178. Ignoring an object can include, for example, not including object data for the object in the input data 170 so that no object-scene embedding 174 is generated for the object.
[0057] Optionally, the lane position adjustment 178 can be compared to one or more constraints to generate a constrained lane position adjustment based on a lane position constraint, a rate of change constraint, or both. The lane position adjustment 178 (or a constrained lane position adjustment) can be used to provide guidance to an operator of the ego vehicle 142 (e.g., verbal instructions or instructions in a navigation display) . Additionally, or alternatively, the lane position adjustment 178 (or a constrained lane position adjustment) can be used to generate vehicle position control signals (e.g., steering control signals) to automatically reposition the ego vehicle 142 within the lane 114. Examples are described in further detail with reference to FIG. 2 and FIG. 3.
[0058] In some embodiments, the traffic path assist system 160 is integrated within the ego vehicle 142; whereas in other embodiments, the traffic path assist system 160 is distinct from and possibly remote from the ego vehicle 142. For example, the traffic path assist system 160 can be included in a mobile device used within the ego vehicle 142 or a server remote from the ego vehicle 142.
[0059] By using the object-scene embeddings 174 associated with a set of objects in the scene 100, the traffic path assist system 160 is able to generate lane position adjustment data (e.g., the lane position adjustment 178 or a constrained lane position adjustment) that accounts for many relevant aspects of the scene 100, rather than, for example, a single other vehicle in the scene 100.
[0060] In some implementations, the device 102 corresponds to or is included in one of various types of devices. In some examples, the processor 108 is integrated in at least one of a mobile phone or a tablet computer device, as described with reference to FIG. 6, a headset device, as described further with reference to FIG. 7, a wearable electronic device, as described with reference to FIG. 8, a mixed reality or augmented reality glasses device, as described with reference to FIG. 9, or a virtual reality, mixed reality, or augmented reality headset, as described with reference to FIG. 10. In another illustrative example, the processor 108 is integrated into a vehicle, such as described further with reference to FIG. 11.
[0061] FIG. 2 depicts an illustrative aspect of components and operations associated with a particular embodiment 200 of the system 160 of FIG. 1, in accordance with some examples of the present disclosure.
[0062] In the embodiment 200, the scene models 164 include an object model 206 and a roadway and location model 208. The object model 206 is configured to process a first portion 202 of the scene data 166 to generate the object data 190, and the roadway and location model 208 is configured to process a second portion 204 of the scene data 166 to generate the roadway data 192 and the localization data 194. For example, the first portion 202 can include perception data (e.g., camera data, active sensor data) for object detection, type determination, relative location and speed determination, etc., and may not include map data, while the second portion 204 can include map data and may not include some (or any) of the perception data (e.g., may include camera data but not include active sensor data) . According to some aspects, the first portion 202 and the second portion 204 are mutually exclusive (e.g., do not share any scene data 166) , while according to some other aspects, the first portion 202 and the second portion 204 are not mutually exclusive. For example, in some embodiments the first portion 202 and the second portion 204 may have one or more data sets in common. A technical advantage of using distinct models to process distinct types of data (e.g., perception data vs. map data) is that the models can be smaller, requiring less power and fewer computational resources during inference, more efficiently replaced or updated, and faster and less expensive to train as compared to a model that processes the combined data.
[0063] According to some aspects, data that is generated at the object model 206, such as object types and relative positions, can be input to the roadway and location model 208 as part of the second portion 204 of the scene data 166, data that is generated at the roadway and location model 208 can be input to the object model 206 as part of the first portion 202 of the scene data 166, or a combination thereof. Although the scene models 164 are illustrated as including the object model 206 and the roadway and location model 208, in other embodiments the object model 206 and the roadway and location model 208 can be combined into a single model, or the scene models 164 may include one or more other, or additional, models. To illustrate, the scene models 164 may include a roadway model and a location model instead of the roadway and location model 208.
[0064] The object data 190 includes object characteristics 220 associated with each object. To illustrate, the object characteristics 220 associated with a first object (e.g., the object 1 data 190A) indicates a type of the first object, a distance and direction of the first object relative to the ego vehicle 142, a movement speed and direction of the first object relative to the ego vehicle 142, or a combination thereof. In the embodiment 200, the object characteristics 220 are illustrated as including an object type 222, a relative distance and direction 224 of the object relative to the position of the ego vehicle 142, and a movement speed and direction 226 of the object relative to the ego vehicle 142.
[0065] According to an aspect, the object type 222 can indicate, for example, whether the object is a vehicle, a pedestrian, or a roadway obstacle (e.g., a caution cone, a sign, debris, an animal) , and a size of the object (e.g., within various ranges such as large, medium, or small as compared to the size of the ego vehicle 142 or the lane 114) . The relative distance and direction 224 can indicate, for example, an offset from the ego vehicle 142 in rectangular coordinates (e.g., X and Y offsets from the ego vehicle 142) , polar coordinates (e.g., radial distance and angle relative to the ego vehicle 142) , etc. The movement speed and direction 226 can similarly indicate a speed and direction of motion of the object relative to the speed and direction of motion of the ego vehicle 142 in one or more of a variety of representations, such as a signed speed (e.g., a positive speed indicates movement in the same roadway direction as the ego vehicle 142, and negative speed indicates movement in the opposite roadway direction) and direction of the relative motion of the object.
[0066] According to an aspect, the roadway data 192 includes roadway characteristics 230 associated with each of the objects. To illustrate, the roadway characteristics 230 are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof. In the embodiment 200, the roadway characteristics 230 include a curvature 232 and lane boundary conditions 234.
[0067] According to an aspect, the localization data 194 associated with a first object indicates a position of the first object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof. In the embodiment 200, the localization data 194 of an object includes a relative position 242 of the object relative to the roadway 110 and a lane identifier (ID) 244 that indicates which lane 112-118 the object is in (e.g., a lane index value, with the index =1 for the leftmost lane and the index incrementing from left to right) .
[0068] The fused scene data 168A associated with a first object (object 1) includes the following data elements: 1) the object type 222 of the first object; 2) an X-position 224A (e.g., a lateral offset from the ego vehicle 142) of the first object; 3) a Y-position 224B (e.g., a fore or aft offset from the ego vehicle 142) of the first object; 4) a signed speed 226A (e.g., positive for speeds along the same roadway direction as the ego vehicle 142 and negative for speeds in the opposite direction) of the first object; 5) a heading angle 226B (e.g., relative to a portion of the roadway 110) of the first object; a lane identifier 244 associated with the first object; 7) a left lane boundary 234A of a lane associated with the first object; 8) a right lane boundary 234B of the lane associated with the first object; 9) a left lane marking 234C of the lane associated with the first object; 10) a right lane marking of the lane associated with the first object; and 11) a lane curvature 232 of the lane associated with the first object. Although the fused scene data 168 for each object includes eleven data elements in the embodiment 200, in other examples, the fused scene data 168 associated with each object can include more, fewer, or different data elements. In general, data elements 1 to 6 can be considered object characteristics of an object, and data elements 7 to 11 can be considered roadway characteristics of a portion of the roadway 110 associated with the object.
[0069] In some cases, the object type 222 can include more specific information about an object, such as that the object is a specific type of vehicle, such as a truck, a car, a motorcycle, etc. The lane boundary conditions 234 of a lane associated with an object can include information indicative of whether the lane boundary is traversable. For example, a right lane boundary (relative to the direction of traffic flow) of the lane 112 is adjacent to the sidewalk 120 which may be considered non-traversable due to a curb height between the lane 112 and the sidewalk 120, due to legal constraints, due to pedestrian safety concerns, or a combination thereof. As another example, a right lane boundary (relative to the direction of traffic flow) of the lane 118 is adjacent to the wall 122, which may be considered non-traversable as a physical barrier. As another example, the left lane boundary (relative to the direction of traffic flow) of the lane 114 is associated with a double line lane marking 134, which may indicate that the left lane boundary is physically traversable, but not legally traversable. In contrast, the right lane boundary (relative to the direction of traffic flow) of the lane 114 is associated with a dashed line lane marking 132, which may indicate that the right lane boundary is physically traversable and legally traversable.
[0070] In a particular aspect, to determine the set of object-scene embeddings 174, the fused scene data 168 associated with the set of objects is provided to an attention network 250. As illustrated, the first network 172 includes or corresponds to the attention network 250, and the attention network 250 is configured to generate the object-scene embeddings 174 using a portion of the input data 170 that is descriptive of the ego vehicle 142 to generate a query input 252 of the attention network 250. According to an aspect, the attention network 250 generates the query input 252 via multiplication by a weight matrix Wk, and generates a key input 254 and a value input 256 via multiplication of the input data 170 by a weight matrix Wk and Wv, respectively, and processes the query input 252 (Q) , the key input 254 (K) , and the value input 256 (V) to generate output having the form (Q*KT) *V that is indicative of interactions between the objects and that corresponds to the object-scene embeddings 174. According to some aspects, the attention network 250 corresponds to a multi-headed attention network, such as a 4-head or 6-head attention network as illustrative, non-limited examples, that may provide more accurate results as compared to a single-head attention network. Although the attention network 250 performs computations that may be described as “interactions” or “similarities” between the objects during generation of the object-scene embeddings 174, it should be understood that the object-scene embeddings 174 do not correspond to physical interactions or physical similarities between the objects.
[0071] In a particular aspect, the set of object-scene embeddings 174 is provided as input to a MLP network 258 to generate the lane position adjustment 178 (e.g., as an offset from a lane center of the lane 114) . As illustrated, the second network 176 includes or corresponds to the MLP network 258.
[0072] According to an aspect, a comparator 260 is configured to compare the lane position adjustment 178 to one or more constraints 262 to generate a constrained lane position adjustment 264. In an example, the constraint 262 can include a lane position constraint, a rate of change constraint, or both. For example, a lane position constraint can include a maximum offset to the left of the lane center, a maximum offset to the right of the lane center, or both. Based on the lane position adjustment 178 indicating a lane position that exceeds the maximum offset to the left or to the right of the lane center, the comparator 260 sets the constrained lane position adjustment 264 to indicate a lane position that matches (or is less than) the maximum offset. Similarly, a lane position constraint can include a lateral offset maximum rate that limits how quickly the ego vehicle 142 can move laterally in the lane. Based on the lane position adjustment 178 indicating a lateral offset rate that exceeds the lateral offset maximum rate, the comparator 260 sets the constrained lane position adjustment 264 to indicate a lateral offset rate that matches (or is less than) the lateral offset maximum rate.
[0073] According to an aspect, the processor 108 is configured to generate one or more vehicle position control signals based on the constrained lane position adjustment 264. For example, a vehicle control system 266 can be configured to translate the constrained lane position adjustment 264 into a control signal 268 that is operable to cause the ego vehicle 142 to maneuver in accordance with the constrained lane position adjustment 264, such as by automatically adjusting a steering mechanism of the ego vehicle 142.
[0074] FIG. 3 depicts an illustrative aspect of components associated with a particular embodiment 300 of the system of FIG. 1, in accordance with some examples of the present disclosure.
[0075] In the embodiment 300, the sensor 104 may be onboard the ego vehicle 142 and configured to generate sensor data 162 corresponding to at least a subset of the scene data 166. As illustrated, the sensor 104 includes cameras 304, one or more active sensor systems 306, or a combination thereof. The cameras 304 are configured to capture images of the scene (e.g., still images, frames of video data, or a combination thereof) . The images are provided to the processor 108 and included in the sensor data 162. The processor 108 is configured to process the images to generate at least a portion of the subset of the scene data 166. In an illustrative example, the scene model 164 (e.g., the object model 206) processes the images from the cameras 304 to generate the object data 190.
[0076] The active sensor system 306 is configured to emit energy and detect returns based on the emitted energy, such as a lidar sensor, an ultrasonic sensor, or a radar sensor, as illustrative, non-limiting examples. Data representing the returns is provided to the processor 108 and included in the sensor data 162. The processor 108 is configured to process the data representing the returns to generate at least a portion of the subset of the scene data 166. In an illustrative example, the scene model 164 (e.g., the object model 206) processes the data representing the returns from the active sensor system 306 to generate the object data 190.
[0077] The device 102 also includes a modem 320 coupled to the processor 108 and configured to receive wireless data 342 from one or more remote devices 340. In an example, the device 102 is configured to receive at least a subset of the scene data 166 from the remote device 340. For example, the remote device 340 can include one or more infrastructure sensors (e.g., a roadway camera) or sensors of another vehicle. To illustrate, at least a portion of the sensor data 162 can be received at the device 102 as the wireless data 342 via transmissions using a vehicle-to-vehicle (V2V) communication protocol, a vehicle-to-everything (V2X) communication protocol, or a combination thereof. In another example, the remote device 340 corresponds to a provider of map data that can be received via the modem 320 and included in the scene data 166. In another example, the wireless data 342 can correspond to instructions, data, or both, that can be received at the device 102 and used in conjunction with operation of the lane position adjuster 180. For example, the wireless data 342 can include parameters of one or more trained ML models (e.g., the scene model 164, the first network 172, the second network 176, etc. ) included in the lane position adjuster 180.
[0078] The processor 108 is configured to process the sensor data 162 at the lane position adjuster 180 to generate lane position adjust data 310. For example, the lane position adjust data 310 can correspond to lane position adjustment 178 or the constrained lane position adjustment 264 of FIG. 2. The processor 108 also includes the vehicle control system 266, which is configured to process the lane position adjust data 310 to generate the control signal 268. The device 102 may provide the control signal 268 to one or more component 370, such as one or more mechanical or electrical components of the ego vehicle 142, to control operation of the ego vehicle 142 in accordance with the lane position adjust data 310.
[0079] In addition to sending the control signal 268 to the component 370, or alternatively, the processor 108 is configured to output instruction data 392 to a user interface device, such as a display device 390. To illustrate, the instruction data 392 may include data to cause the display device 390 to display a graphical depiction of a lane position adjustment (e.g., as a trajectory overlay on an image of the roadway 110, as a text instruction, as a graphical icon, etc. ) , such as on a navigational display or a head-up display of the ego vehicle 142, or on a wearable device, mobile device, or augmented reality device of an operator of the ego vehicle 142, to instruct the operator of the ego vehicle to perform the lane position adjustment. Alternatively, or in addition, the output instruction data 392 may be provided to a speech interface device (e.g., a speaker) to audibly instruct the operator of the ego vehicle 142, to a haptic device to signal the instruction to a wearer of the haptic device, or to any other type of user interface device.
[0080] FIG. 4 is a diagram of an illustrative aspect of operations 400 associated with the system of FIG. 1, in accordance with some examples of the present disclosure. In a particular embodiment, the operations 400 are performed by the traffic path assist system 160 (e.g., the device 102) .
[0081] The operations 400 include, at block 402, obtaining object data and lane data from perception. For example, the lane position adjuster 180 processes the sensor data 162 to generate the object data 190, the roadway data 192, and the localization data 194. The operations 400 include, at block 404, obtaining a network input including fused object data, roadway data, and localization data. For example, the lane position adjuster 180 performs data fusion to generate the fused scene data 168 that is provided as the input data 170 to the first network 172 (e.g., the attention network 250) .
[0082] The operations 400 include, at block 406, obtaining object-scene embeddings using an attention network and determining a lane position adjustment. For example, the attention network 250 processes the input data 170 to generate the object-scene embeddings 174. The object-scene embeddings 174 are processed at the second network 176 (e.g., the MLP network 258) to generate the lane position adjustment 178.
[0083] The operations 400 include, at block 408, saturating and / or constraining the lane position adjustment based on maximum host-to-left / right lane gaps and a maximum lateral offset rate. For example, the comparator 260 determines whether the lane position adjustment 178 indicates a lane position that exceeds the maximum offset to the left or to the right of the lane center, a lateral offset rate that exceeds a lateral offset maximum rate, or both, and saturates / constrains the lane position adjustment to not exceed the constraints.
[0084] The operations 400 include, at block 410, applying the lateral offset to lane center as traffic assist path. For example, the lateral offset can be applied via the control signal 268 of FIG. 2 or FIG. 3, or via the instruction data 392 to an operator of the ego vehicle, or both.
[0085] FIG. 5 depicts an implementation 500 of the device 102 as an integrated circuit 502 that includes the one or more processors 108. The integrated circuit 502 also includes input circuitry 504, such as one or more bus interfaces, to receive input data 505, such as the sensor data 162, the scene data 166, the wireless data 342, or a combination thereof. The integrated circuit 502 also includes output circuitry 506, such as a bus interface, to enable sending of output data 507, such as the lane position adjustment 178, the constrained lane position adjustment 264, the control signal 268, the instruction data 392, or a combination thereof. The integrated circuit 502 including the lane position adjuster 180 enables implementation of lane position adjustment as a component in a system, such as a mobile phone or tablet as depicted in FIG. 6, a headset as depicted in FIG. 7, a wearable electronic device as depicted in FIG. 8, a mixed reality or augmented reality glasses device as described with reference to FIG. 9, a virtual reality, mixed reality, or augmented reality headset as depicted in FIG. 10, or a vehicle as depicted in FIG. 11.
[0086] FIG. 6 depicts an example 600 in which the device 102 includes a mobile device 602, such as a phone or tablet, as illustrative, non-limiting examples. The mobile device 602 includes a display screen 604, one or more sensors (e.g., cameras) 610, and a speaker 612. The lane position adjuster 180 is integrated in the mobile device 602, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the mobile device 602. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from the one or more sensors 610, from one or more sensors of a vehicle in which the mobile device 602 is mounted (e.g., the ego vehicle 142) , from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the mobile device 602 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and transmit the resulting control signal 268 to the vehicle, display the instruction data 392 at the display screen 604, play out the instruction data 392 at the speaker 612, or any combination thereof.
[0087] FIG. 7 depicts an example 700 in which the device 102 includes a headset device 702. The headset device 702 includes one or more speakers 712. The lane position adjuster 180 is integrated in the headset device 702, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the headset device 702. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from one or more sensors of a vehicle in which the headset device 702 is worn (e.g., the ego vehicle 142) , from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the headset device 702 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and transmit the resulting control signal 268 to the vehicle, display the instruction data 392 at a display screen of the vehicle, play out the instruction data 392 at the speaker 712, or any combination thereof.
[0088] FIG. 8 depicts an example 800 in which the device 102 includes a wearable electronic device 802, illustrated as a “smart watch. ” The wearable electronic device 802 includes a display screen 804 and a speaker 812. The lane position adjuster 180 is integrated in the wearable electronic device 802, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the wearable electronic device 802. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from one or more sensors of a vehicle in which the wearable electronic device 802 is worn (e.g., the ego vehicle 142) , from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the wearable electronic device 802 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and transmit the resulting control signal 268 to the vehicle, display the instruction data 392 at the display screen 804, play out the instruction data 392 at the speaker 812, or any combination thereof. In a particular example, the wearable electronic device 802 includes a haptic device that provides a haptic notification (e.g., vibrates) associated with display of the instruction data 392. For example, the haptic notification can cause a user to look at the wearable electronic device 802 to view the instruction data 392 at the display screen 804.
[0089] FIG. 9 depicts an example 900 in which the device 102 includes a portable electronic device that corresponds to an extended reality device, such as augmented reality or mixed reality glasses 902. The glasses 902 include a holographic projection unit 904 configured to project visual data onto a surface of a lens 906 or to reflect the visual data off of a surface of the lens 906 and onto the wearer's retina. The glasses 902 include one or more sensors (e.g., cameras) 910, and one or more speakers 912. The lane position adjuster 180 is integrated in the glasses 902, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the glasses 902. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from the one or more sensors 910, from one or more sensors of a vehicle in which the glasses 902 are worn (e.g., the ego vehicle 142) , from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the glasses 902 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and transmit the resulting control signal 268 to the vehicle, display the instruction data 392 via a projection onto the surface of the lens 906, play out the instruction data 392 at the speaker 912, or any combination thereof.
[0090] FIG. 10 depicts an example 1000 in which the device 102 includes a portable electronic device that corresponds to a virtual reality, augmented reality, or mixed reality headset 1002. The headset 1002 includes a visual display device 1004, one or more sensors (e.g., cameras) 1010, and one or more speakers 1012. The lane position adjuster 180 is integrated in the headset 1002, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the headset 1002. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from the one or more sensors 1010, from one or more sensors of a vehicle in which the headset 1002 is worn (e.g., the ego vehicle 142) , from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the headset 1002 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and transmit the resulting control signal 268 to the vehicle, display the instruction data 392 at the display device 1004, play out the instruction data 3102 at the speaker 1012, or any combination thereof.
[0091] FIG. 11 depicts a second example 1100 in which the device 102 corresponds to, or is integrated within, a vehicle 1102, illustrated as a car. The vehicle 1102 includes a display screen 1104, one or more sensors (e.g., cameras or active sensors) 1110, and one or more speakers 1112. The lane position adjuster 180 is integrated in the vehicle 1102, such as in the integrated circuit 502, which is illustrated using dashed lines to indicate internal components that are not generally visible to a user of the vehicle 1102. In a particular example, the lane position adjuster 180 operates to determine a lane position adjustment based on sensor data 162 obtained from the one or more sensors 1110, from one or more remote devices (e.g., infrastructure sensors or remote vehicles) , or any combination thereof. For example, the vehicle 1102 may process the sensor data 162 using the lane position adjuster 180 to generate the lane position adjust data 310, and send the resulting control signal 268 to one or more components of the vehicle (e.g., the component 370) , display the instruction data 392 at the display screen 1104, play out the instruction data 3102 at the speaker 1112, or any combination thereof.
[0092] Referring to FIG. 12, a particular implementation of a method 1200 of determining a lane position adjustment using fused scene data is shown. In a particular aspect, one or more operations of the method 1200 are performed by at least one of the scene model (s) 164, the data fusion component 182, the first network 172, the second network 176, the lane position adjuster 180, the processor 108, the device 102, the system 160 of FIG. 1, or a combination thereof.
[0093] The method 1200 includes, at block 1202, processing, at a device, at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene. For example, the scene model (s) 164 (e.g., the object model 206) process at least a portion of the scene data 166 (e.g., the first portion 202) to generate the object data 190. In some embodiments, the object characteristics associated with a first object indicate a type of the first object, a distance and direction of the first object relative to the vehicle, a movement speed and direction of the first object, or a combination thereof. For example, the object characteristics 220 of the object 1 data 190A indicate the object type 222, the relative distance and direction 224, and the movement speed and direction 226 associated with a first object detected in the scene 100.
[0094] The method 1200 includes, at block 1204, processing, at the device, at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects. For example, the scene model (s) 164 (e.g., the roadway and location model 208) process at least a portion of the scene data 166 (e.g., the second portion 204) to generate the roadway data 192 and the localization data 194. In some embodiments, the roadway characteristics are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof. For example, the roadway characteristics 230 include the curvature 232 and the lane boundary conditions 234. In some embodiments, the localization data associated with a first object indicate a position of the object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof. For example, the localization data 194 of FIG. 2 includes the relative position 242 and the lane ID 244.
[0095] The method 1200 includes, at block 1206, generating, at the device, fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. For example, the data fusion component 182 of the lane position adjuster 180 generates the fused scene data 168.
[0096] The method 1200 includes, at block 1208, determining, at the device and based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects. For example, the first network 172 processes the input data 170 to generate the set of object-scene embeddings 174. In some embodiments, determining the set of object-scene embeddings includes providing the fused scene data associated with the set of objects to an attention network, such as the attention network 250, using data descriptive of the vehicle as query input.
[0097] The method 1200 includes, at block 1210, determining, at the device and based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway. For example, the second network 176 processes the set of object-scene embeddings 174 to generate the lane position adjustment 178. In some embodiments, determining the lane position adjustment includes providing the set of object-scene embeddings as input to a multilayer perceptron, such as the MLP network 258, to generate the lane position adjustment as an offset from a lane center of the lane.
[0098] The method 1200 of FIG. 12 may be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC) , a processing unit such as a central processing unit (CPU) , a digital signal processor (DSP) , a controller, another hardware device, firmware device, or any combination thereof. As an example, the method 1200 of FIG. 12 may be performed by a processor that executes instructions, such as described with reference to FIG. 13.
[0099] Referring to FIG. 13, a block diagram of a particular illustrative implementation of a device is depicted and generally designated 1300. In various implementations, the device 1300 may have more or fewer components than illustrated in FIG. 13. In an illustrative implementation, the device 1300 may correspond to the device 102. In an illustrative implementation, the device 1300 may perform one or more operations described with reference to FIGS. 1-12.
[0100] In a particular implementation, the device 1300 includes a processor 1306 (e.g., a central processing unit (CPU) ) . The device 1300 may include one or more additional processors 1310 (e.g., one or more DSPs) . In a particular aspect, the processor 108 of FIG. 1 corresponds to the processor 1306, the processors 1310, or a combination thereof. The processors 1310 may include a speech and music coder-decoder (CODEC) 1308 that includes a voice coder ( “vocoder” ) encoder 1336, a vocoder decoder 1338, the lane position adjuster 180, or a combination thereof.
[0101] In this context, the term “processor” refers to an integrated circuit consisting of logic cells, interconnects, input / output blocks, clock management components, memory, and optionally other special purpose hardware components, designed to execute instructions and perform various computational tasks. Examples of processors include, without limitation, central processing units (CPUs) , digital signal processors (DSPs) , neural processing units (NPU) , graphics processing units (GPUs) , field programmable gate arrays (FPGAs) , microcontrollers, quantum processors, coprocessors, vector processors, other similar circuits, and variants and combinations thereof. In some cases, a processor can be integrated with other components, such as communication components, input / output components, etc. to form a system on a chip (SOC) device or a packaged electronic device.
[0102] Taking CPUs as a starting point, a CPU typically includes one or more processor cores, each of which includes a complex, interconnected network of transistors and other circuit components defining logic gates, memory elements, etc. A core is responsible for executing instructions to, for example, perform arithmetic and logical operations. Typically, a CPU includes an Arithmetic Logic Unit (ALU) that handles mathematical operations and a Control Unit that generates signals to coordinate the operation of other CPU components, such as to manage operations a fetch-decode-execute cycle.
[0103] CPUs and / or individual processor cores generally include local memory circuits, such as registers and cache to temporarily store data during operations. Registers include high-speed, small-sized memory units intimately connected to the logic cells of a CPU. Often registers include transistors arranged as groups of flip-flops, which are configured to store binary data. Caches include fast, on-chip memory circuits used to store frequently accessed data. Caches can be implemented, for example, using Static Random-Access Memory (SRAM) circuits.
[0104] Operations of a CPU (e.g., arithmetic operations, logic operations, and flow control operations) are directed by software and firmware. At the lowest level, the CPU includes an instruction set architecture (ISA) that specifies how individual operations are performed using hardware resources (e.g., registers, arithmetic units, etc. ) . Higher level software and firmware is translated into various combinations of ISA operations to cause the CPU to perform specific higher-level operations. For example, an ISA typically specifies how the hardware components of the CPU move and modify data to perform operations such as addition, multiplication, and subtraction, and high-level software is translated into sets of such operations to accomplish larger tasks, such as adding two columns in a spreadsheet. Generally, a CPU operates on various levels of software, including a kernel, an operating system, applications, and so forth, with each higher level of software generally being more abstracted from the ISA and usually more readily understandable by human users.
[0105] GPUs, NPUs, DSPs, microcontrollers, coprocessors, FPGAs, ASICS, and vector processors include components similar to those described above for CPUs. The differences among these various types of processors are generally related to the use of specialized interconnection schemes and ISAs to improve a processor’s ability to perform particular types of operations. For example, the logic gates, local memory circuits, and the interconnects therebetween of a GPU are specifically designed to improve parallel processing, sharing of data between processor cores, and vector operations, and the ISA of the GPU may define operations that take advantage of these structures. As another example, ASICs are highly specialized processors that include similar circuitry arranged and interconnected for a particular task, such as encryption or signal processing. As yet another example, FPGAs are programmable devices that include an array of configurable logic blocks (e.g., interconnect sets of transistors and memory elements) that can be configured (often on the fly) to perform customizable logic functions.
[0106] The device 1300 may include a memory 1386 and a CODEC 1334. The memory 1386 may include instructions 1356, that are executable by the one or more additional processors 1310 (or the processor 1306) to implement the functionality described with reference to the lane position adjuster 180, or both. The device 1300 may include a modem 1348 coupled, via a transceiver 1350, to an antenna 1352. According to an aspect, the modem 1348 corresponds to the modem 320.
[0107] The device 1300 may include a display 1328 coupled to a display controller 1326. In some implementations, the device 1300 may include one or more sensors 1394, such as the sensor (s) 104. For example, the sensor (s) 1394 can include one or cameras 1396 (e.g., the cameras 304) , one or more other sensors such as active sensor (s) 1398 (e.g., the active sensor system (s) 306) , or a combination thereof. One or more speakers 1392, microphone (s) 1390, or both, may be coupled to the CODEC 1334. The CODEC 1334 may include a digital-to-analog converter (DAC) 1302, an analog-to-digital converter (ADC) 1304, or both. In a particular implementation, the CODEC 1334 may receive analog signals from the microphone (s) 1390, convert the analog signals to digital signals using the analog-to-digital converter 1304, and provide the digital signals to the speech and music codec 1308. In a particular implementation, the speech and music codec 1308 may provide digital signals to the CODEC 1334. The CODEC 1334 may convert the digital signals to analog signals using the digital-to-analog converter 1302 and may provide the analog signals to the speaker (s) 1392.
[0108] In a particular implementation, the device 1300 may be included in a system-in-package or system-on-chip device 1322. In a particular implementation, the memory 1386, the processor 1306, the processors 1310, the display controller 1326, the CODEC 1334, the transceiver 1350, and the modem 1348 are included in the system-in-package or system-on-chip device 1322. In a particular implementation, an input device 1330 and a power supply 1344 are coupled to the system-in-package or the system-on-chip device 1322. Moreover, in a particular implementation, as illustrated in FIG. 13, the display 1328, the input device 1330, the speaker (s) 1392, the microphone (s) 1390, the sensor (s) 1394, the antenna 1352, and the power supply 1344 are external to the system-in-package or the system-on-chip device 1322. In a particular implementation, each of the display 1328, the input device 1330, the speaker (s) 1392, the microphone (s) 1390, the sensor (s) 1394, the antenna 1352, and the power supply 1344 may be coupled to a component of the system-in-package or the system-on-chip device 1322, such as an interface or a controller.
[0109] The device 1300 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a gaming console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aerial vehicle, a home automation system, a voice-activated device, a wireless speaker and voice activated device, a portable electronic device, a car, a computing device, a communication device, an internet-of-things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.
[0110] In conjunction with the described implementations, an apparatus includes means for processing at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene. For example, the means for processing at least a first portion of scene data that represent a scene associated with a vehicle can include the scene model (s) 164, the lane position adjuster 180, the processor (s) 108, the device 102, the system 160, the object model 206, the integrated circuit 502, the processor 1306, the processor (s) 1310, the system-in-package or the system-on-chip device 1322, the device 1300, other circuitry configured to process at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene, or a combination thereof.
[0111] The apparatus also includes means for processing at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects. For example, the means for processing at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects can include the scene model (s) 164, the lane position adjuster 180, the processor (s) 108, the device 102, the system 160, the roadway and location model 208, the integrated circuit 502, the processor 1306, the processor (s) 1310, the system-in-package or the system-on-chip device 1322, the device 1300, other circuitry configured to process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects, or a combination thereof.
[0112] The apparatus also includes means for generating fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. For example, the means for generating fused scene data for each object of the set of objects can include the data fusion component 182, the lane position adjuster 180, the processor (s) 108, the device 102, the system 160, the integrated circuit 502, the processor 1306, the processor (s) 1310, the system-in-package or the system-on-chip device 1322, the device 1300, other circuitry configured to generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object, or a combination thereof.
[0113] The apparatus includes means for determining, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects. For example, the means for determining, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects can include the first network 172, the lane position adjuster 180, the processor (s) 108, the device 102, the system 160, the attention network 250, the integrated circuit 502, the processor 1306, the processor (s) 1310, the system-in-package or the system-on-chip device 1322, the device 1300, other circuitry configured to determine, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects, or a combination thereof.
[0114] The apparatus also includes means for determining, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway. For example, the means for determining, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway can include the second network 176, the lane position adjuster 180, the processor (s) 108, the device 102, the system 160, the MLP network 258, the comparator 260, the integrated circuit 502, the processor 1306, the processor (s) 1310, the system-in-package or the system-on-chip device 1322, the device 1300, other circuitry configured to determine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway, or a combination thereof.
[0115] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as the memory 106 or the memory 1386) includes instructions (e.g., the instructions 1356) that, when executed by one or more processors (e.g., the one or more processors 108, the one or more processors 1310, or the processor 1306) , cause the one or more processors to process at least a first portion (e.g., the first portion 202) of scene data (e.g., the scene data 166) that represent a scene (e.g., the scene 210) associated with a vehicle (e.g., the ego vehicle 142) to determine object data (e.g., the object data 190) , indicative of object characteristics (e.g., the object characteristics 220) , for each object of a set of objects (e.g., the vehicles 140, 142, 144, 146, 148, and 150, and the pedestrian 152) detected in the scene. The instructions, when executed by the one or more processors, cause the one or more processors to process at least a second portion (e.g., the second portion 204) of the scene data to determine roadway data (e.g., the roadway data 192) representing roadway characteristics (e.g., the roadway characteristics 230) associated with a roadway (e.g., the roadway 110) and localization data (e.g., the localization data 194) representing map locations of objects of the set of objects. The instructions, when executed by the one or more processors, cause the one or more processors to generate fused scene data (e.g., the fused scene data 168) for each object of the set of objects, the fused scene data for a particular object (e.g., the fused scene data 168A) including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object. The instructions, when executed by the one or more processors, cause the one or more processors to determine, based on the fused scene data, a set of object-scene embeddings (e.g., the object-scene embeddings 174) , the set of object-scene embeddings including an object-scene embedding for each object of the set of objects, and determine, based on the set of object-scene embeddings, a lane position adjustment (e.g., the lane position adjustment 178) for a lateral position of the vehicle within a lane (e.g., the lane 114) of the roadway.
[0116] Particular aspects of the disclosure are described below in sets of interrelated Examples:
[0117] According to Example 1, a device includes a memory configured to store scene data that represent a scene associated with a vehicle; and one or more processors configured to process at least a first portion of the scene data to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene; process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects; generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object; based on the fused scene data, determine a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; and determine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0118] Example 2 includes the device of Example 1, wherein the object characteristics associated with a first object indicate a type of the first object.
[0119] Example 3 includes the device of Example 1 or Example 2, wherein the object characteristics associated with a first object indicate a distance and direction of the first object relative to the vehicle.
[0120] Example 4 includes the device of any of Examples 1 to 3, wherein the object characteristics associated with a first object indicate a movement speed and direction of the first object.
[0121] Example 5 includes the device of any of Examples 1 to 4, wherein the object characteristics associated with a first object indicate a type of the first object, a distance and direction of the first object relative to the vehicle, a movement speed and direction of the first object, or a combination thereof.
[0122] Example 6 includes the device of any of Examples 1 to 5, wherein the roadway characteristics are descriptive of curvature of the roadway.
[0123] Example 7 includes the device of any of Examples 1 to 6, wherein the roadway characteristics are descriptive lane boundary conditions associated with the roadway.
[0124] Example 8 includes the device of any of Examples 1 to 7, wherein the roadway characteristics are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof.
[0125] Example 9 includes the device of any of Examples 1 to 8, wherein the localization data associated with a first object indicate a position of the object relative to the roadway.
[0126] Example 10 includes the device of any of Examples 1 to 9, wherein the localization data associated with a first object indicate a lane identifier associated with a lane of the roadway in which the first object is located.
[0127] Example 11 includes the device of any of Examples 1 to 10, wherein the localization data associated with a first object indicate a position of the object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof.
[0128] Example 12 includes the device of any of Examples 1 to 11, wherein, to determine the set of object-scene embeddings, the one or more processors are configured to provide the fused scene data associated with the set of objects to an attention network using data descriptive of the vehicle as query input.
[0129] Example 13 includes the device of any of Examples 1 to 12, wherein, to determine the lane position adjustment, the one or more processors are configured to provide the set of object-scene embeddings as input to a multilayer perceptron to generate the lane position adjustment as an offset from a lane center of the lane.
[0130] Example 14 includes the device of any of Examples 1 to 13, wherein the one or more processors are configured to compare the lane position adjustment to a lane position constraint, a rate of change constraint, or both, to generate a constrained lane position adjustment; and generate one or more vehicle position control signals based on the constrained lane position adjustment.
[0131] Example 15 includes the device of any of Examples 1 to 14, wherein the one or more processors are integrated within the vehicle.
[0132] Example 16 includes the device of any of Examples 1 to 15 and further includes a modem coupled to the one or more processors and configured to receive at least a subset of the scene data from a remote device.
[0133] Example 17 includes the device of any of Examples 1 to 16 and further includes one or more sensors onboard the vehicle and coupled to the one or more processors, wherein the one or more sensors are configured to generate sensor data corresponding to at least a subset of the scene data.
[0134] Example 18 includes the device of Example 17, wherein the one or more sensors comprise one or more cameras configured to capture images of the scene, wherein the one or more processors are configured to process the images to generate at least a portion of the subset of the scene data.
[0135] Example 19 includes the device of Example 17 or Example 18, wherein the one or more sensors comprise one or more active sensor systems configured to emit energy and detect returns based on the emitted energy, wherein the one or more processors are configured to process data representing the returns to generate at least a portion of the subset of the scene data.
[0136] According to Example 20, a method includes processing, at a device, at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene; processing, at the device, at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects; generating, at the device, fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object; determining, at the device and based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; and determining, at the device and based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0137] Example 21 includes the method of Example 20, wherein the object characteristics associated with a first object indicate a type of the first object.
[0138] Example 22 includes the method of Example 20 or Example 21, wherein the object characteristics associated with a first object indicate a distance and direction of the first object relative to the vehicle.
[0139] Example 23 includes the method of any of Examples 20 to 22, wherein the object characteristics associated with a first object indicate a movement speed and direction of the first object.
[0140] Example 24 includes the method of any of Examples 20 to 23, wherein the object characteristics associated with a first object indicate a type of the first object, a distance and direction of the first object relative to the vehicle, a movement speed and direction of the first object, or a combination thereof.
[0141] Example 25 includes the method of any of Examples 20 to 24, wherein the roadway characteristics are descriptive of curvature of the roadway.
[0142] Example 26 includes the method of any of Examples 20 to 25, wherein the roadway characteristics are descriptive lane boundary conditions associated with the roadway.
[0143] Example 27 includes the method of any of Examples 20 to 26, wherein the roadway characteristics are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof.
[0144] Example 28 includes the method of any of Examples 20 to 27, wherein the localization data associated with a first object indicate a position of the object relative to the roadway.
[0145] Example 29 includes the method of any of Examples 20 to 28, wherein the localization data associated with a first object indicate a lane identifier associated with a lane of the roadway in which the first object is located.
[0146] Example 30 includes the method of any of Examples 20 to 29, wherein the localization data associated with a first object indicate a position of the object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof.
[0147] Example 31 includes the method of any of Examples 20 to 30, wherein determining the set of object-scene embeddings includes providing the fused scene data associated with the set of objects to an attention network using data descriptive of the vehicle as query input.
[0148] Example 32 includes the method of any of Examples 20 to 31, wherein determining the lane position adjustment includes providing the set of object-scene embeddings as input to a multilayer perceptron to generate the lane position adjustment as an offset from a lane center of the lane.
[0149] Example 33 includes the method of any of Examples 20 to 32 and further includes comparing the lane position adjustment to a lane position constraint, a rate of change constraint, or both, to generate a constrained lane position adjustment; and generating one or more vehicle position control signals based on the constrained lane position adjustment.
[0150] Example 34 includes the method of any of Examples 20 to 33 and further includes receiving at least a subset of the scene data from a remote device.
[0151] Example 35 includes the method of any of Examples 20 to 34 and further includes generating, at one or more sensors onboard the vehicle, sensor data corresponding to at least a subset of the scene data.
[0152] Example 36 includes the method of Example 35, wherein the one or more sensors include one or more cameras that capture images of the scene, and further includes processing the images to generate at least a portion of the subset of the scene data.
[0153] Example 37 includes the method of Example 35 or Example 36, wherein the one or more sensors include one or more active sensor systems that emit energy and detect returns based on the emitted energy, and further includes processing data representing the returns to generate at least a portion of the subset of the scene data.
[0154] According to Example 38, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method of any of Examples 20 to 37.
[0155] According to Example 39, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform the method of any of Examples 20 to 37.
[0156] According to Example 40, an apparatus includes means for carrying out the method of any of Examples 20 to 37.
[0157] According to Example 41, a non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to process at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene; process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects; generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object; determine, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; and determine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0158] According to Example 42, an apparatus includes means for processing at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene; means for processing at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects; means for generating fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object; means for determining, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; and means for determining, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.
[0159] Those of skill would further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the implementations disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. Various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor executable instructions depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, such implementation decisions are not to be interpreted as causing a departure from the scope of the present disclosure.
[0160] The steps of a method or algorithm described in connection with the implementations disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module may reside in random access memory (RAM) , flash memory, read-only memory (ROM) , programmable read-only memory (PROM) , erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , registers, hard disk, a removable disk, a compact disc read-only memory (CD-ROM) , or any other form of non-transient storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor may read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC) . The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or user terminal.
[0161] The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope possible consistent with the principles and novel features as defined by the following claims.
Claims
1.A device comprising:a memory configured to store scene data that represent a scene associated with a vehicle; andone or more processors configured to:process at least a first portion of the scene data to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene;process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects;generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object;based on the fused scene data, determine a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; anddetermine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.2.The device of claim 1, wherein the object characteristics associated with a first object indicate a type of the first object, a distance and direction of the first object relative to the vehicle, a movement speed and direction of the first object, or a combination thereof.3.The device of claim 1, wherein the roadway characteristics are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof.4.The device of claim 1, wherein the localization data associated with a first object indicate a position of the object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof.5.The device of claim 1, wherein, to determine the set of object-scene embeddings, the one or more processors are configured to provide the fused scene data associated with the set of objects to an attention network using data descriptive of the vehicle as query input.6.The device of claim 1, wherein, to determine the lane position adjustment, the one or more processors are configured to provide the set of object-scene embeddings as input to a multilayer perceptron to generate the lane position adjustment as an offset from a lane center of the lane.7.The device of claim 1, wherein the one or more processors are configured to:compare the lane position adjustment to a lane position constraint, a rate of change constraint, or both, to generate a constrained lane position adjustment; andgenerate one or more vehicle position control signals based on the constrained lane position adjustment.8.The device of claim 1, wherein the one or more processors are integrated within the vehicle.9.The device of claim 1, further comprising a modem coupled to the one or more processors and configured to receive at least a subset of the scene data from a remote device.10.The device of claim 1, further comprising one or more sensors onboard the vehicle and coupled to the one or more processors, wherein the one or more sensors are configured to generate sensor data corresponding to at least a subset of the scene data.11.The device of claim 10, wherein the one or more sensors comprise one or more cameras configured to capture images of the scene, wherein the one or more processors are configured to process the images to generate at least a portion of the subset of the scene data.12.The device of claim 10, wherein the one or more sensors comprise one or more active sensor systems configured to emit energy and detect returns based on the emitted energy, wherein the one or more processors are configured to process data representing the returns to generate at least a portion of the subset of the scene data.13.A method comprising:processing, at a device, at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene;processing, at the device, at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects;generating, at the device, fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object;determining, at the device and based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; anddetermining, at the device and based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.14.The method of claim 13, wherein the object characteristics associated with a first object indicate a type of the first object, a distance and direction of the first object relative to the vehicle, a movement speed and direction of the first object, or a combination thereof.15.The method of claim 13, wherein the roadway characteristics are descriptive of curvature of the roadway, lane boundary conditions associated with the roadway, or a combination thereof.16.The method of claim 13, wherein the localization data associated with a first object indicate a position of the object relative to the roadway, a lane identifier associated with a lane of the roadway in which the first object is located, or a combination thereof.17.The method of claim 13, wherein determining the set of object-scene embeddings includes providing the fused scene data associated with the set of objects to an attention network using data descriptive of the vehicle as query input.18.The method of claim 13, wherein determining the lane position adjustment includes providing the set of object-scene embeddings as input to a multilayer perceptron to generate the lane position adjustment as an offset from a lane center of the lane.19.The method of claim 13, further comprising:comparing the lane position adjustment to a lane position constraint, a rate of change constraint, or both, to generate a constrained lane position adjustment; andgenerating one or more vehicle position control signals based on the constrained lane position adjustment.20.A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to:process at least a first portion of scene data that represent a scene associated with a vehicle to determine object data, indicative of object characteristics, for each object of a set of objects detected in the scene;process at least a second portion of the scene data to determine roadway data representing roadway characteristics associated with a roadway and localization data representing map locations of objects of the set of objects;generate fused scene data for each object of the set of objects, the fused scene data for a particular object including object data associated with the particular object, roadway data associated with the particular object, and localization data associated with the particular object;determine, based on the fused scene data, a set of object-scene embeddings, the set of object-scene embeddings including an object-scene embedding for each object of the set of objects; anddetermine, based on the set of object-scene embeddings, a lane position adjustment for a lateral position of the vehicle within a lane of the roadway.