System and method for generating a vehicle-integrated environment representation

A hybrid method combining conventional and learning-based sensor fusion techniques generates a comprehensive vehicle environment representation, addressing the challenge of supporting all driving automation levels and domains by enhancing feature extraction and ensuring safety in ADAS and autonomous driving.

JP2026509853APending Publication Date: 2026-03-25MERCEDES BENZ GROUP AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing advanced driver assistance systems (ADAS) and autonomous driving technologies face challenges in generating accurate and comprehensive vehicle environment representations that can support all levels of driving automation and operational design domains, requiring versatile models to capture complex environmental patterns for scene understanding and safety-critical sensor data processing.

Method used

A hybrid approach combining conventional grid-based sensor fusion with learned sensor data processing, using inverse sensor modeling and machine learning techniques to generate a fused environment representation, integrating classical and learning-based methods to enhance feature extraction and data optimization.

Benefits of technology

This hybrid method provides a more robust and scalable environment representation, enabling advanced ADAS and autonomous driving functions by capturing high-level features and ensuring safety through interpretable sensor data processing, supporting all SAE levels and operational design domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026509853000001_ABST
    Figure 2026509853000001_ABST
Patent Text Reader

Abstract

The vehicle computing system can receive raw sensor data in both a conventional sensor data processing module and a trained sensor data processing module. Each module can reproject the sensor data in the BEV space and optionally perform sensor fusion when multiple sensor data types are being processed. The system can then combine a trained BEV grid map or volume with a conventional BEV grid map or volume to generate a hybrid BEV representation of the vehicle's surrounding environment, and process the hybrid BEV representation of the surrounding environment to derive a fused representation of the vehicle's surrounding environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] An advanced driver assistance system (ADAS) for a vehicle automatically performs driver assistance functions such as collision warning, blind spot monitoring, inter-vehicle distance control, emergency braking, automatic parking, vehicle lane centering, and lane following by using sensor information. The Society of Automotive Engineers (SAE) provides multiple levels of driving automation, with level 0 corresponding to no driving automation and level 5 corresponding to full driving automation.

Summary of the Invention

[0002] In this specification, a system and method for dynamically generating a sensor fusion surrounding environment representation of a vehicle are described using (i) a conventional grid such as a bird's eye view (BEV) grid map or volume that is generated from raw sensor data and in which sensor measurements are reprojected into a grid or volume using geometric formulas, and (ii) a learned BEV volume generated by a learned sensor data processing module. A computing system of a vehicle can receive raw sensor data from a sensor suite of the vehicle that can include a set of multiple sensor types such as LIDAR sensors, radar sensors, cameras or other image sensors, ultrasonic sensors, etc. In a particular example, the computing system does not use a machine learning model to generate a conventional BEV grid map or volume, and processes raw sensor data measurements from the vehicle's sensor suite so that a set of sensor data captured within a conventional BEV grid from a conventional reprojection is represented, and can execute a conventional reprojection and grid sensor fusion module to generate a conventional BEV grid map or volume in a classical manner.

[0003] Simultaneously, the computing system can execute a trained sensor data processing module to generate one or more trained 2D or 3D BEV grid maps or volumes based on raw sensor data. In some embodiments, the trained BEV grid map or volume may include a single, sensor-fused trained BEV volume generated from each of several sensor types (fused LiDAR, radar, and image data). In variants, the trained BEV grid map or volume may be a single trained BEV grid map or volume based on a single sensor data type (e.g., image data), or it may be a combined (e.g., concatenated) BEV grid map or volume from separate sensor data maps or volumes for each of one or more sensor data types, such as a trained BEV grid map or volume composed of image data, a trained BEV grid map or volume composed of radar data, and / or a trained BEV grid map or volume composed of LiDAR data.

[0004] In certain implementations, a conventional sensor data processing module can perform inverse sensor modeling on raw sensor measurements to generate a conventional BEV grid map or volume. In further implementations, a computing system can combine a conventional BEV grid map or volume with a trained BEV map or volume along a spatial dimension such that features from the BEV grid map or volume and trained BEV maps or volumes are spatially correlated. The computing system can then process the resulting hybrid BEV representation of the vehicle's surrounding environment to, for example, derive a fused representation of the surrounding environment, derive various aspects of the road infrastructure on which the vehicle operates (e.g., lane markings, road topology, lane topology, crosswalks, etc.), perform scene understanding tasks (e.g., object detection and classification, determination of traffic priority rules, etc.), determine grid occupancy, perform motion prediction, perform driver assistance tasks, or autonomously operate the vehicle along a driving route.

[0005] The disclosures herein are illustrative and not limiting, and in the accompanying drawings, similar reference numbers refer to similar elements. [Brief explanation of the drawing]

[0006] [Figure 1] This is a block diagram illustrating an exemplary computing system for generating a vehicle fusion environment representation, as described in this specification. [Figure 2] This is a block diagram illustrating an exemplary vehicle computing system, including a dedicated module for generating a sensor-fused environment representation of a vehicle, as described in this specification. [Figure 3] This figure shows an example of a vehicle that acquires sensor data from multiple sensor types to generate a fused representation of the surrounding environment, as described in this specification. [Figure 4]This flowchart illustrates an exemplary method for combining a conventional BEV grid map or volume with a learned BEV map or volume to perform an automated vehicle task, as described in the examples of this specification. [Figure 5] This flowchart illustrates an exemplary method for combining a conventional BEV grid map or volume with a learned BEV map or volume to perform an automated vehicle task, as described in the examples of this specification. [Modes for carrying out the invention]

[0007] This specification describes systems and methods for generating vehicle environmental representations that are suitable for supporting all SAE levels and all operational design domains (ODDs) and are certifiable for all required functional safety standards. Operational planning and driver assistance techniques in real-world environments require versatile models of the environment so that high-level features of the environment representing complex patterns are captured to perform scene understanding tasks, and so that sensor measurements can be processed in an interpretable manner for safety purposes. A vehicle computing system may include conventional reprojection and grid sensor data processing modules and trained sensor data processing modules, each capable of receiving raw sensor data from various sensors in the vehicle. Conventional sensor data processing modules can use captured sensor data to generate a conventional BEV grid map using an inverse sensor model of the vehicle's surrounding environment (e.g., a 2D grid map or a 3D grid volume), while trained sensor data processing modules can generate a trained BEV map or volume of the vehicle's surrounding environment based on one or more downstream tasks (e.g., ADAS tasks, autonomous driving tasks, scene understanding tasks, grid occupancy determination tasks, motion prediction, motion planning, etc.).

[0008] In various implementations, computing systems can concatenate or otherwise combine conventional BEV grid maps or volumes with trained BEV maps or volumes to generate a hybrid grid-based BEV representation of the vehicle's surrounding environment. The use of classical grid maps can provide a versatile representation of the vehicle's surrounding environment, and as a result, integrating grid maps or volumes with trained maps or volumes (e.g., using conventional inverse sensor modeling) is intended to ensure that occupancy information and other sensor-based details are available, for example, to the vehicle's motion planning unit or environmental analysis module.

[0009] In further implementations, the computing system may utilize a hybrid grid-based BEV representation to derive a fused representation of the vehicle's surrounding environment. For example, the computing system may implement a machine learning decoder on the hybrid grid-based BEV representation to generate a fused representation of the surrounding environment. The computing system may then utilize this fused representation of the vehicle's surrounding environment to perform various tasks such as object detection, scene understanding, and grid occupancy determination in order to facilitate (facilitate) assisted driving or autonomous vehicle operation.

[0010] Among other advantages, the examples described herein achieve the technical effect of dynamically creating a sensor-fused hybrid grid-based BEV representation of the vehicle's surrounding environment using both classical and learning-based approaches. As provided herein, the classical or conventional approach utilizes an inverse sensor model to reproject measurements from sensors onto a BEV grid or volume using geometric formulas, which can be advantageous, for example, for LIDAR and radar data. Furthermore, the learning-based approach utilizes data to learn a mapping from raw sensor measurements to a grid or volume, which can be advantageous, for example, for image data. In various implementations, the conventional and learning-based approaches can generate a BEV map or volume of a single sensor data type, or sensor fusion can be performed using multiple sensor data types. Then, the conventional BEV map or volume and the learning-based BEV map or volume can be combined to generate a hybrid BEV map or volume that incorporates both approaches. This hybrid integration can provide a higher level of features necessary for scene understanding and data-driven feature optimization, which is intended to facilitate the scalability of the various examples described herein to support advanced SAE levels and ODD.

[0011] As provided herein, the reprojection operation corresponds to mapping from a sensor reference frame to a BEV grid or volume reference frame and can be learned using an inverse sensor model and / or coordinate transformation or performed conventionally. As further provided herein, the sensor fusion operation can be learned using a probabilistic formula or performed conventionally to combine sensor data from multiple sensors and sensor types and to aggregate individual sensor measurements at the cell level. Thus, raw sensor data may be transferred to BEV space by reprojection using an inverse sensor model and coordinate transformation, and then sensor fusion may be performed.

[0012] As provided herein, each vehicle may encode the collected sensor data using an autoencoder that includes a neural network capable of reducing the dimensionality of the sensor data. In various applications, the neural network may include a bottleneck architecture that provides an autoencoder that reduces the dimensionality of the sensor data. A decoder may be used for scene reconstruction to allow comparison between the encoded sensor data and the original sensor data, and the loss may include the difference between the original sensor data and the reconstruction. Data "compression" may include the smaller resulting dimensionality from the autoencoder, which may result in fewer units of memory required to store the data. In some examples, variational autoencoders are mounted and run on the vehicle, and as a result, normal distribution resampling with KL divergence loss may be made to rely only on the likelihood of the observed data to compress the encoded data. The use of variational autoencoders on the vehicle is intended to further enhance compression due to the fact that memory is not wasted on useless data or data that resembles noise. Therefore, throughout this disclosure, the use of the terms “autoencoder” or “machine learning encoder” may refer to a neural network (any general autoencoder, variational autoencoder, or other learning-based encoder) that performs the dimensionality reduction and / or data compression techniques described herein.

[0013] As further provided herein, a sensor data “volume,” a BEV “volume,” or a trained BEV “volume” refers to captured sensor data and / or reconstructed sensor data that provide a three-dimensional representation of an environment captured by a set of sensors. A machine learning model (e.g., a machine learning encoder / decoder) can generate a trained BEV volume from raw sensor data that may include fused sensor data from multiple sensor types (e.g., LiDAR, radar, and imaging). In doing so, the machine learning model can discard certain captured data, compress and / or encode the sensor data, expand the sensor data, and / or generate a reconstruction of a real-world three-dimensional environment based on the compressed and encoded sensor data. The examples described herein may further refer to a BEV grid map or BEV grid volume, which may include a two-dimensional, three-dimensional, or any n-dimensional discretized space (e.g., a space further including a temporal dimension). Such terms may be used interchangeably throughout this disclosure.

[0014] In certain implementations, a computing system may perform one or more of the functions described herein using a learning-based approach, such as by running an artificial neural network (e.g., a recurrent neural network, a convolutional neural network, etc.) or one or more machine learning models to process each set of trajectories and classify the driving behavior of vehicles driven by each person passing through an intersection. Such a learning-based approach can further be adapted to a computing system that stores or includes one or more machine learning models. In one embodiment, the machine learning models may include unsupervised learning models. In one embodiment, the machine learning models may include neural networks (e.g., deep neural networks) or other types of machine learning models, including nonlinear and / or linear models. The neural networks may include feedforward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some exemplary machine learning models may leverage attentional mechanisms such as self-attention. For example, some exemplary machine learning models may include multi-head self-attention models (e.g., transformer models).

[0015] As provided herein, “Network” or “one or more networks” can include any type of network or combination of networks that enable communication between devices. In one embodiment, a network may include one or more of the following: a local area network, a wide area network, the Internet, a secure network, a cellular network, a mesh network, a peer-to-peer communication link, or a combination thereof, and may include any number of wired or wireless links. Communication over a network(s) can be achieved via a network interface using, for example, any type of protocol, protection scheme, encoding, format, packaging, etc.

[0016] As further provided herein, an “autonomous map” or “autonomous driving map” may include a ground truth map recorded by a mapping vehicle using various sensors (e.g., LiDAR sensors and / or a set of cameras or other imaging devices) and labeled (manually or automatically) to indicate traffic objects and / or traffic right-of-way rules at any given location. In a variant, the autonomous map may include a scene reconstructed using a decoder from encoded sensor data recorded and compressed by the vehicle. For example, a given autonomous map may be labeled by a human based on observed traffic signs, traffic signals and lane markings within the ground truth map. In a further example, reference points or other points of interest may be further labeled on the autonomous map for additional assistance to the autonomous vehicle. The autonomous vehicle or self-driving vehicle may then utilize the labeled autonomous map to perform positioning, attitude determination, change detection and various other actions required for autonomous driving on public roads. For example, an autonomous vehicle can refer to an autonomous map to determine traffic rules (e.g., speed limits) at its current location, and can dynamically compare live sensor data from its onboard sensor suite with the corresponding autonomous map to safely navigate along its current route.

[0017] One or more examples described herein provide that methods, techniques, and operations performed by a computing device are performed programmatically or as computer implementations. “Programmatically,” as used herein, means through the use of code or computer executable instructions. These instructions can be stored in one or more memory resources of the computing device. Steps performed programmatically may or may not be automatic.

[0018] One or more examples described herein can be implemented using a programmatic module, engine, or component. A programmatic module, engine, or component may include a program, a subroutine, a portion of a program, or a software or hardware component capable of performing one or more defined tasks or functions. Where used herein, a module or component may reside on a hardware component independent of other modules or components. Alternatively, a module or component may be a shared element or shared process of other modules, programs, or machines.

[0019] Some of the examples described herein may generally require the use of a computing device, including processing and memory resources. For example, one or more of the examples described herein may be implemented entirely or partially using network equipment (e.g., a router) on a computing device such as a server and / or a personal computer. Memory resources, processing resources, and network resources may all be used in connection with the establishment, use, or implementation of any of the examples described herein (including any implementation of any method or any implementation of any system).

[0020] Furthermore, one or more examples described herein may be implemented through the use of instructions executable by one or more processors. These instructions may be carried on non-temporary computer-readable media. The machines shown in, or described using, the following figures provide examples of processing resources and computer-readable media capable of carrying and / or executing instructions for implementing the examples disclosed herein. In particular, many of the machines shown with the examples of the present invention include processors and various forms of memory for holding data and instructions. Examples of non-temporary computer-readable media include persistent memory storage devices such as hard drives in personal computers or servers. Other examples of computer storage media include portable storage units such as flash memory or magnetic memory. Computers, terminals, and network-enabled devices are all examples of machines and devices that utilize processors, memory, and instructions stored on computer-readable media. Furthermore, examples may be implemented in the form of computer programs or computer-usable carrying media capable of carrying such programs.

[0021] Exemplary computing system

[0022] Figure 1 is a block diagram illustrating an exemplary computing system for generating a fused environment representation of a vehicle, according to an example described herein. In one embodiment, the computing system 100 may include a control circuit 110, which may include one or more processors (e.g., microprocessors), one or more processing cores, a programmable logic circuit (PLC), or a programmable logic / gate array (PLA / PGA), a field programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or any other control circuit. In some implementations, the control circuit 110 and / or the computing system 100 may be part of, or form part of, a vehicle control unit (also referred to as a vehicle controller) that is embedded in or otherwise disposed in a vehicle (e.g., a Mercedes-Benz® car or van). For example, the vehicle controller may be an infotainment system controller (e.g., an infotainment head unit), a telematics control unit (TCU), an electronic control unit (ECU), a central powertrain controller (CPC), a central exterior and interior controller (CEIC), a zone controller, or any other controller (the term "or" is used herein interchangeably with "and / or"), or may include them. In a variant, the control circuit 110 and / or the computing system 100 may be contained in one or more servers (e.g., backend servers).

[0023] In one embodiment, the control circuit 110 may be programmed by one or more computer-readable instructions or computer-executable instructions stored in a non-temporary computer-readable medium 120. The non-temporary computer-readable medium 120 may be a memory device also called a data storage device, and the memory device may include an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. The non-temporary computer-readable medium 120 may, for example, form a floppy disk, a hard disk drive (HDD), a solid state drive (SDD), solid state integrated memory, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), dynamic random access memory (DRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), and / or memory stick. In some cases, the non-temporary computer-readable medium 120 may store computer-executable instructions or computer-readable instructions, such as instructions that perform the methods described below in relation to Figures 4A, 4B, 5A, and 5B.

[0024] In various embodiments, the terms "computer-readable instructions" and "computer-executable instructions" are used to describe software instructions or computer code configured to perform various tasks and operations. In various embodiments, when computer-readable instructions or computer-executable instructions form a module, the term "module" broadly refers to a set of software instructions or code configured to cause the control circuit 110 to perform one or more functional tasks. Modules and computer-readable instructions / computer-executable instructions may be described as performing various operations or tasks when the control circuit 110 or other hardware components execute the module or computer-readable instructions.

[0025] In a further embodiment, the computing system 100 can include a communication interface 140 that enables communication via one or more networks 150 for sending and receiving data. The communication interface 140 may include any circuitry, component, software, etc. for communicating via one or more networks 150 (e.g., local area network, wide area network, Internet, secure network, cellular network, mesh network, and / or peer-to-peer communication link). In some implementations, the communication interface 140 may include one or more of, for example, a communication controller, receiver, transceiver, transmitter, port, conductor, software, and / or hardware for communicating data / information.

[0026] As an illustrative embodiment, computing system 100 can be present on a vehicle-mounted computing system, can receive raw sensor data from a vehicle's sensor suite, the sensor suite including a plurality of sensor types (e.g., a combination of LIDAR, image, and radar data). The computing system can execute a conventional reprojection and grid sensor fusion module to generate a conventional BEV grid map using the raw sensor data. The computing system can further execute a learned sensor fusion module to generate at least one learned BEV map or volume based on the raw sensor data. The computing system can then combine one or more learned BEV volumes with a conventional BEV grid map or volume to generate a hybrid BEV representation of the vehicle's surrounding environment, and can process the hybrid BEV representation of the surrounding environment to derive a fused representation of the surrounding environment, which can be utilized to perform various assistance and / or autonomous driving tasks described throughout this disclosure.

[0027] Description of the System

[0028] FIG. 2 is a block diagram showing an exemplary vehicle computing system including a dedicated module for generating a sensor fusion environmental representation of a vehicle, according to an example described herein. In a particular example, vehicle computing system 200 can include any vehicle manufactured by an OEM, or any vehicle modified to include a sensor suite 205 including a set of sensors such as an image sensor (e.g., a camera), a LIDAR sensor, a radar sensor, an ultrasonic sensor, etc., and can be included in a vehicle driven by a consumer. Additionally or alternatively, vehicle computing system 200 can be included in a special mapping vehicle that operates to collect and / or encode ground truth sensor data of a road network (road grid) for scene understanding tasks and / or to generate an occupancy map for autonomous vehicle operation on the road network.

[0029] As the vehicle operates across the road network, the sensor suite 205 can collect sensor data of multiple sensor types, including LIDAR data, image or video data, radar data, etc. In various examples, the vehicle computing system 200 may include a trained sensor data processing module 210 that processes the sensor data for one or more purposes, such as semi-autonomous or fully autonomous driving, ADAS actions, or scene understanding for autonomous mapping. In various implementations, the trained sensor data processing module 210 may transfer raw sensor data to the BEV space (e.g., using inverse sensor models and / or coordinate transformations), perform sensor fusion on multiple types of sensor data, and / or encode the sensor fusion-based data (e.g., reducing spatial dimension, discarding specific parts of the sensor data, and / or compressing the sensor data). In a particular example, the trained sensor data processing module 210 may process the sensor data to generate a set of feature maps according to a set of rules or filters (e.g., a set of feature detectors). In some examples, the trained sensor data processing module 210 may perform further coordinate transformations to output a trained sensor-fused BEV volume or a set of trained BEV volumes of different sensor data types.

[0030] For example, in a particular implementation, the trained sensor data processing module 210 can generate image-based BEV volumes from image data, LIDAR-based BEV volumes from LIDAR data, radar-based BEV volumes from radar data, and / or additional BEV volumes based on additional sensor data. The separate BEV volumes may be linked together by the hybrid BEV module 230 of the vehicle computing system 200, or otherwise combined. In a variant, the trained sensor data processing module 210 can transfer raw sensor data into BEV space by reprojection using an inverse sensor model and / or coordinate transformation, and then fuse sensor data from multiple sensor types to generate a sensor-fused BEV volume (e.g., including image data, LIDAR data, radar data, etc.) to be processed by the hybrid BEV module 230.

[0031] As illustrated herein, the vehicle computing system 200 may include a conventional sensor data processing module 220 that performs reprojection and inverse sensor modeling on raw sensor data from multiple sensor types of the sensor suite 205 in a classical manner (non-learning approach). For example, the conventional sensor data processing module 220 may use geometric formulas to reproject sensor measurements onto a BEV grid or volume, and then utilize an inverse sensor model to fuse sensor data from image sensors, LiDAR sensors, and / or radar sensors to generate a sensor-fused BEV grid map or volume. In further examples, the conventional sensor data processing module 200 may perform coordinate transformations to generate a conventional BEV grid map, which may include a conventional BEV grid based on classical techniques. As provided herein, the conventional BEV grid map may include a two-dimensional BEV grid map or a three-dimensional BEV grid volume. The conventional sensor data processing module 220 may include a conventional reprojection and grid module in which a single sensor modality (e.g., LiDAR data) is reprojected into the BEV space, or it may include a conventional reprojection and grid sensor fusion module that can output a sensor fusion BEV grid map based on multiple types of raw sensor data (e.g., LiDAR and radar data) to the hybrid BEV module 230.

[0032] In certain implementations, the conventional sensor data processing module 220 and the trained sensor data processing module 210 can perform sensor fusion for all sensor modalities (e.g., image, LiDAR, and radar data) using both classical and trained approaches, respectively. In modified implementations, modules 210 and 220 can perform reprojection techniques for a single sensor data type (without fusion) or for a subset of selected sensor data types (with sensor fusion). For example, the conventional sensor data processing module 220 can perform reprojection into BEV space using signal types of sensor data (e.g., LiDAR data) where sensor fusion is not performed, or a subset of multiple types of sensor data (e.g., LiDAR and radar data), and these sensor data types are then fused into a BEV grid map or volume. In an additional example, the trained sensor data processing module 210 can reproject a single sensor data type (e.g., image data) into the BEV space (where sensor fusion is not performed), or it can reproject a subset of multiple sensor data types into the BEV space and perform sensor fusion accordingly.

[0033] Therefore, the conventional BEV grid or volume output by the conventional sensor fusion module is

[0034] The hybrid BEV module 230 can concatenate or otherwise combine learned BEV volumes (one or more) from the learned sensor data processing module 210 and BEV grid maps from the conventional sensor data processing module 220 to generate a hybrid BEV grid volume. Thus, the hybrid BEV grid volume may include higher-level features representing complex patterns in the vehicle's surrounding environment from the learned BEV volumes (one or more) and conventional sensor data grids from the sensor-fused BEV grid maps, which can be physically verified, for example, to support advanced SAE level proof. As provided herein, such higher-level features may include information that would otherwise be lost from the sensor data encoding process by the learned sensor data processing module 210.

[0035] In various implementations, the vehicle computing system 200 may include a machine learning decoder 240 that runs on a hybrid BEV grid map to derive a fused environment representation (FER) (e.g., a grid-based representation) of the area surrounding the vehicle. In a particular example, the machine learning decoder 240 processes encoded sensor data within a hybrid BEV grid volume to generate or otherwise derive a fused environment representation, which in a particular example may include a three-dimensional reconstruction of the surrounding environment. Thus, the fused environment representation may include a combination of a learning-based BEV grid and a conventional BEV grid (using an inverse sensor model) of the vehicle's surrounding environment to perform any number of functions.

[0036] In various examples, the vehicle computing system 200 may include a vehicle control module 250 that can dynamically analyze a fused environment representation. In some embodiments, the vehicle control module 250 may include an advanced driver assistance system (ADAS) that can analyze the fused environment representation to perform driver assistance functions such as inter-vehicle distance control, emergency braking assistance, lane keeping, lane centering, highway driver assistance, autonomous obstacle avoidance, and / or autonomous parking tasks. Thus, the vehicle control module 250 may operate a set of vehicle control mechanisms 260 to perform these tasks. As provided herein, the control mechanisms 260 may include the vehicle's steering system, braking system, acceleration system, and / or signaling system, as well as auxiliary systems.

[0037] In a modified form, the vehicle control module 250 may include an autonomous action planning module that automatically determines a series of immediate action plans for a vehicle along a travel route based on a fused environment representation, and operates the vehicle control mechanism 260 to drive the vehicle autonomously along the travel route according to the immediate action plan. Thus, the vehicle control module 250 can perform scene understanding tasks such as determining occupancy in a grid-based representation of the surrounding environment (for example, predicting whether each grid or 3D voxel in a fused grid-based representation does not contain an object or is occupied by an object), performing object detection and classification tasks, determining priority right-of-way rules in any given situation, and / or generally determining various aspects of road infrastructure, such as detecting lane markings, road signs, traffic signals, etc., determining lane and road topology, and identifying pedestrian crossings, bicycle lanes, etc.

[0038] Integrated sensor environment

[0039] Figure 3 shows an example of a vehicle 310 that acquires sensor data from multiple sensor types to generate a fused representation of the surrounding environment 300 of the vehicle 310, as described herein. In the example of Figure 3, the vehicle 310 may include various sensors such as one or more LiDAR sensors 322, cameras 324, and radar sensors 330. As provided herein, the raw sensor data from these various sensors can be reprojected into the BEV space and then combined or fused by the vehicle 310's computing system 200 to generate a conventional BEV grid map or volume. In a further implementation, the vehicle 310's computing system 200 may further process the raw sensor data using a trained sensor data processing module 210 to reproject the raw sensor measurements into the BEV space and generate a sensor-fused trained BEV map or volume of the surrounding environment 300 of the vehicle 310.

[0040] For example, the computing system 200 of the vehicle 310 performs a reprojection and / or coordinate transformation into the BEV space and then uses multiple sensor views 303 (e.g., stereoscopic or 3D image streams of the environment 300, one or more 3D LiDAR point cloud maps, and / or radar sensor views) based on the various sensor types included in the vehicle 310 to implement the sensor fusion and hybrid BEV techniques described herein. In various examples, conventional BEV grid maps using classical methods can capture higher levels of features that might be lost in the encoding process by the trained sensor data processing module 210. As one example, the machine learning decoder 240 can use the trained BEV volume output by the machine learning encoder 310 to generate a representation of the surrounding environment 300, which may include objects such as pedestrians 304, parking meters 327, other vehicles 325, traffic signs, and signals.

[0041] Using a conventional BEV grid map or volume generated by a conventional sensor data processing module 220, the machine learning decoder 240 can further generate a representation of the surrounding environment 300 to include more fine-grained features, or to enable the vehicle's computing system 200 to identify and / or classify such fine-grained features. Thus, the hybrid BEV grid volume output by the hybrid BEV module 230 in Figure 2 can be processed by the machine learning decoder 240 to generate a sensor-fused representation of the surrounding environment so that features such as sidewalks 321, roadside curbs 329, lane markings 328, and pedestrian crossings 315 can be identified and their fine-grained features can be determined. In such an example, the computing system 200 can determine the properties of the lane markings 328 (e.g., dashed vs. solid lines, curved arrows, bicycle lanes, etc.), the characteristics of the curbs 329 (e.g., whether the curbs 329 have a 90-degree structure, or whether they have a roadway approach structure or a parking lot approach structure), and so on. Therefore, the embodiments described herein can facilitate the capture and classification of these finer features that may be desired or necessary for scene understanding and feature optimization, which can further facilitate the scalability of the various examples described herein to support, for example, advanced SAE levels and ODD.

[0042] methodology

[0043] Figures 4 and 5 are flowcharts illustrating an exemplary method for combining a conventional BEV grid map or volume with a learned BEV volume to perform an automated vehicle task, as described herein. In the following description of Figures 4 and 5, reference numerals may be used to refer to various features illustrated and described in relation to Figures 1 and 2. Furthermore, the processes described in relation to Figures 4 and 5 can be performed by an exemplary computing system 200, such as the one described in relation to Figure 2. Moreover, certain steps described in relation to the flowcharts of Figures 4 and 5 may be performed before, simultaneously with, or after any other step, and do not necessarily have to be performed in each illustrated sequence.

[0044] Referring to Figure 4, in block 400, the vehicle computing system 200 can receive raw sensor data from the vehicle's sensor suite 205 by both the conventional sensor data processing module 220 and the trained sensor data processing module 210. As described throughout this disclosure, the sensor suite 205 may include one or more sensor types such as LIDAR sensors, cameras, and radar sensors. In one example, the sensor suite 205 may include an array of a single sensor type (e.g., cameras). In a variant, the sensor suite 205 may include multiple sensor types in any combination and arrangement. In block 405, the conventional sensor data processing module 220 of the computing system 200 can generate a BEV grid map or volume using the raw sensor data and a conventional reprojection model. As provided herein, the BEV grid map or volume may include a conventional grid map or volume using an inverse sensor model. As further provided herein, including a conventional grid map or volume may facilitate safety certification for the vehicle computing system 200 (e.g., suitable for a specific ODD level or SAE level).

[0045] In block 410, the trained sensor data processing module 210 can generate at least one trained BEV volume based on raw sensor data from the sensor suite 205. In certain embodiments, the trained sensor data processing module 210 can generate a trained BEV volume for each sensor data type, or it can generate a trained sensor fusion BEV volume containing one sensor type or a subset of multiple sensor data types. In various examples, the trained BEV volume can be encoded based on a set of feature filters of a machine learning model (an encoder / decoder model for semi-autonomous or fully autonomous driving). Thus, the trained sensor data processing module 210 can filter, discard, encode, and / or compress various parts of the raw sensor data in order to generate a trained BEV volume.

[0046] In block 415, the vehicle computing system 200 can concatenate or otherwise combine learned BEVV volumes (one or more) with classical BEV grid maps / volumes to generate a hybrid BEV representation of the vehicle's surrounding environment. In block 420, the vehicle computing system 200 can then process the hybrid BEV representation to derive a fused representation of the vehicle's surrounding environment. The computing system 200 can dynamically analyze the fused representation to perform any number and combination of tasks, as described below in relation to Figure 5.

[0047] Figure 5 is another flowchart illustrating how a learned BEV volume is combined with a classic BEV grid map to implement automated vehicle functions, as described in the examples herein. Referring to Figure 5, in block 500, the vehicle computing system 200 can receive raw sensor data from various sensors of the vehicle's sensor suite 205. In various examples, the sensor suite 205 may include different sensor types (LIDAR, camera, radar, ultrasound, etc.) that output LIDAR data in block 501, radar data in block 502, image data in block 503, and / or ultrasonic data in block 504.

[0048] In block 505, the computing system 200 may run the conventional sensor data processing module 220 on raw sensor data to generate a conventional BEV grid map or volume using an inverse sensor model and / or grid sensor fusion, as described herein. As further described herein, the conventional BEV grid map or volume may include a two-dimensional or three-dimensional grid map. In block 510, the computing system 200 may run the trained sensor data processing module 210 on raw sensor data to generate a set of feature maps, which may include sensor data encoded based on a set of feature filters of the trained sensor data processing module 210. In block 515, the trained sensor data processing module 210 may further generate one or more trained BEV grid maps and / or volumes using the set of feature maps.

[0049] As described herein, a conventional sensor data processing module 220 can reproject a single sensor data type into BEV space to generate a BEV grid or volume of a single sensor data type (e.g., a LIDAR grid map or volume), or can reproject multiple sensor data types into BEV space to generate a sensor-fused BEV grid map or volume (e.g., based on LIDAR and radar data). In a further implementation, a trained sensor data processing module 210 can reproject a single sensor data type into BEV space to generate a single sensor data type BEV grid or volume (e.g., an image-based BEV grid map or volume), or can reproject multiple sensor data types into BEV space to generate a sensor-fused BEV grid map or volume based on multiple sensor data types.

[0050] In various examples, in block 520, the computing system 200 can generate a hybrid BEV representation of the vehicle's surrounding environment using a conventional BEV grid map or volume using classical methods and a trained BEV grid map or volume. As described throughout this disclosure, the hybrid BEV representation of the vehicle's surrounding environment may include grid integration using conventional inverse sensor modeling and a learning-based approach for feature optimization that enables proof from a functional safety standpoint. In block 525, the computing system 200 can further derive various aspects of the road infrastructure in the vehicle's surrounding environment based on the hybrid BEV representation.

[0051] Various aspects of road infrastructure can include road topology, lane topology, lane boundaries, road markings, crosswalks, sidewalks, parking spaces, bicycle lanes, road signs and traffic signs, traffic signals, right-of-way rules, etc. In a further example, the computing system 200 can dynamically analyze the hybrid BEV representation and / or fused grid representation of the surrounding environment to perform autonomous tasks such as object detection, object classification, instance segmentation, motion prediction, or traffic rule determination tasks for semi-autonomous or fully autonomous driving. Thus, in block 530, the computing system 200 can dynamically implement advanced driver assistance functions for the vehicle driver, such as inter-vehicle distance control, emergency braking assistance, lane keeping, lane centering, highway driver assistance, autonomous obstacle avoidance, and / or autonomous parking functions. Additionally or alternatively, in block 535, the computing system 200 can dynamically generate an action plan for autonomously operating the vehicle's control mechanisms 260 along the driving route.

[0052] It is intended that the examples described herein be extended to individual elements and concepts described herein, independently of other concepts, ideas, or systems, and that combinations of elements described in any part of this application be included as examples. While examples are described in detail herein with reference to the accompanying drawings, it should be understood that the concepts are not limited to those exact examples. Therefore, many modifications and variations will be apparent to those skilled in the art. Accordingly, the scope of the concepts is intended to be defined by the following claims and their equivalents. Furthermore, it is intended that any specific feature described individually or as part of an example can be combined with any other feature or part of another example described individually, even if other features and examples do not refer to that particular feature.

Claims

1. A computing system for autonomous driving or assisted driving, wherein the computing system is One or more processors, The system comprises a memory that stores instructions, and when an instruction is executed by one or more processors, it is executed by the computing system. A conventional sensor data processing module and a trained sensor data processing module receive raw sensor data from a vehicle sensor suite, which includes multiple sensor types. The conventional sensor data processing module generates a conventional bird's-eye view (BEV) grid map or volume based on the raw sensor data. The aforementioned trained sensor data processing module generates a trained BEV grid map or volume based on the raw sensor data. In order to generate a hybrid BEV representation of the surrounding environment of the vehicle, the learned BEV grid map or volume is combined with the conventional BEV grid map or volume. In order to derive a fused representation of the surrounding environment of the vehicle, the hybrid BEV representation of the surrounding environment is processed. Computing system.

2. The computing system according to claim 1, wherein the plurality of sensor types include a plurality of LIDAR sensor types, image sensor types, radar sensor types, and ultrasonic sensor types.

3. The computing system according to claim 1, wherein the conventional sensor data processing module performs inverse sensor modeling on raw sensor measurements in order to generate the conventional BEV grid map.

4. The computing system according to claim 1, wherein the trained sensor data processing module generates a set of feature maps using the raw sensor data, and generates the trained BEV grid map or volume using the set of feature maps.

5. The computing system according to claim 1, wherein the instructions to be executed cause the computing system to process the hybrid BEV representation of the surrounding environment in order to derive an embodiment of the road infrastructure of the route on which the vehicle operates in real time, wherein the embodiment of the road infrastructure includes one or more of road topology, lane topology, lane boundaries, road markings, crosswalks, sidewalks, parking spaces, bicycle lanes, road signs and traffic signs, traffic signals, or right-of-way rules.

6. The computing system according to claim 1, wherein the instructions to be executed cause the computing system to process the hybrid BEV representation of the surrounding environment in order to perform a scene understanding task.

7. The computing system according to claim 6, wherein the scene understanding task includes at least one of object detection, object classification, instance segmentation, motion prediction, or traffic rule determination tasks.

8. The computing system according to claim 1, wherein the conventional BEV grid map or volume includes one of a two-dimensional BEV grid map, a three-dimensional grid volume, or an arbitrary n-dimensional discretized space.

9. The computing system according to claim 1, wherein the learned BEV grid map or volume includes a sensor-fused learned BEV grid map or volume based on the raw sensor data from the plurality of sensor types.

10. The computing system according to claim 1, wherein the trained BEV grid map or volume is generated using image data, and the conventional BEV grid map or volume is generated using at least one of LIDAR data or radar data.

11. The vehicle includes an autonomous vehicle, and the instructions to be executed are further transmitted to the computing system, The computing system according to claim 1, which dynamically analyzes the fused representation of the surrounding environment in order to autonomously operate the set of control mechanisms of the autonomous vehicle along the driving route.

12. The computing system according to claim 11, wherein the set of control mechanisms includes a plurality of the following of the autonomous vehicle: (i) an acceleration system, (ii) a braking system, (iii) a steering system, or (iv) a signaling system.

13. The computing system includes an advanced driver assistance system (ADAS), and the instructions to be executed are further transmitted to the computing system. The computing system according to claim 1, which dynamically analyzes the fused representation of the surrounding environment in order to assist the driver of the vehicle while the driver is operating the vehicle.

14. The computing system according to claim 13, wherein the command to be executed causes the ADAS to automatically perform one or more of the following tasks: inter-vehicle distance control, emergency braking assistance, lane keeping, lane centering, highway assistance, autonomous obstacle avoidance, or autonomous parking task, thereby assisting the driver of the vehicle.

15. A non-temporary computer-readable medium storing instructions, wherein, when an instruction is executed by one or more processors of a computing system, the computing system... A conventional sensor data processing module and a trained sensor data processing module receive raw sensor data from a vehicle sensor suite, which includes multiple sensor types. The conventional sensor data processing module generates a conventional bird's-eye view (BEV) grid map or volume based on the raw sensor data. The aforementioned trained sensor data processing module generates a trained BEV grid map or volume based on the raw sensor data. In order to generate a hybrid BEV representation of the surrounding environment of the vehicle, the learned BEV grid map or volume is combined with the conventional BEV grid map or volume. In order to derive a fused representation of the surrounding environment of the vehicle, the hybrid BEV representation of the surrounding environment is processed. Non-temporary computer-readable media.

16. The non-temporary computer-readable medium according to claim 15, wherein the plurality of sensor types include a plurality of LIDAR sensor types, image sensor types, radar sensor types, and ultrasonic sensor types.

17. The conventional sensor data processing module performs inverse sensor modeling on raw sensor measurements to generate the conventional BEV grid map, the non-temporary computer-readable medium according to claim 15.

18. The non-temporary computer-readable medium according to claim 15, wherein the trained sensor data processing module generates a set of feature maps using the raw sensor data, and generates the trained BEV grid map or volume using the set of feature maps.

19. The non-temporary computer-readable medium according to claim 15, wherein the instructions to be executed cause the computing system to process the hybrid BEV representation of the surrounding environment in order to derive an embodiment of the road infrastructure of the route on which the vehicle operates in real time, the embodiment of the road infrastructure includes one or more of road topology, lane topology, lane boundaries, road markings, crosswalks, sidewalks, parking spaces, bicycle lanes, road signs and traffic signs, traffic signals, or right-of-way rules.

20. A computer implementation method for autonomous driving or assisted driving, wherein the method is carried out by one or more processors. Conventional sensor data processing modules and trained sensor data processing modules receive raw sensor data from a vehicle sensor suite which includes multiple sensor types, The conventional sensor data processing module generates a conventional bird's-eye view (BEV) grid map or volume based on the raw sensor data, The aforementioned trained sensor data processing module generates a trained BEV grid map or volume based on the raw sensor data, In order to generate a hybrid BEV representation of the surrounding environment of the vehicle, the learned BEV grid map or volume is combined with the conventional BEV grid map or volume, To derive a fused representation of the surrounding environment of the vehicle, the hybrid BEV representation of the surrounding environment is processed. A computer implementation method for autonomous driving or driver assistance, including the above.