Multi-modal encoder channel fusion with cross-modal awareness

By training a single-modal encoder in a multimodal fusion pipeline, and using a LiDAR or RADAR encoder to determine point cloud data features and fuse them with the camera encoder channel, the problems of distraction and inaccurate recognition in environmental perception of vehicle driver assistance systems are solved, thereby improving object recognition and vehicle safety.

CN121444091APending Publication Date: 2026-01-30QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480045037.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-18
Filing Date
2024-05-30
Publication Date
2026-01-30

AI Technical Summary

Technical Problem

Existing driver assistance systems for vehicles suffer from distraction and slow response in terms of environmental perception, especially in complex environments where they struggle to accurately identify objects, leading to an increased risk of traffic accidents.

Method used

By training a single-modal encoder in a multimodal fusion pipeline, using a LiDAR or RADAR encoder to determine the perspective and bird's-eye view features of point cloud data, and adding them in layers with the channels of the camera encoder, image frame feature generation is improved, thus achieving early multimodal channel fusion.

Benefits of technology

It improves the vehicle's situational awareness of its surroundings, enhances the accuracy of object recognition and vehicle safety, reduces collision risks, and improves the response speed and accuracy of driver assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121444091A_ABST
    Figure CN121444091A_ABST
Patent Text Reader

Abstract

This disclosure provides systems, methods, and apparatus for a vehicle driving assistance system that supports image processing. In a first aspect, a method includes receiving an image frame representing a scene; receiving point cloud data representing a scene; determining a first set of image frame features; determining a second set of point cloud data features based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on a first feature set in the plurality of first feature sets of the image frame and a second feature set in the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frames and the fourth feature set of the point cloud data. Other aspects and features are also claimed and described.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of U.S. Patent Application No. 18 / 354,074, filed July 18, 2023, entitled “MULTI-MODAL ENCODER CHANNELFUSION WITH CROSS-MODALITY AWARENESS,” the entire contents of which are expressly incorporated herein by reference. Technical Field

[0002] The aspects of this disclosure generally relate to driver-operated or driver-assisted vehicles, and more specifically to methods and systems suitable for providing driving assistance or for autonomous driving. Background Technology

[0003] Vehicles come in many shapes and sizes, are propelled by various propulsion technologies, and carry goods, including people, animals, or objects. These machines are capable of moving goods over long distances, at high speeds, and with larger loads than humans can move. Vehicles were initially driven by humans to control the speed and direction of goods to their destination. Human operation of vehicles has resulted in numerous unfortunate accidents caused by collisions between vehicles, between vehicles and objects, between vehicles and humans, or between vehicles and animals. With advancements in vehicle automation research, various driver assistance systems have been developed and introduced. These include GPS-based navigation guidance, adaptive cruise control, lane change assist, collision avoidance systems, night vision, parking assist, and blind spot detection. Summary of the Invention

[0004] The following summary outlines some aspects of this disclosure to provide a basic understanding of the techniques discussed. This summary is not an exhaustive overview of all the intended features of this disclosure, nor is it intended to identify key or essential elements of all aspects of this disclosure, nor to define the scope of any or all aspects of this disclosure. The sole purpose of this summary is to present, in a general form, some concepts of one or more aspects of this disclosure as a prelude to the more detailed description that follows.

[0005] Distractions can occur to human operators of vehicles, a contributing factor to many vehicle collisions. Driver distraction can include changing the radio, observing events outside the vehicle, and using electronic devices. Sometimes, environmental factors can prevent even attentive drivers from detecting situations in time to prevent a collision. Aspects of this disclosure provide improved systems for enhancing situational awareness to assist drivers in vehicles while driving on roads.

[0006] The example implementation provides a technique for early multimodal channel fusion at the encoder level, which improves object recognition. The technique involves training a single-modal encoder in a fusion pipeline with cross-modal awareness. For example, in some implementations, a camera encoder is trained to determine perspective features based on corresponding perspective views and bird's-eye view (BEV) features from point cloud data for one or more image frames. The point cloud data provides multiple views of the scene represented by the image frames, which improves the camera encoder's feature generation from the image frames. Object recognition is then improved through the improved image frame feature generation.

[0007] More specifically, the technique involves a ranging (e.g., LiDAR, RADAR) encoder determining voxels based on point cloud data and determining perspective features and BEV features of the point cloud data from the voxels. Representative perspective features and BEV features are selected, and the selected features are added as channels of a camera encoder. In at least some embodiments, channel addition is hierarchical, as channel addition can be repeated for each level (e.g., convolutional group) of feature generation by the camera encoder. Each level can be a layer of a model implementing the camera encoder. In such embodiments, the set of features determined by the ranging encoder at a particular level is determined to correspond to the set of features determined by the camera encoder at a particular level. In this way, the perspective features and BEV features corresponding to the particular level of feature generation by the ranging encoder are added as channels to the corresponding level of feature generation by the camera encoder. Perspective features of the image data are then determined. The remaining operations of the pipeline implementing the provided technique are typical operations of a multimodal fusion pipeline for generating fused data.

[0008] In one aspect of this disclosure, an image processing method for use in a vehicle assistance system includes: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on a first feature set in the plurality of first feature sets of the image frame and a second feature set in the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and a fourth feature set of the point cloud data.

[0009] In an additional aspect of this disclosure, an apparatus includes at least one processor and a memory coupled to the at least one processor. The at least one processor is configured to perform operations including: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on a first feature set from the plurality of first feature sets of the image frame and a second feature set from the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and a fourth feature set of the point cloud data.

[0010] In an additional aspect of this disclosure, a non-transitory computer-readable medium stores instructions that, when executed by a processor, cause the processor to perform operations. The operations include: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on the first feature set of the plurality of first feature sets of the image frame and the second feature set of the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and the fourth feature set of the point cloud data.

[0011] In an additional aspect of this disclosure, a vehicle includes: an image sensor; a ranging sensor; a processor; and a memory storing instructions that, when executed by the processor, cause the processor to perform operations including: receiving an image frame representing a scene from the image sensor; receiving point cloud data representing the scene from the ranging sensor; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on a first feature set of the plurality of first feature sets of the image frame and a second feature set of the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and a fourth feature set of the point cloud data.

[0012] The features and technical advantages of the examples according to this disclosure have been summarized rather broadly above in order to better understand the detailed description below. Additional features and advantages will be described below. The disclosed concepts and specific examples can be readily used as the basis for modifying or designing other structures for achieving the same purpose of this disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein (both their organization and operation) and their associated advantages will be better understood from the following description when considered in conjunction with the accompanying drawings. Each figure in the drawings is provided for illustrative and descriptive purposes and not as a definition of limitation of the claims.

[0013] In various specific implementations, the technologies and devices can be used in wireless communication networks such as Code Division Multiple Access (CDMA) networks, Time Division Multiple Access (TDMA) networks, Frequency Division Multiple Access (FDMA) networks, Orthogonal FDMA (OFDMA) networks, Single Carrier FDMA (SC-FDMA) ng networks, LTE networks, GSM networks, fifth-generation (5G) or new radio (NR) networks (sometimes referred to as "5G NR" networks, systems, or devices), and other communication networks. As described herein, the terms "network" and "system" are used interchangeably.

[0014] For example, CDMA networks can implement radio technologies such as Universal Terrestrial Radio Access (UTRA) and CDMA2000. UTRA includes Wideband CDMA (W-CDMA) and Low Chip Rate (LCR). CDMA2000 covers the IS-2000, IS-95, and IS-856 standards.

[0015] For example, TDMA networks can implement radio technologies such as the Global System for Mobile Communications (GSM). The 3rd Generation Partnership Project (3GPP) defines the standard for the GSM EDGE (Enhanced Data Rate GSM Evolution) Radio Access Network (RAN) (also known as GERAN). GERAN is the radio component of the network that connects GSM / EDGE base stations (e.g., Ater and Abis interfaces) and base station controllers (A interface, etc.). The radio access network represents a component of the GSM network through which telephone calls and packet data are routed from the Public Switched Telephone Network (PSTN) and the Internet to subscriber handsets (also known as user terminals or user equipment (UEs)) and from subscriber handsets to the PSTN and the Internet. A mobile phone operator's network may include one or more GERANs, which may be coupled with UTRAN in the case of UMTS / GSM networks. Additionally, the operator's network may also include one or more LTE networks, or one or more other networks. Different network types may use different radio access technologies (RATs) and RANs.

[0016] OFDMA networks can implement radio technologies such as Evolved UTRA (E-UTRA), IEEE 802.11, IEEE 802.16, IEEE 802.20, and flash-OFDM. UTRA, E-UTRA, and GSM are part of the Universal Mobile Telecommunications System (UMTS). Specifically, Long Term Evolution (LTE) is a version of UMTS that uses E-UTRA. UTRA, E-UTRA, GSM, UMTS, and LTE are described in documents provided by an organization called the 3rd Generation Partnership Project (3GPP), and cdma2000 is described in documents from an organization called 3rd Generation Partnership Project 2 (3GPP2). 5G networks include diverse deployments, diverse spectrum, and diverse services and devices that can be achieved using a unified OFDM-based air interface.

[0017] This disclosure may refer to LTE, 4G, or 5G NR technologies to describe certain aspects; however, this description is not intended to be limited to any particular technology or application, and one or more aspects described with reference to one technology may be understood to be applicable to another technology. Additionally, one or more aspects of this disclosure may relate to shared access to radio spectrum between networks using different radio access technologies or radio air interfaces.

[0018] Devices, networks, and systems can be configured to communicate via one or more portions of the electromagnetic spectrum. The electromagnetic spectrum is typically subdivided into various categories, bands, channels, etc., based on frequency or wavelength. In 5G NR, two initial operating bands have been designated as frequency ranges FR1 (410MHz-7.125GHz) and FR2 (24.25GHz-52.6GHz). The frequencies between FR1 and FR2 are generally referred to as mid-band frequencies. Although a portion of FR1 is greater than 6GHz, in various documents and articles, FR1 is often (interchangeably) referred to as the “sub-6GHz” band. A similar naming issue sometimes arises for FR2, which in documents and articles is often (interchangeably) referred to as the “millimeter wave” (mmWave) band, although this is different from the extremely high frequency (EHF) band (30GHz-300GHz) designated as “mmWave” by the International Telecommunication Union (ITU).

[0019] In light of the above, unless otherwise specifically stated, it should be understood that when the term "below 6 GHz" is used herein, it can broadly refer to frequencies that are less than 6 GHz, within FR1, or may include intermediate frequency band frequencies. Furthermore, unless otherwise specifically stated, it should be understood that when the term "mmWave" is used herein, it can broadly refer to frequencies that may include intermediate frequency band frequencies, within FR2, or within the EHF band.

[0020] 5G NR devices, networks, and systems can be implemented using optimized OFDM-based waveform characteristics. These characteristics may include scalable parameter sets and transmission time intervals (TTIs); a general, flexible framework for efficiently multiplexing services and features using dynamic, low-latency Time Division Duplex (TDD) or Frequency Division Duplex (FDD) designs; and advanced radio technologies such as massive MIMO, robust mmWave transmission, advanced channel decoding, and device-centric mobility. The scalability of parameter sets and subcarrier spacing in 5G NR efficiently addresses the operation of various services across different spectrums and deployments. For example, in various outdoor and macro coverage deployments implementing FDD or TDD below 3 GHz, subcarrier spacing may occur at 15 kHz, for example, over bandwidths of 1 MHz, 5 MHz, 10 MHz, 20 MHz, etc. For other various outdoor and small cell coverage deployments with TDD above 3 GHz, subcarrier spacing may occur at 30 kHz over 80 MHz / 100 MHz bandwidths. For various other indoor broadband implementations using TDD on the unlicensed portion of the 5 GHz band, the subcarrier spacing might appear at 60 kHz over a 160 MHz bandwidth. Finally, for various deployments transmitting via mmWave components under TDD at 28 GHz, the subcarrier spacing might appear at 120 kHz over a 500 MHz bandwidth.

[0021] For clarity, certain aspects of the apparatus and technology may be described below with reference to example 5G NR implementations or in a 5G-centric manner, and 5G terminology may be used as illustrative examples in the sections described below; however, this description is not intended to be limited to 5G applications.

[0022] Furthermore, it should be understood that, in operation, wireless communication networks adapted according to the concepts herein may operate using any combination of licensed or unlicensed spectrum, depending on load and availability. Therefore, it will be apparent to those skilled in the art that the systems, apparatuses, and methods described herein can be applied to other communication systems and applications besides the specific examples provided.

[0023] While aspects and implementations are described herein by way of example, those skilled in the art will understand that additional implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and package arrangements. For example, implementations or uses may be via integrated chip implementations or other devices based on non-modular components (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail or purchasing devices, medical devices, AI-enabled devices, etc.). While some examples may or may not specifically point to use cases or applications, the applicability of various types of the described innovations is evident.

[0024] The scope of implementations can range from chip-level or modular components to non-modular, non-chip-level implementations, and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems that incorporate one or more of the described aspects. In some practical settings, devices combining the described aspects and features may also necessary include additional components and features for implementation and practice that are claimed and described. The innovations described herein are expected to be implemented in a wide variety of implementations of different sizes, shapes, and constructions, including both large and small devices, chip-level components, multi-component systems (e.g., radio frequency (RF) chains, communication interfaces, processors), distributed arrangements, end-user equipment, etc.

[0025] In the following description, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of this disclosure. As used herein, the term "coupled" means a direct connection or a connection via one or more intermediate components or circuits. Furthermore, specific terminology is set forth in the following description and for purposes of explanation to provide a thorough understanding of this disclosure. However, it will be apparent to those skilled in the art that these specific details may not be necessary to practice the teachings disclosed herein. In other instances, known circuits and devices are illustrated in block diagram form to avoid obscuring the teachings of this disclosure.

[0026] Certain portions of the following detailed description are presented using other symbolic representations of procedures, logic blocks, processes, and data bit operations within computer memory. In this disclosure, procedures, logic blocks, processes, etc., are conceived as a self-consistent sequence of steps or instructions that produce a desired result. These steps are those that require physical operations on physical quantities. Although not strictly necessary, these physical quantities typically take the form of electrical or magnetic signals that can be stored, transferred, combined, compared, and otherwise manipulated within a computer system.

[0027] In the accompanying drawings, a single block can be described as performing one or more functions. The one or more functions performed by this block can be performed in a single component or across multiple components, and / or can be performed using hardware, software, or a combination of hardware and software. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps are described below in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the system as a whole. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Additionally, the example device may include components other than those shown, including well-known components such as processors, memory, etc.

[0028] Unless otherwise specifically stated, it will be apparent from the following discussion that, throughout this application, the use of terms such as “access,” “receive,” “transmit,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “set,” and “generate” refers to the actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented as physical quantities in the registers, memories, or other such information storage, transmission, or display devices of the computer system.

[0029] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (such as a smartphone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more components that can implement at least some parts of this disclosure. Although the term "device" is used in the following description and examples to describe various aspects of this disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. As used herein, an apparatus can include a device or part of a device for performing the described operations.

[0030] As used herein (including the claims), the term "or" in a list of two or more items means that any one of the listed items may be used alone, or any combination of two or more listed items may be used. For example, if a composition is described as containing component A, B, or C, the composition may contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C.

[0031] Additionally, as used herein (including the claims), the word “or” in a list of items beginning with “at least one of” indicates a separate list such that a list such as “at least one of A, B or C” refers to A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.

[0032] Additionally, as used herein, the term “substantially” is defined as being largely but not necessarily entirely what is specified (and includes what is specified; for example, substantially 90 degrees includes 90 degrees and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any specific implementation of the disclosure, the term “substantially” may be used in place of the specified content within “[percentage]”, where percentage includes 0.1%, 1%, 5%, or 10%.

[0033] Additionally, as used herein, relative terms, unless otherwise specified, can be understood as a quantity relative to a reference. For example, terms such as “higher” or “lower” or “more” or “less” can be understood as a threshold amount higher, lower, more, or less than a reference value. Attached Figure Description

[0034] A further understanding of the nature and advantages of this disclosure can be achieved by referring to the following figures. In the figures, similar components or features may have the same reference numerals. Furthermore, various components of the same type can be distinguished by adding a dash after the reference numeral and a second numeral for differentiation between similar components. If only the first reference numeral is used in the specification, the description applies to any one of the similar components having the same first reference numeral, regardless of the second reference numeral.

[0035] Figure 1 This is a perspective view of a motor vehicle with a driver monitoring system according to an embodiment of this disclosure.

[0036] Figure 2 A block diagram of an example image processing configuration for a vehicle according to one or more aspects of this disclosure is shown.

[0037] Figure 3 It is a block diagram illustrating details of an example wireless communication system based on one or more aspects.

[0038] Figure 4 This is a block diagram illustrating an example multimodal sensor fusion pipeline according to one or more aspects of this disclosure.

[0039] Figure 5 This is a block diagram illustrating an example pipeline for channel fusion of a multimodal encoder according to one or more aspects of this disclosure.

[0040] Figure 6 This is a flowchart illustrating an example method for multimodal sensor fusion according to one or more aspects of this disclosure.

[0041] Similar reference numerals and names in different figures indicate similar elements. Detailed Implementation

[0042] The detailed description set forth below in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to limit the scope of this disclosure. Rather, the detailed description includes specific details for providing a thorough understanding of the subject matter of the invention. It will be apparent to those skilled in the art that these specific details are not necessary in every situation, and in some cases, well-known structures and components are shown in block diagram form for clarity of presentation.

[0043] In typical multimodal fusion systems, such as camera and LiDAR fusion systems, the single-modal encoders for each modality are trained independently without requiring any knowledge of the other modality. Independent training in such multimodal fusion systems leaves room for improvements in the performance of the multimodal fusion system and in systems that depend on the output of the multimodal fusion system (e.g., object detection). This disclosure provides systems, apparatus, methods, and computer-readable media that support techniques for early multimodal channel fusion at the encoder level. The techniques relate to training single-modal encoders in a multimodal fusion pipeline with cross-modal awareness. For example, in this technique, a camera encoder is trained to determine features of one or more image frames representing a scene based on features determined by a ranging encoder (e.g., a LiDAR or RADAR encoder) from point cloud data representing the scene.

[0044] More specifically, the ranging encoder determines voxels based on point cloud data and determines perspective features and BEV features of the point cloud data from the voxels. Representative perspective features and BEV features are selected and added as channels of the camera encoder. In at least some embodiments, channel addition is hierarchical because channel addition can be repeated for each level (e.g., convolutional group) of feature generation by the camera encoder. Each level can be a layer of the model implementing the camera encoder. In such embodiments, the set of features determined by the ranging encoder at a particular level is determined to correspond to the set of features determined by the camera encoder at a particular level. In this way, the perspective features and BEV features corresponding to the particular level of feature generation by the ranging encoder are added as channels as corresponding levels of feature generation by the camera encoder. Perspective features of the image data are then determined. The remaining operations of the pipeline implementing the provided technique are typical operations of a multimodal fusion pipeline for generating fused data.

[0045] Specific embodiments of the subject matter described in this disclosure can be implemented to achieve one or more of the following potential advantages or benefits. In some aspects, this disclosure provides techniques for image processing that can be particularly beneficial in intelligent transportation applications. For example, the proposed techniques improve feature generation from one or more image frames by providing the camera encoder with additional information determined by the ranging encoder. Specifically, point cloud data provides multiple views of a scene represented by one or more image frames and thus provides the camera encoder with information for feature generation (e.g., perspective features and BEV features), which would otherwise not be beneficial to the camera encoder in a typical multimodal fusion system. The resulting fused data is thus improved by improved camera feature generation. The provided techniques further improve the accuracy of downstream perception tasks that utilize such fused data to provide transportation assistance services. Specifically, these techniques enable more accurate tracking of vehicles, pedestrians, obstacles, road signs, road markings, etc.

[0046] One benefit of improved tracking is that it allows vehicle control systems to navigate vehicles more accurately around obstacles. This is particularly useful in situations where unexpected obstacles or road conditions may exist that could pose a danger to the driver. Additionally, improved tracking can help improve overall road safety by reducing vehicle collisions. With better tracking capabilities, vehicles can respond to nearby obstacles more quickly and navigate around detected obstacles more efficiently. These improvements can also be extended to driver assistance systems, which can benefit from increased monitoring capabilities. By expanding the number, type, and variety of detectable surrounding objects, these systems can provide drivers with more accurate warnings and assistance when necessary, without generating unnecessary notifications or distractions.

[0047] Figure 1 This is a perspective view of a motor vehicle with a driver monitoring system according to an embodiment of the present disclosure. The vehicle 100 may include a forward-facing camera 112 mounted in the cabin and visible through the windshield 102. The vehicle may also include a cabin-facing camera 114 mounted in the cabin facing the occupants of the vehicle 100, and particularly the driver of the vehicle 100. Although a set of mounting locations for cameras 112 and 114 is shown for the vehicle 100, other mounting locations may also be used for cameras 112 and 114. For example, one or more cameras may be mounted on one of the driver or passenger B-pillars 126 or one of the driver or passenger C-pillars 128, such as near the top of pillars 126 or 128. Alternatively, one or more cameras may be mounted at the front of the vehicle 100, such as behind the radiator grille 130 or integrated with the bumper 132. Also, one or more cameras may be mounted as part of a driver or passenger side mirror assembly 134.

[0048] Camera 112 may be oriented such that its field of view captures the scene in front of vehicle 100 in the direction in which vehicle 100 is moving when in drive mode or forward. In some embodiments, an additional camera may be located at the rear of vehicle 100 and oriented such that its field of view captures the scene behind vehicle 100 in the direction in which vehicle 100 is moving when in reverse. Although embodiments of this disclosure may be described with reference to a “forward-facing” camera (reference camera 112), aspects of this disclosure may be similarly applied to a “rear-facing” camera facing the reverse direction of vehicle 100. Thus, the benefits obtained when the operator drives vehicle 100 in the forward direction may be similarly obtained when the operator drives vehicle 100 in the reverse direction.

[0049] Furthermore, although embodiments of this disclosure may be described with reference to a “forward-facing” camera (reference camera 112), aspects of this disclosure can be similarly applied to input received from a camera array mounted around vehicle 100 to provide a larger field of view, which may be up to 360 degrees in a direction parallel to the ground and / or up to 360 degrees in a vertical direction perpendicular to the ground. For example, additional cameras may be mounted around the exterior of vehicle 100, such as mounted on or integrated into doors, mounted on or integrated into wheels, mounted on or integrated into bumpers, mounted on or integrated into hoods, and / or mounted on or integrated into the roof.

[0050] Camera 114 can be oriented such that the field of view of camera 114 captures the scene in the cockpit of the vehicle and includes the user operator of the vehicle, and in particular the face of the user operator of the vehicle, with sufficient detail to distinguish the direction of the user operator's gaze.

[0051] Each of cameras 112 and 114 may include one, two, or more image sensors, such as a first image sensor. When multiple image sensors are present, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have a different sensitivity or a different dynamic range than the second image sensor. In one example, the first image sensor may be a wide-angle image sensor, and the second image sensor may be a telephoto image sensor. In another example, the first sensor is configured to acquire an image through a first lens having a first optical axis, and the second sensor is configured to acquire an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification, and the second lens may have a second magnification different from the first magnification. This configuration may occur in a camera module with a lens group, wherein multiple image sensors and associated lenses are located at offset positions within the camera module. Additional image sensors with larger, smaller, or the same field of view may be included.

[0052] Each image sensor may include components for capturing data representing a scene, such as image sensors (including charge-coupled devices (CCDs), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal-oxide-semiconductor (CMOS) sensors) and / or time-of-flight detectors. The device may also include components for focusing and / or converging light onto one or more of the image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components may be controlled to capture a first image frame, a second image frame, and / or more image frames. The image frames may be processed to form a single output image frame (e.g., through a fusion operation), and the output image frame may be further processed according to the aspects described herein.

[0053] As used herein, an image sensor can refer to the image sensor itself and any specific other components coupled to the image sensor for generating image frames for processing by an image signal processor or other logic circuitry, or for storage in memory (whether short-term buffers or long-term non-volatile memory). For example, an image sensor can include other components of a camera, including a shutter, buffers, or other readout circuitry for accessing the individual pixels of the image sensor. An image sensor can also refer to an analog front-end or other circuitry for converting analog signals into a digital representation of an image frame, which is provided to digital circuitry coupled to the image sensor.

[0054] Figure 2A block diagram of an example image processing configuration for a vehicle according to one or more aspects of this disclosure is shown. The vehicle 100 may include or be otherwise coupled to an image signal processor 212 for processing image frames from one or more image sensors, such as a first image sensor 201, a second image sensor 202, and a depth sensor 240. In some specific embodiments, the vehicle 100 may also include or be coupled to a processor (e.g., CPU) 204 and a memory 206 storing instructions 208. The device 100 may also include or be coupled to a display 214 and an input / output (I / O) component 216. The I / O component 216 may be used for user interaction, such as a touchscreen interface and / or physical buttons. The I / O component 216 may also include a network interface for communicating with other devices such as other vehicles, an operator's mobile device, and / or a remote monitoring system. The network interface may include one or more of a wide area network (WAN) adapter 252, a local area network (LAN) adapter 253, and / or a personal area network (PAN) adapter 254. Example WAN adapter 252 is a 4G LTE or 5G NR wireless network adapter. Example LAN adapter 253 is an IEEE 802.11 WiFi wireless network adapter. Example PAN adapter 254 is a Bluetooth wireless network adapter. Each of adapters 252, 253, and / or 254 may be coupled to an antenna comprising multiple antennas configured for main and diversity reception and / or configured to receive a specific frequency band. Vehicle 100 may also include or be coupled to a power source 218, such as a battery or alternator. Vehicle 100 may also include or be coupled to... Figure 2 Additional features or components not shown. In one example, a wireless interface that may include one or more transceivers and an associated baseband processor may be coupled to or included in a WAN adapter 252 for a wireless communication device. In another example, an analog front end (AFE) for converting analog image frame data into digital image frame data may be coupled between image sensors 201 and 202 and image signal processor 212.

[0055] Vehicle 100 may include a sensor hub 250 for interfacing with sensors to receive data on the movement of vehicle 100, data on the environment surrounding vehicle 100, and / or other non-camera sensor data. One example non-camera sensor is a gyroscope, a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example non-camera sensor is an accelerometer, a device configured to measure acceleration, which can also be used to determine velocity and distance traveled by appropriately integrating the measured acceleration, and one or more of acceleration, velocity, and / or distance may be included in the generated motion data. In other examples, the non-camera sensor may be a Global Positioning System (GPS) receiver, a LiDAR system, a RADAR system, or another ranging system. For example, sensor hub 250 may be connected to a vehicle bus for transmitting configuration commands and / or receiving information from vehicle sensors 272, such as distance (e.g., ranging) sensors or vehicle-to-vehicle (V2V) sensors (e.g., sensors for receiving information from nearby vehicles).

[0056] Image signal processor (ISP) 212 can receive image data, such as data used to form image frames. In one embodiment, a local bus connection couples the image signal processor 212 to a local bus connection that corresponds to... Figure 1 The first camera 203 of camera 112 and the camera that can correspond to Figure 1 The second camera 205 of camera 114 contains image sensors 201 and 202. In another embodiment, a wired interface may couple image signal processor 212 to an external image sensor. In yet another embodiment, a wireless interface may couple image signal processor 212 to image sensors 201 and 202.

[0057] The first camera 203 may include a first image sensor 201 and a corresponding first lens 231. The second camera 205 may include a second image sensor 202 and a corresponding second lens 232. Each of lenses 231 and 232 may be controlled by an associated autofocus (AF) algorithm 233 executed in the ISP 212, which adjusts lenses 231 and 232 to focus on a specific focal plane at a certain scene depth from image sensors 201 and 202. The AF algorithm 233 may be assisted by a depth sensor 240. In some embodiments, lenses 231 and 232 may have a fixed focal length.

[0058] First image sensor 201 and second image sensor 202 are configured to capture one or more image frames. Lenses 231 and 232 focus light onto image sensors 201 and 202 respectively via one or more apertures for receiving light, one or more shutters for blocking light outside the exposure window, one or more color filter arrays (CFAs) for filtering light outside a specific frequency range, one or more analog front ends for converting analog measurements into digital information, and / or other suitable components for imaging.

[0059] In some embodiments, the image signal processor 212 may execute instructions from memory, such as instructions 208 from memory 206, instructions stored in a separate memory coupled to or included in the image signal processor 212, or instructions provided by processor 204. Additionally or alternatively, the image signal processor 212 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in this disclosure. For example, the image signal processor 212 may include one or more image front-ends (IFEs) 235, one or more image post-processing engines (IPEs) 236, and one or more automatic exposure compensation (AEC) engines 234. AF 233, AEC 234, IFE 235, and IPE 236 may each include dedicated circuitry embodied as software code executed by the ISP 212 and / or a combination of hardware within the ISP 212 and software code executed on the ISP.

[0060] In some embodiments, memory 206 may include a nontransitory or nontransitory computer-readable medium storing computer-executable instructions 208 to perform all or part of one or more of the operations described in this disclosure. In some embodiments, instructions 208 include a camera application (or other suitable application) to generate images or videos for execution during operation of vehicle 100. Instructions 208 may also include other applications or programs, such as operating systems, mapping applications, or entertainment applications, for execution by vehicle 100. For example, execution of a camera application by processor 204 may enable vehicle 100 to generate images using image sensors 201 and 202 and image signal processor 212. Memory 206 may also be accessed by image signal processor 212 to store processed frames or may be accessed by processor 204 to obtain processed frames. In some embodiments, vehicle 100 includes a system-on-a-chip (SoC) that integrates image signal processor 212, processor 204, sensor hub 250, memory 206, and input / output components 216 into a single package.

[0061] In some embodiments, at least one of the image signal processor 212 or processor 204 executes instructions to perform various operations described herein, including object detection, risk map generation, driver monitoring, and driver alert operations. For example, the execution of instructions may instruct the image signal processor 212 to begin or end the capture of image frames or sequences of image frames. In some embodiments, processor 204 may include one or more general-purpose processor cores 204A capable of executing scripts or instructions (such as instructions 208 stored in memory 206) of one or more software programs. For example, processor 204 may include one or more application processors configured to execute a camera application (or other suitable application for generating images or videos) stored in memory 206.

[0062] When executing a camera application, processor 204 may be configured to instruct image signal processor 212 to perform one or more operations in reference to image sensor 201 or 202. For example, the camera application may receive a command to start a video preview display, and upon receiving the command, capture and process a sequence of video including image frames from one or more image sensors 201 or 202, and display the video on an information display on display 114 in the cockpit of vehicle 100.

[0063] In some embodiments, in addition to the ability to execute software to enable vehicle 100 to perform multiple functions or operations (such as those described herein), processor 204 may also include an IC or other hardware (e.g., an artificial intelligence (AI) engine 224). In some other embodiments, vehicle 100 does not include processor 204, such as when all the described functionalities are configured in image signal processor 212.

[0064] In some embodiments, display 214 may include one or more suitable displays or screens that allow user interaction and / or present items (such as previews of image frames captured by image sensors 201 and 202) to the user. In some embodiments, display 214 is a touch-sensitive display. I / O component 216 may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via display 214. For example, I / O component 216 may include (but is not limited to) a graphical user interface (GUI), keyboard, mouse, microphone, speaker, squeezable bezel, one or more buttons (such as a power button), slider, switch, etc. In some embodiments involving autonomous driving, I / O component 216 may include an interface to a vehicle bus for providing commands and information to and receiving information from a vehicle system 270, which includes a propulsion system (e.g., commands for increasing or decreasing speed or applying braking) and a steering system (e.g., commands for turning wheels, changing route, or changing final destination). According to embodiments of this disclosure, the accuracy of commands output to the vehicle system 270 can be improved by enabling a multimodal sensor fusion pipeline with cross-modal perception at the encoder level. Cross-modal perception improves the generation of camera features from image data, which improves the accuracy of fused data determined in part based on camera features, and thus improves object recognition performance that can affect commands transmitted to the vehicle system 270.

[0065] Although shown as coupled to each other via processor 204, components such as processor 204, memory 206, image signal processor 212, display 214, and I / O components 216 may be coupled to each other in various other arrangements, such as via one or more local buses, which are not shown for simplicity. While image signal processor 212 is illustrated as separate from processor 204, image signal processor 212 may be the core of processor 204, which is an application processor unit (APU) included in a system-on-a-chip (SoC), or otherwise included in processor 204. Although vehicle 100 is mentioned in the examples herein for the purpose of encompassing aspects of this disclosure, some device components may not be included. Figure 2 The details are shown to prevent obscuring aspects of this disclosure. Additionally, other components, the number of components, or combinations of components may be included in a suitable vehicle for performing aspects of this disclosure. Therefore, this disclosure is not limited to the configuration of specific equipment or components, but includes vehicle 100.

[0066] The vehicle 100 can communicate as a user equipment (UE) within the wireless network 300, such as via WAN adapter 252, etc. Figure 3 As shown. Figure 3 This is a block diagram illustrating details of an example wireless communication system according to one or more aspects. Wireless network 300 may, for example, include a 5G wireless network. As those skilled in the art will recognize, Figure 3 The components appearing in this network are likely to have corresponding components in other network arrangements (including, for example, cellular network arrangements and non-cellular network arrangements (e.g., device-to-device, peer-to-peer, or self-organizing network arrangements)).

[0067] Figure 3 The illustrated wireless network 300 includes base station 305 and other network entities. A base station can be a station communicating with a UE and may also be referred to as an evolved Node B (eNB), a next-generation eNB (gNB), and an access point, etc. Each base station 305 may provide communication coverage for a specific geographic area. In 3GPP, the term "cell" may refer to a specific geographic coverage area of ​​a base station or a base station subsystem serving that coverage area, depending on the context in which the term is used. In the specific implementation of the wireless network 300 herein, base station 305 may be associated with the same operator or different operators (e.g., the wireless network 300 may include multiple operator wireless networks). Additionally, in the specific implementation of the wireless network 300 herein, base station 305 may use one or more frequencies (e.g., one or more bands of licensed spectrum, unlicensed spectrum, or combinations thereof) from the same frequencies as neighboring cells to provide wireless communication. In some examples, a single base station 305 or UE 315 may be operated by more than one network operating entity. In some other examples, each base station 305 and UE 315 may be operated by a single network operating entity.

[0068] Base stations can provide communication coverage for macro cells, small cells (such as pico cells or femto cells), or other types of cells. Macro cells typically cover a relatively large geographic area (e.g., a radius of several kilometers) and allow unrestricted access by UEs with service subscriptions to a network provider. Small cells (such as pico cells) typically cover a relatively small geographic area and allow unrestricted access by UEs with service subscriptions to a network provider. Small cells (such as femto cells) also typically cover a relatively small geographic area (e.g., a home) and, in addition to unrestricted access, provide restricted access by UEs associated with the femto cell (e.g., UEs in a Closed Subscriber Group (CSG), UEs of users in a home, etc.). A base station for a macro cell may be referred to as a macro base station. A base station for a small cell may be referred to as a small cell base station, pico base station, femto base station, or home base station. Figure 3In the example shown, base stations 305d and 305e are conventional macro base stations, while base stations 305a-305c are macro base stations implemented using one of three-dimensional (3D), full-dimensional (FD), or massive MIMO. Base stations 305a to 305c utilize their higher-dimensional MIMO capabilities to employ 3D beamforming in elevation and azimuth beamforming to increase coverage and capacity. Base station 305f is a small cell base station, which can be a home node or a portable access point. A base station can support one or more (e.g., two, three, and four cells, etc.) cells.

[0069] Wireless Network 300 can support synchronous or asynchronous operation. For synchronous operation, base stations can have similar frame timings, and transmissions from different base stations can be roughly aligned in time. For asynchronous operation, base stations can have different frame timings, and transmissions from different base stations may not be aligned in time. In some scenarios, the network can be enabled or configured to handle dynamic switching between synchronous and asynchronous operations.

[0070] UE 315 is distributed throughout the wireless network 300, and each UE may be stationary or mobile. It should be understood that although mobile devices are generally referred to as UEs in the standards and specifications issued by 3GPP, such devices may additionally or otherwise be referred to by those skilled in the art as mobile station (MS), subscriber station, mobile unit, subscriber unit, radio unit, remote unit, mobile device, radio device, wireless communication device, remote device, mobile subscriber station, access terminal (AT), mobile terminal, radio terminal, remote terminal, mobile phone, terminal, user agent, mobile client, client, gaming device, augmented reality device, vehicle component, vehicle equipment or vehicle module, or some other suitable term.

[0071] Non-limiting examples of mobile devices include specific implementations of one or more of UE 315, including mobile phones, cellular phones, smartphones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, laptop computers, personal computers (PCs), notebooks, netbooks, smartbooks, tablets, personal digital assistants (PDAs), and vehicles. While UEs 315a to 305j are specifically shown as vehicles, vehicles may employ the communication configurations described with reference to any of UEs 315a to 315k.

[0072] In one respect, a UE can be a device that includes a Universal Integrated Circuit Card (UICC). In another respect, a UE can be a device that does not include a UICC. In some respects, a UE that does not include a UICC may also be referred to as an IoE device. Figure 3The illustrated UEs 315a to 315d are examples of mobile smartphone-type devices accessing the wireless network 300. The UE can also be a machine specifically configured for connected communications, including machine-type communications (MTC), enhanced MTC (eMTC), and narrowband IoT (NB-IoT). Figure 3 The illustrated UEs 315e to 315k are examples of various machines configured for communication that access the wireless network 300.

[0073] Mobile devices (such as UE 315) can communicate with any type of base station (whether macro base station, pico base station, femto base station, or relay station). Figure 3 In this context, a communication link (represented by a lightning bolt) indicates radio transmission or expected transmission between the UE and a serving base station (which is designated to serve the UE on the downlink or uplink) and backhaul transmission between base stations. The UE may operate as a base station or other network node in some scenarios. Backhaul communication between base stations of the wireless network 300 can be performed using wired or wireless communication links.

[0074] In operation, at wireless network 300, base stations 305a to 305c use 3D beamforming and cooperative spatial technologies such as Cooperative Multipoint (CoMP) or Multi-Connection to serve UEs 315a and 315b. Macro base station 305d performs backhaul communication with base stations 305a to 305c and the small cell (base station 305f). Macro base station 305d also transmits multicast services subscribed to and received by UEs 315c and 315d. Such multicast services may include mobile TV or streaming video, or may include other services for providing community information, such as weather emergencies or alerts, such as Amber Alerts or Grey Alerts.

[0075] The specific implementation of the wireless network 300 supports communication with highly reliable and redundant links for such devices. Redundant communication links with UE 315e include links from macro base stations 305d and 305e, and small cell base station 305f. Other machine-type devices, such as UE 315f (thermometer), UE 315g (smart meter), and UE 315h (wearable device), can communicate directly with base stations such as small cell base station 305f and macro base station 305e via the wireless network 300, or in a multi-hop configuration by communicating with another user equipment relaying its information to the network, such as UE 315f relaying temperature measurement information to smart meter UE 315g, which then reports it to the network via small cell base station 305f. The wireless network 300 can also provide additional network efficiency through dynamic, low-latency TDD or low-latency FDD communication, such as in vehicle-to-vehicle (V2V) mesh networks between UEs 315i to 315k communicating with macro base station 305e.

[0076] refer to Figure 1 , Figure 2 and Figure 3 The aspects of the transportation system described and illustrated herein may include techniques for early multimodal channel fusion at the encoder level. These techniques involve training a single-modal encoder in a multimodal fusion pipeline with cross-modal awareness. For example, in this technique, a camera encoder is trained to determine features of one or more image frames representing a scene based on features determined by a ranging encoder (e.g., a LiDAR or RADAR encoder) from point cloud data representing the scene.

[0077] Figure 4 This is a block diagram illustrating a pipeline 400 for multimodal sensor fusion that implements the provided technique. Pipeline 400 can be constructed from the above-described... Figure 2 and Figure 3The pipeline 400 is implemented using one or more components. In an illustrated embodiment, the pipeline 400 includes receiving one or more image frames 402 from an image sensor (e.g., a first image sensor 201 or a second image sensor 202) of a camera (e.g., a first camera 203 or a second camera 205). An encoder 404 determines perspective (PV) features of the one or more image frames 402 at multiple levels (e.g., groups of convolutions). Each level has different levels of specificity (e.g., granularity) of the features of the one or more image frames 402. The encoder 404 can be implemented as one or more machine learning models, including supervised learning models, unsupervised learning models, other types of machine learning models, and / or other types of predictive models. For example, the model implementing the encoder 404 can be implemented as one or more of neural networks, transformer models, decision tree models, support vector machines, Bayesian networks, classifier models, regression models, etc. Each level of the encoder 404 can be a different layer of the model implementing the encoder 404.

[0078] The pipeline 400 also includes receiving point cloud data 406 from a ranging sensor (e.g., depth sensor 240). For example, point cloud data 406 can be received from a LiDAR sensor or a RADAR sensor. Encoder 408 generates voxels 409 from the point cloud data 406 at multiple levels (e.g., groups of convolutions). Each level has a different level of specificity (e.g., granularity) of the features of the point cloud data 406. As known to those skilled in the art, a voxel is a three-dimensional volume unit in a grid-based spatial representation. Similar to how pixels represent a single point in a two-dimensional image, a voxel represents a point in three-dimensional space. At each level in the encoder 408, a set of PV features and BEV features of the point cloud data 406 are determined from the voxels 409. For example, one dimension of the voxels 409 can be flattened to determine various perspective views, and if the view is a top view, a BEV view is determined. In this example, determining the PV features and BEV features of the point cloud data 406 may involve global max pooling. The encoder 408 can be implemented as one or more machine learning models capable of determining voxels from point cloud data 406, including supervised learning models, unsupervised learning models, other types of machine learning models, and / or other types of prediction models. For example, the model implementing encoder 408 can be implemented as one or more of neural networks, transformer models, decision tree models, support vector machines, Bayesian networks, classifier models, regression models, etc. Each level of encoder 408 can be a different layer of the model implementing encoder 408.

[0079] The provided technique for early multimodal channel fusion at the encoder level is implemented via model 410 and encoder 404. Model 410 is trained to determine which set of PV features and BEV features of point cloud data 406 corresponds to which specific level of encoder 404. Model 410 is further trained to select one or more pairs of PV features and BEV features from the determined set of PV features and BEV features to serve as additional channel inputs to encoder 404 for the corresponding level. This process can be repeated such that one or more pairs of PV features and BEV features are served as additional channel inputs to encoder 404 for each level of encoder 404. The PV features of one or more image frames 402 determined by encoder 404 are fused (e.g., cascaded) with one or more pairs of PV features and BEV features of point cloud data 406 at each level via fusion 405, such that encoder 404 generates improved PV features 412 compared to a typical multimodal sensor fusion system (e.g., compared to PV features determined by encoder 404 before considering the features of point cloud data 406). The improvement to PV feature 412 is based on modifying the conventionally determined PV features of the image frame by considering the corresponding features of point cloud data 406. The following section combines... Figure 5 Further details are provided regarding cross-modal sensing of encoders 404 and 408, and the operation of model 410.

[0080] Model 410 can be implemented as one or more machine learning models, including supervised learning models, unsupervised learning models, other types of machine learning models, and / or other types of prediction models. For example, model 410 can be implemented as one or more of neural networks, transformer models, decision tree models, support vector machines, Bayesian networks, classifier models, regression models, etc. Model 410 can be trained based on training data to input one or more pairs of PV features and BEV features of point cloud data 406 as additional channels into encoder 404. For example, one or more training data sets containing a set of PV features and BEV features of point cloud data and a set of PV features of image data can be used. The training data sets can specify one or more desired outputs. For example, the desired output could be a set of PV features of image data, with the set of PV features and BEV features of point cloud data expected to be associated with that set of PV features of image data. In another example, the desired output could be one or more pairs of PV features and BEV features from the set of PV features and BEV features of point cloud data expected to be added to encoder 404. The parameters of model 410 can be updated based on whether model 410 generates the correct output when compared with the expected output. Specifically, model 410 can receive one or more input data from a training dataset associated with multiple expected outputs. Model 410 can generate a predicted output based on the current configuration of model 410. The predicted output can be compared with the expected output, and one or more parameter updates can be computed based on the difference between the predicted output and the expected output. Specifically, parameters can include weights (e.g., priorities) of different features and feature combinations (e.g., PV features or BEV features of point cloud data). Parameter updating of model 410 can include updating one or more features among the analyzed features and / or weights assigned to different features or feature combinations (e.g., relative to the current configuration of model 410). Alternatively, model 410 can be a layer of a multimodal sensor fusion model.

[0081] The remaining boxes of pipeline 400 are typical boxes of a multimodal sensor fusion pipeline known to those skilled in the art. For example, at circle 414, PV features 412 are projected onto the BEV plane to determine BEV features 416 of one or more image frames 402. Continuing the example, encoder 408 determines 3D sparse features 418 of point cloud data 406, and at circle 420, the 3D sparse features 418 are flattened to determine BEV features 422 of point cloud data 406. Fusion 424 occurs such that BEV features 416 and BEV features 422 are fused (e.g., cascaded) to generate fused data 426. Fusion 405 and fusion 424 are both included in pipeline 400, allowing fusion at different granularities. In some embodiments, at box 428, fused data 426 may be input into a 3D object detection pipeline for object detection. In other embodiments, fused data 426 may be used for purposes other than object detection.

[0082] Figure 5 This is a block diagram of an example pipeline 500 for multimodal encoder channel fusion. Encoder 404 includes stages 502, 504, 506, 508, and 510. In at least some aspects, each of stages 502, 504, 506, 508, and 510 can be a layer of a model implementing encoder 504. Each of stages 502, 504, 506, 508, and 510 is associated with a set of PV features of one or more image frames 402. Encoder 408 includes stages 512, 514, 516, 518, and 520. In at least some aspects, each of stages 512, 514, 516, 518, and 520 can be a layer of a model implementing encoder 408. While example pipeline 500 includes an equal number of channels for encoder 404 and encoder 408, in other examples, encoder 404 may have a different number of channels than encoder 408. Each of levels 512, 514, 516, 518, and 520 is associated with the PV feature set and BEV feature set of point cloud data 406. For example, Figure 5 A PV feature set 524 and a BEV feature set 526, both associated with level 516, are depicted. The PV feature set 524 includes PV features 524A, 524B, 524C, 524D, and 524E. The BEV feature set 526 includes BEV features 526A, 526B, 526C, 526D, and 526E. In at least some respects, arrow 522 indicates that global max pooling was used to determine the PV feature set 524 and the BEV feature set 526.

[0083] Model 410 is trained to determine which of the levels 502, 504, 506, 508, and 510 of encoder 404 corresponds to the PV feature set 524 and the BEV feature set 526. In one example, model 410 makes this determination via statistical analysis. More specifically, in such an example, PV features 524A, 524B, 524C, 524D, and 524E of the PV feature set 524 are input into model 410, and model 410 determines the mean and / or variance of PV features 524A, 524B, 524C, 524D, and 524E. Alternatively, the mean and / or variance of PV features 524A, 524B, 524C, 524D, and 524E may be determined separately from model 410 and then input into model 410. Continuing the example, the PV feature set for each of levels 502, 504, 506, 508, and 510 is input into model 410, and model 410 determines the mean and / or variance of the PV feature set for each of levels 502, 504, 506, 508, and 510. Alternatively, the mean and / or variance of the PV feature set for each of levels 502, 504, 506, 508, and 510 can be determined separately from model 410 and then input into model 410. In this example, model 410 is trained to determine levels 502, 504, 506, 508, or 510 that have the most similar mean and / or variance to the mean and / or variance of PV features 524A, 524B, 524C, 524D, and 524E. Figure 5 The illustrated example shows that the mean and / or variance of PV features 524A, 524B, 524C, 524D, and 524E are most similar to the mean and / or variance of the PV features of level 506, such that the PV feature set 524 and the associated BEV feature set 526 correspond to level 506.

[0084] At box 528, model 410, trained with attention weights, determines which pair of features from PV feature set 524 and BEV feature set 526 is closest to the feature set of level 506. In the example, group channel normalization in encoder 404 is used to facilitate feature learning in groups with similar properties. A pair includes PV feature 530 and BEV feature 532. For example, model 410 can determine that a pair including PV feature 524B and BEV feature 526B is most closely related to the feature set of level 506. In such an example, PV feature 530 is PV feature 524B, and BEV feature 532 is BEV feature 526B. PV feature 530 and BEV feature 532 are each added as additional channels of encoder 404. Pipeline 500 also includes determining a pair of PV features 530 and BEV features 532 to be added to each of the stages 502, 504, 508, and 510 of encoder 404, based on the PV feature sets and BEV feature sets of each stage 512, 514, 518, and 520 of encoder 408. In this way, the hierarchical features of the image data are enhanced by corresponding hierarchical 3D information from the ranging sensor. Although the PV feature set 524 and BEV feature set 526 of stage 516 are shown in the illustrated example to each include five features, the PV feature sets and BEV feature sets of stages 512, 514, 518, and 520 can each include any suitable number of features. In this way, encoder 404 and encoder 408 perceive each other across modalities by adding channels to encoder 404 based on the features determined by encoder 404 and the features determined by encoder 408.

[0085] Figure 6 The image processing method described above is illustrated. Figure 6 This is a flowchart illustrating an example method 600 for image processing used in a vehicle assistance system. Method 600 includes receiving image frames (e.g., one or more image frames 402) representing a scene at block 602. One or more image frames 402 may be received from an image sensor (e.g., a first image sensor 201) of a camera (e.g., a first camera 203). At block 604, point cloud data (e.g., point cloud data 406) representing the scene is received. In various respects, point cloud data 406 may be received from a ranging sensor (e.g., a depth sensor 240), such as a LiDAR sensor or a RADAR sensor.

[0086] At box 606, multiple first feature sets of one or more image frames 402 are determined (e.g., feature sets for each of levels 502, 504, 506, 508, and 510). At box 608, multiple second feature sets of point cloud data 406 are determined based on multiple voxels representing point cloud data 406 (e.g., voxel 409). For example, the second feature set of level 516 includes PV features 524A, 524B, 524C, 524D, and 524E and BEV features 526A, 526B, 526C, 526D, and 526E. In various aspects, the multiple second feature sets are each determined from voxel 409 using global max pooling. In various aspects, the second feature sets include multiple pairs of PV features and BEV features of point cloud data 406. For example, the second feature set of level 516 includes five pairs: PV feature 524A and BEV feature 526A, PV feature 524B and BEV feature 526B, PV feature 524C and BEV feature 526C, PV feature 524D and BEV feature 526D, and PV feature 524E and BEV feature 526E.

[0087] In various aspects, each of the plurality of first feature sets of one or more image frames 402 corresponds to a corresponding level in a plurality of first levels (e.g., levels 502, 504, 506, 508, and 510) of a first encoder (e.g., encoder 404), and each of the plurality of second feature sets of point cloud data 406 corresponds to a corresponding level in a plurality of second levels (e.g., levels 512, 514, 516, 518, and 520) of a second encoder (e.g., encoder 408). In such aspects, the corresponding level (e.g., level 506) corresponding to the feature set of level 506 associated with one or more image frames 402 corresponds to the corresponding level (e.g., level 516) corresponding to the PV feature set 524 and the BEV feature set 526.

[0088] At box 610, a third feature set (e.g., PV feature 412) of one or more image frames 402 is determined based on a first feature set (e.g., PV features of level 506) from a plurality of first feature sets of one or more image frames 402 and a second feature set (e.g., PV feature 524B and BEV feature 526B) from a plurality of second feature sets of point cloud data 406. In various aspects, the first and second feature sets are determined based on a first statistical indicator (e.g., mean, variance) associated with the plurality of first feature sets of one or more image frames 402 and a second statistical indicator (e.g., mean, variance) associated with the plurality of second feature sets of point cloud data 406.

[0089] In various aspects, the PV feature 412 of one or more image frames 402 is determined by an encoder (e.g., encoder 404). In various examples of such aspects, method 600 also includes adding a perspective feature (e.g., PV feature 524B) from one of a plurality of pairs as a first channel of encoder 404, and adding a BEV feature (e.g., BEV feature 526B) from that pair as a second channel of encoder 404.

[0090] At box 612, fused data (e.g., fused data 426) is output, which combines PV features 412 of one or more image frames 402 with a second set of features (e.g., BEV features 422) of point cloud data 406. In some aspects, method 600 also includes detecting objects represented in one or more image frames 402 based on fused data 426. In such aspects, method 600 may also include functionality to control a vehicle (e.g., vehicle 100) based on the detected objects. For example, steering or braking of the vehicle may be controlled to avoid collisions with objects.

[0091] Note that this is for reference only. Figures 4 to 6 One or more boxes (or operations) described may be combined with one or more boxes (or operations) described in another drawing with reference to the accompanying drawings. For example, Figures 4 to 6 One or more boxes (or operations) can be connected with Figures 1 to 3 Combine one or more boxes (or operations). For example, with... Figure 5 One or more associated boxes can be connected with and Figure 4 Group one or more related boxes together.

[0092] In one or more aspects, the technology for supporting vehicle operation may include additional aspects, such as any single aspect or any combination of aspects described below or in conjunction with one or more other processes or devices described elsewhere herein. In a first aspect, the apparatus is configured to perform operations including: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on the first feature set of the plurality of first feature sets of the image frame and the second feature set of the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and the fourth feature set of the point cloud data. In some embodiments, the apparatus includes a wireless device, such as a UE. In some embodiments, the apparatus may include at least one processor and memory coupled to the processor. The processor may be configured to perform the operations described herein with respect to the apparatus. In some other embodiments, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some embodiments, the apparatus may include one or more components configured to perform the operations described herein. In some embodiments, the method of wireless communication may include one or more operations described herein with reference to the apparatus.

[0093] In a second aspect, in conjunction with the first aspect, the apparatus is further configured to determine the first feature set and the second feature set based on a first statistical indicator associated with a plurality of first feature sets of an image frame and a second statistical indicator associated with a plurality of second feature sets of point cloud data.

[0094] In a third aspect, in combination with one or more of the first or second aspects, each of the plurality of first feature sets of the image frame corresponds to a corresponding level in the plurality of first levels of the first encoder, and each of the plurality of second feature sets of the point cloud data corresponds to a corresponding level in the plurality of second levels of the second encoder.

[0095] In the fourth aspect, combined with the third aspect, the corresponding level corresponding to the first feature set of the image frame corresponds to the corresponding level corresponding to the second feature set of the point cloud data.

[0096] In the fifth aspect, in combination with one or more of the first to fourth aspects, the second feature set includes multiple pairs of perspective features of point cloud data and BEV features of point cloud data.

[0097] In the sixth aspect, in conjunction with the fifth aspect, the third set of features of the image frame is determined by the encoder, and the apparatus is further configured to add perspective features from one of a plurality of pairs as the first channel of the encoder, and add BEV features from that pair as the second channel of the encoder.

[0098] In the seventh aspect, combined with the fifth aspect, multiple perspective features and multiple BEV features are each determined from multiple voxels using global max pooling.

[0099] In the eighth aspect, in combination with one or more of the first to seventh aspects, the point cloud data is received from a ranging sensor.

[0100] In the ninth aspect, in combination with one or more of the first to eighth aspects, the device is also configured to detect objects represented in an image frame based on fused data.

[0101] In the tenth aspect, in conjunction with the ninth aspect, the device is also configured to control the function of the vehicle based on the detected objects.

[0102] In an eleventh aspect, a means of transportation includes: an image sensor; a ranging sensor; a processor; and a memory storing instructions that, when executed by the processor, cause the processor to perform operations including: receiving an image frame representing a scene from the image sensor; receiving point cloud data representing the scene from the ranging sensor; determining a plurality of first feature sets of the image frame; determining a plurality of second feature sets of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third feature set of the image frame based on a first feature set of the plurality of first feature sets of the image frame and a second feature set of the plurality of second feature sets of the point cloud data; and outputting fused data that combines the third feature set of the image frame and a fourth feature set of the point cloud data.

[0103] In the twelfth aspect, in conjunction with the eleventh aspect, each of the plurality of first feature sets of the image frame corresponds to a corresponding level in the plurality of first levels of the first encoder, and each of the plurality of second feature sets of the point cloud data corresponds to a corresponding level in the plurality of second levels of the second encoder. In the twelfth aspect, the corresponding level corresponding to the first feature set of the image frame corresponds to the corresponding level corresponding to the second feature set of the point cloud data.

[0104] In the thirteenth aspect, in combination with one or more of the eleventh to twelfth aspects, the second feature set includes multiple pairs of perspective features of point cloud data and BEV features of point cloud data, the third feature set of the image frame is determined by the encoder, and the operation further includes: adding perspective features from one of the multiple pairs as a first channel of the encoder; and adding BEV features from that pair as a second channel of the encoder.

[0105] In the fourteenth aspect, in combination with one or more of aspects eleven to thirteen, the operation further includes: detecting objects represented in an image frame based on fused data; and controlling the functions of a vehicle based on the detected objects.

[0106] This article is about Figures 1 to 4 The components, functional blocks, and modules described include processors, electronic devices, hardware devices, electronic components, logic circuits, memory, software code, firmware code, and any combination thereof. Software should be interpreted broadly as instructions, instruction sets, code, code segments, program code, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, and / or functions, regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description languages, or other terms. Furthermore, the features discussed herein can be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.

[0107] Those skilled in the art will further understand that the various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such specific implementation decisions should not be construed as departing from the scope of this disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein is merely illustrative, and that components, methods, or interactions of various aspects of this disclosure can be combined or performed in ways other than those illustrated and described herein.

[0108] The various exemplary logics, logic blocks, modules, circuits, and algorithmic processes described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. The interchangeability of hardware and software has been broadly described in terms of functionality and illustrated in the various exemplary components, blocks, modules, circuits, and processes described above. Whether such functionality is implemented in hardware or software depends on the specific application and the design constraints imposed on the overall system.

[0109] Hardware and data processing means for implementing the various exemplary logic, logic blocks, modules, and circuits described herein can be implemented or executed using general-purpose single-chip or multi-chip processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some embodiments, the processor may be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In some embodiments, specific processes and methods may be performed by circuitry specific to a given function.

[0110] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuits, computer software, firmware, including the structures disclosed in this specification and their structural equivalents or any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus.

[0111] If implemented in software, the functionality can be stored as one or more instructions or code on or transmitted through a computer-readable medium. The processes of the methods or algorithms disclosed herein can be implemented in a processor-executable software module that can reside on a computer-readable medium. Computer-readable media include both computer storage media and communication media, including any medium that can be implemented to transfer a computer program from one location to another. Storage media can be any available medium accessible to a computer. By way of example and not limitation, such computer-readable media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible to a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. As used herein, disks and optical discs include compact optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operation of a method or algorithm may reside as a set of code and instructions or any combination of code and instructions on a machine-readable medium and a computer-readable medium that may be incorporated into a computer program product.

[0112] Various modifications to the specific embodiments described in this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to some other specific embodiments without departing from the spirit or scope of this disclosure. Therefore, the claims are not intended to be limited to the specific embodiments shown herein, but are to be accorded the widest scope consistent with this disclosure, the principles disclosed herein, and the novel features thereof.

[0113] Certain features described in this specification in the context of a single embodiment may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as operating in certain combinations and even originally claimed in this way, one or more features from the claimed combination may be removed from that combination in some cases, and the claimed combination may be for sub-combinations or variations thereof.

[0114] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the indicated specific order or sequential order, or to perform all illustrated operations to achieve the desired result. Furthermore, the drawings may schematically depict one or more example processes in the form of flowcharts. However, other operations not depicted may be combined with the schematically illustrated example processes. For example, one or more additional operations may be performed before, after, simultaneously with, or between any illustrated operation. In some contexts, multitasking and parallel processing are advantageous. Moreover, the separation of the various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result.

[0115] The prior description of this disclosure is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples and designs described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for image processing for use in a vehicle assist system, the method comprising: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first sets of features of the image frame; determining a plurality of second sets of features of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third set of features of the image frame based on a first set of features of the plurality of first sets of features of the image frame and a second set of features of the plurality of second sets of features of the point cloud data; and outputting fusion data combining the third set of features of the image frame and a fourth set of features of the point cloud data.

2. The method of claim 1, the method further comprising: determining the first set of features and the second set of features based on first statistical indicators associated with the plurality of first sets of features of the image frame and second statistical indicators associated with the plurality of second sets of features of the point cloud data.

3. The method of claim 1, wherein each first set of features of the plurality of first sets of features of the image frame corresponds to a respective level of a plurality of first levels of a first encoder, and wherein each second set of features of the plurality of second sets of features of the point cloud data corresponds to a respective level of a plurality of second levels of a second encoder.

4. The method of claim 3, wherein the respective level corresponding to the first set of features of the image frame corresponds to the respective level corresponding to the second set of features of the point cloud data.

5. The method of claim 1, wherein the second set of features includes pairs of perspective view features of the point cloud data and BEV features of the point cloud data.

6. The method of claim 5, wherein the third set of features of the image frame is determined by an encoder, the method further comprising: adding a perspective view feature in a pair of the pairs as a first channel of the encoder; and adding a BEV feature in the pair as a second channel of the encoder.

7. The method of claim 5, wherein the plurality of perspective view features and the plurality of BEV features are each determined from the plurality of voxels using global max pooling.

8. The method of claim 1, wherein the point cloud data is received from a ranging sensor.

9. The method of claim 1, the method further comprising detecting an object represented in the image frame based on the fusion data.

10. The method of claim 9, the method further comprising controlling a function of a vehicle based on the detected object.

11. An apparatus, the apparatus comprising: a memory storing processor-readable code; and at least one processor coupled to the memory, the at least one processor configured to execute the processor-readable code to cause the at least one processor to perform operations comprising: receiving an image frame representing a scene; receiving point cloud data representing the scene; ​ ​ ​ determining a plurality of first sets of features of the image frame; determining a plurality of second sets of features of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third set of features of the image frame based on a first set of features of the plurality of first sets of features of the image frame and a second set of features of the plurality of second sets of features of the point cloud data; and outputting fusion data that combines the third set of features of the image frame and a fourth set of features of the point cloud data.

12. The apparatus of claim 11, the operations further comprising: determining the first set of features and the second set of features based on first statistical indicators associated with the plurality of first sets of features of the image frame and second statistical indicators associated with the plurality of second sets of features of the point cloud data.

13. The apparatus of claim 11, wherein each first set of features of the plurality of first sets of features of the image frame corresponds to a respective level of a plurality of first levels of a first encoder, and wherein each second set of features of the plurality of second sets of features of the point cloud data corresponds to a respective level of a plurality of second levels of a second encoder.

14. The apparatus of claim 13, wherein the respective level corresponding to the first set of features of the image frame corresponds to the respective level corresponding to the second set of features of the point cloud data.

15. The apparatus of claim 12, wherein the second set of features includes pairs of perspective view features of the point cloud data and BEV features of the point cloud data.

16. The apparatus of claim 15, wherein the third set of features of the image frame is determined by an encoder, the operations further comprising: adding a perspective view feature of a pair of the pairs as a first channel of the encoder; and adding a BEV feature of the pair as a second channel of the encoder.

17. The apparatus of claim 15, wherein the plurality of perspective view features and the plurality of BEV features are each determined from the plurality of voxels using global max pooling.

18. The apparatus of claim 11, wherein the point cloud data is received from a LiDAR sensor or a radar sensor.

19. The apparatus of claim 11, wherein the operations further comprise detecting an object represented in the image frame based on the fusion data.

20. The apparatus of claim 19, wherein the operations further comprise controlling a function of a vehicle based on the detected object.

21. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform operations comprising: receiving an image frame representing a scene; receiving point cloud data representing the scene; determining a plurality of first sets of features of the image frame; determining a plurality of second sets of features of the point cloud data based on a plurality of voxels representing the point cloud data; determine a third set of features of the image frame based on a first set of features of the plurality of first sets of features of the image frame and a second set of features of the plurality of second sets of features of the point cloud data; and output fusion data that combines the third set of features of the image frame and a fourth set of features of the point cloud data.

22. The non-transitory computer-readable medium of claim 21, the operations further comprising: determine the first set of features and the second set of features based on first statistical indicators associated with the plurality of first sets of features of the image frame and second statistical indicators associated with the plurality of second sets of features of the point cloud data.

23. The non-transitory computer-readable medium of claim 21, wherein each of the plurality of third sets of features of the image frame corresponds to a respective level of a plurality of first levels of a first encoder, and wherein each of a plurality of fourth sets of features of the point cloud data corresponds to a respective level of a plurality of second levels of a second encoder.

24. The non-transitory computer-readable medium of claim 22, wherein the second set of features includes pairs of perspective view features of the point cloud data and BEV features of the point cloud data.

25. The non-transitory computer-readable medium of claim 24, wherein the third set of features of the image frame is determined by an encoder, the method further comprising: adding a perspective view feature of a pair of the pairs as a first channel of the encoder; and adding a BEV feature of the pair as a second channel of the encoder.

26. A vehicle, the vehicle comprising: an image sensor; a ranging sensor; a processor; and a memory storing instructions that, when executed by the processor, cause the processor to perform operations comprising: receiving, from the image sensor, an image frame representing a scene; receiving, from the ranging sensor, point cloud data representing the scene; determining a plurality of first sets of features of the image frame; determining a plurality of second sets of features of the point cloud data based on a plurality of voxels representing the point cloud data; determining a third set of features of the image frame based on a first set of features of the plurality of first sets of features of the image frame and a second set of features of the plurality of second sets of features of the point cloud data; and outputting fusion data that combines the third set of features of the image frame and a fourth set of features of the point cloud data.

27. The vehicle of claim 26, the operations further comprising: determining the first set of features and the second set of features based on first statistical indicators associated with the plurality of first sets of features of the image frame and second statistical indicators associated with the plurality of second sets of features of the point cloud data.

28. The vehicle of claim 26, wherein each first feature set of the plurality of first feature sets of the image frame corresponds to a respective level of a plurality of first levels of a first encoder, wherein each second feature set of the plurality of second feature sets of the point cloud data corresponds to a respective level of a plurality of second levels of a second encoder, and wherein the respective level corresponding to the first feature set of the image frame corresponds to the respective level corresponding to the second feature set of the point cloud data.

29. The vehicle of claim 26, wherein the second feature sets include pairs of perspective view features of the point cloud data and BEV features of the point cloud data, wherein the third feature set of the image frame is determined by an encoder, and the operations further comprise: adding a perspective view feature in a pair of the pairs as a first channel of the encoder; and adding a BEV feature in the pair as a second channel of the encoder.

30. The vehicle of claim 26, wherein the operations further comprise: detecting an object represented in the image frame based on the fused data; and controlling a function of the vehicle based on the detected object.

30. The vehicle of claim 29, wherein the function is one of: a steering function, a braking function, a throttle function, a speed limit function, a lane keeping function, a collision avoidance function, a pedestrian detection function, a traffic sign detection function, a traffic light detection function, a pedestrian crossing detection function, a vehicle detection function, a vehicle classification function, a vehicle tracking function, a vehicle speed estimation function, a vehicle distance estimation function, a vehicle direction estimation function, a vehicle acceleration estimation function, a vehicle deceleration estimation function, a vehicle trajectory estimation function, a vehicle behavior estimation function, a vehicle state estimation function, a vehicle intent estimation function, a vehicle type estimation function, a vehicle type classification function, a vehicle type tracking function, a vehicle type speed estimation function, a vehicle type distance estimation function, a vehicle type direction estimation function, a vehicle type acceleration estimation function, a vehicle type deceleration estimation function, a vehicle type trajectory estimation function, a vehicle type behavior estimation function, a vehicle type state estimation function, a vehicle type intent estimation function, a vehicle type estimation function, a vehicle type classification function, a vehicle type tracking function, a vehicle type speed estimation function, a vehicle type distance estimation function, a vehicle type direction estimation function, a vehicle type acceleration estimation function, a vehicle type deceleration estimation function, a vehicle type trajectory estimation function, a vehicle type behavior estimation function, a vehicle type state estimation function, a vehicle type intent estimation function, a vehicle type estimation function, a vehicle type classification function, a vehicle type tracking function, a vehicle type speed estimation function, a vehicle type distance estimation function, a vehicle type direction estimation function, a vehicle type acceleration estimation function, a vehicle type deceleration estimation function, a vehicle type trajectory estimation function, a vehicle type behavior estimation function