Cross-modal sensor training
By combining the cross-modal training method of LIDAR and camera systems in the vehicle, the problem of low processing efficiency of high-resolution image data in the vehicle is solved, more accurate object detection and situational awareness are achieved, and road safety is improved.
Patent Information
- Application Number
- CN202380076816.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-09-14
- Filing Date
- 2023-09-15
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively assist drivers in vehicles to enhance situational awareness, especially when processing high-resolution image data, resulting in increased processing volume and reduced efficiency.
By combining LIDAR system and camera system, cross-modal training methods are adopted to train LIDAR-based models and use their output to improve camera-based models, thereby improving the accuracy and efficiency of object detection.
It realizes more accurate object detection and tracking in the vehicle-assisted driving system, improves the vehicle's situational awareness of the surrounding environment, reduces the risk of vehicle collisions and improves road safety.
Smart Images

Figure CN120153408A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Patent Application No. 18 / 467,455, entitled "INTERMODAL SENSOR TRAINING", filed on September 14, 2023, and U.S. Provisional Patent Application No. 63 / 383,024, entitled "INTERMODAL SENSOR TRAINING", filed on November 9, 2022, the entire contents of both of which are hereby expressly incorporated by reference. Technical Field
[0003] Aspects of the present disclosure generally relate to image processing and, more particularly, to multi - object detection. Some features enable and provide improved image processing, including training models for performing object detection.
[0004] Introduction
[0005] Vehicles come in many shapes and sizes, are propelled by various propulsion technologies, and carry cargo including people, animals, or objects. These machines are capable of moving cargo over long distances, moving cargo at high speeds, and moving larger cargo than can be moved by human power. Vehicles were initially driven by humans to control the speed and direction of the cargo's arrival at the destination. The human operation of vehicles has led to many unfortunate accidents caused by collisions between vehicles and vehicles, vehicles and objects, vehicles and humans, or vehicles and animals. With the progress of research on vehicle automation, a variety of driving assistance systems have been produced and introduced. These include navigation guidance via GPS, adaptive cruise control, lane - change assistance, collision - avoidance systems, night vision, parking assistance, and blind - spot detection.
[0006] An image - capture device is a device that can capture one or more digital images (either still images for photos or image sequences for videos). The capture device can be incorporated into various devices. By way of example, the image - capture device can include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, video surveillance camera), or other devices having digital imaging or video capabilities.
[0007] The amount of image data captured by an image sensor has increased through successive generations of image capture devices. The amount of information captured by an image sensor is related to the number of pixels in the image sensor of the image capture device, and the number of pixels can be measured as the number of megapixels indicating the number of millions of sensors in the image sensor. For example, a 12 - megapixel image sensor has 12 million pixels. A higher megapixel value generally represents a higher - resolution image that is more suitable for user viewing.
[0008] The increased amount of image data captured by an image capture device has some negative impacts as the resolution increases with the additional image data. The additional image data increases the amount of processing performed by the image capture device when determining image frames and videos from the image data and performing other operations related to the image data. SUMMARY
[0009] Some aspects of the present disclosure are summarized below to provide a basic understanding of the technologies discussed. This summary is not an exhaustive overview of all the expected features of the present disclosure, and is neither intended to identify the key or important elements of all aspects of the present disclosure, nor to depict the scope of any or all aspects of the present disclosure. The sole purpose of this summary is to present some concepts of one or more aspects of the present disclosure in a general form as a prelude to the more specific embodiments that are given later.
[0010] The human operator of a vehicle can be distracted, which is a factor in many vehicle collision accidents. Driver distraction can include changing the radio, observing events outside the vehicle, and using electronic devices, etc. Sometimes situations occur that even a careful driver cannot detect in time to prevent a vehicle collision. Aspects of the present disclosure provide improved systems for assisting a driver in a vehicle to enhance situational awareness while driving on a road.
[0011] In some aspects, different modes of a sensor system can be used to improve the training of a model for object detection in an image processing system. For example, in some systems, a LIDAR system with the ability to sense scene position and scene awareness can be used to improve the training for understanding input image data from a camera - based 3D object detection system. A camera - based system may have difficulty precisely locating objects in 3D space because camera - based processing involves projection from a perspective view (PV) to a bird's - eye view (BEV), and due to the lack of explicit guidance, it is difficult to train a camera - based system. The LIDAR system provides a better 3D object detection model because the LIDAR system provides more accurate geometric information from the processed LIDAR point cloud.
[0012] In one example of cross-modal training of a camera-based system with input data from a LIDAR system, a method may include training a LIDAR-based model; receiving first and second time-synchronized input data samples from the LIDAR system and the camera system, respectively; processing the first input data sample through the LIDAR-based model to obtain a first output; processing the second input data sample through the camera-based model to obtain a second output; and training the camera-based model during the processing of the second input data sample using the first output of the LIDAR-based model and / or intermediate features of the LIDAR-based model. With the trained camera-based model, the LIDAR system may be deactivated to save power and / or removed to save cost. The method may then include receiving a third input data sample from the camera system and processing the third input data sample using the camera-based model based on the training from the LIDAR-based model.
[0013] In some embodiments, compared to a camera-only model, a first method for an image processing system learns a set of additional features that are spread out from the camera as they are highly sparse when visualized. In some embodiments, a second method for a camera-to-BEV elevation operation for an image processing system includes more accurate predictions of small and distant objects.
[0014] The image processing methods described herein may be performed by an image capture device and / or on image data captured by one or more image capture devices. An image capture device (a device that may capture one or more digital images, whether a still image photograph or an image sequence of a video) may be incorporated into a variety of devices. By way of example, an image capture device may include a standalone digital camera or digital video camera, a wireless communication device phone equipped with a camera (such as a mobile phone, cellular or satellite radiotelephone), a personal digital assistant (PDA), a panel or tablet device, a gaming device, a computing device (such as a webcam, video surveillance camera), or other devices having digital imaging or video capabilities.
[0015] The image processing techniques described herein may relate to a digital camera having an image sensor and processing circuitry (e.g., an application specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), or a central processing unit (CPU)). The image signal processor (ISP) may include one or more of these processing circuits and be configured to perform operations to obtain image data for processing according to the image processing techniques described herein and / or involved in the image processing techniques described herein. The ISP may be configured to control the capture of image frames from one or more image sensors and determine one or more image frames from the one or more image sensors to generate a view of a scene in an output image frame. The output image frame may be part of a sequence of image frames forming a video sequence. The video sequence may include other image frames received from the image sensor or other image sensors.
[0016] In an example application, the image signal processor (ISP) may receive instructions for capturing a sequence of image frames in response to the loading of software, such as a camera application, to produce a preview display from an image capture device. The image signal processor may be configured to produce a single output image frame stream based on the image frames received from one or more image sensors. The single output image frame stream may include raw image data from the image sensor, merged image data from the image sensor, or corrected image data processed by one or more algorithms within the image signal processor. For example, the image frames may be processed by an image post - processing engine (IPE) and / or other image processing circuitry to process the image frames obtained from the image sensor (which may have had some processing of the data performed on it before being output to the image signal processor) to perform one or more of tone mapping, portrait illumination, contrast enhancement, gamma correction, etc. The output image frame from the ISP may be stored in a memory and retrieved by an application processor executing the camera application, which may perform further processing on the output image frame to adjust the appearance of the output image frame and reproduce the output image frame on a display for user viewing.
[0017] After an output frame representing a scene is determined by an image signal processor and / or an application processor (such as through the image processing techniques described in various embodiments herein), the output image frame can be displayed on a device display as a single still image and / or as part of a video sequence, saved to a storage device as a picture or video sequence, sent over a network, and / or printed to an output medium. For example, an image signal processor (ISP) can be configured to obtain an input frame of image data (e.g., pixel values) from one or more image sensors and, in turn, produce a corresponding output image frame (e.g., a preview display frame, a still image capture, a frame for video, a frame for object tracking, etc.). In other examples, the image signal processor can output the image frame to various output devices and / or camera modules for further processing, such as for 3A parameter synchronization (e.g., autofocus (AF), auto white balance (AWB), and automatic exposure control (AEC)), generating a video file via the output frame, configuring the frame for display, configuring the frame for storage, sending the frame over a network connection, etc. Generally, an image signal processor (ISP) can obtain incoming frames from one or more image sensors, produce a stream of output frames, and output the stream of output frames to various output destinations.
[0018] In some aspects, the output image frame can be produced by combining aspects of the image correction of the present disclosure with other computational photography techniques such as high dynamic range (HDR) photography or multi-frame noise reduction (MFNR). In the case of HDR photography, a first image frame and a second image frame are captured using different exposure times, different apertures, different lenses, and / or other characteristics that can result in an improved dynamic range of the fused image when combining the two image frames. In some aspects, the method can be performed for MFNR photography, where the first image frame and the second image frame are captured using the same or different exposure times and the first image frame and the second image frame are fused to generate a corrected first image frame that has reduced noise compared to the captured first image frame.
[0019] In some aspects, the device can include an image signal processor or a processor (e.g., an application processor) that includes specific functionality for camera control and / or processing, such as enabling or disabling a merging module or otherwise controlling aspects of the image correction. The methods and techniques described herein can be performed entirely by the image signal processor or the processor, or the various operations can be split between the image signal processor and the processor and, in some aspects, across additional processors.
[0020] The device may include one, two, or more image sensors, such as a first image sensor. When there are multiple image sensors, the configurations of these image sensors may be different. For example, the first image sensor may have a larger field of view (FOV) than the second image sensor, or the first image sensor may have different sensitivity or different dynamic range from the second image sensor. In one example, the first image sensor may be a wide-angle image sensor, and the second image sensor may be a tele image sensor. In another example, the first sensor is configured to obtain an image through a first lens having a first optical axis, and the second sensor is configured to obtain an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens may have a first magnification, and the second lens may have a second magnification different from the first magnification. Any of these or other configurations may be part of a lens cluster on a mobile device, such as where multiple image sensors and associated lenses are located at offset positions on the front or rear side of the mobile device. Additional image sensors with larger, smaller, or the same field of view may be included. The image processing techniques described herein may be applied to image frames captured from any of the image sensors in a multi-sensor device.
[0021] In an additional aspect of the present disclosure, a device configured for image processing and / or image capture is disclosed. The apparatus includes components for capturing image frames. The apparatus also includes one or more components for capturing data representative of a scene, such as image sensors (including charge-coupled devices (CCDs), Bayer filter sensors, infrared (IR) detectors, ultraviolet (UV) detectors, complementary metal-oxide-semiconductor (CMOS) sensors) and time-of-flight detectors. The apparatus may also include one or more components for focusing and / or concentrating light onto one or more of the image sensors (including simple lenses, compound lenses, spherical lenses, and aspherical lenses). These components may be controlled to capture a first image frame and / or a second image frame input to the image processing techniques described herein.
[0022] For those of ordinary skill in the art, other aspects, features, and specific implementations will become apparent when the following description of specific exemplary aspects is reviewed in conjunction with the accompanying drawings. Although the features may be discussed below with respect to certain aspects and drawings, various aspects may include one or more of the advantageous features discussed herein. In other words, although one or more aspects may be discussed as having certain advantageous features, one or more of such features may also be used according to various aspects. In a similar manner, although the exemplary aspects may be discussed below as device, system, or method aspects, the exemplary aspects may be implemented in various devices, systems, and methods.
[0023] The method can be embedded as computer program code in a computer-readable medium, the computer program code including instructions to cause a processor to perform the steps of the method. In some embodiments, the processor can be part of a mobile device that includes: a first network adapter configured to send data, such as an image or video as recorded data or as streaming data, over a first network connection of a plurality of network connections; and a processor coupled to the first network adapter and a memory. The processor can cause the output image frames described herein to be sent over a wireless communication network, such as a 5G NR communication network.
[0024] The features and technical advantages of examples in accordance with the present disclosure have been outlined rather broadly above so that the detailed description that follows may be better understood. Additional features and advantages will be described below. The disclosed concepts and specific examples may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. The characteristics of the concepts disclosed herein, both as to their organization and method of operation, as well as associated advantages, will be better understood when considered in conjunction with the accompanying drawings. Each of the drawings provided is for the purpose of illustration and description only and not as a definition of the limits of the claims.
[0025] Although aspects and specific implementations are described by way of some examples in this application, those skilled in the art will understand that additional specific implementations and use cases may arise in many different arrangements and scenarios. The innovations described herein can be implemented across many different platform types, devices, systems, shapes, sizes, and packaging arrangements. For example, aspects and / or uses can be implemented via integrated chips and other non-module-component-based devices (e.g., end-user devices, vehicles, communication devices, computing devices, industrial equipment, retail / purchase devices, medical devices, artificial intelligence (AI)-enabled devices, etc.). Although some examples may or may not specifically point to use cases or applications, the applicability of various types of the described innovations can occur. The scope of specific implementations can range from chip-level or module components to non-module, non-chip-level implementations and further to aggregated, distributed, or original equipment manufacturer (OEM) devices or systems that incorporate one or more aspects of the described innovations. In some practical environments, devices that incorporate the described aspects and features may also necessarily include additional components and features for implementing and practicing the claimed and described aspects. For example, the transmission and reception of wireless signals necessarily includes multiple components for analog and digital purposes (e.g., hardware components including antennas, radio frequency (RF) chains, power amplifiers, modulators, buffers, processors, interleavers, adders / summers, etc.). The innovations described herein are intended to be practiced in a variety of devices, chip-level components, systems, distributed arrangements, end-user devices, etc., having different sizes, shapes, and configurations. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] A further understanding of the nature and advantages of the present disclosure can be realized by reference to the following drawings. In the drawings, like components or features may have the same reference numeral. Additionally, various components of the same type can be distinguished by adding a dash and a second label used to differentiate between like components after the reference numeral. If only the first reference numeral is used in the specification, the description applies to any one of the like components having the same first reference numeral, regardless of the second reference numeral.
[0027] Figure 1A is a perspective view of a motor vehicle having a driver monitoring system according to an embodiment of the present disclosure.
[0028] Figure 1B shows a block diagram of an example device for performing image capture from one or more image sensors.
[0029] Figure 2 is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure.
[0030] Figure 3is a block diagram illustrating details of an example wireless communication system in accordance with one or more aspects.
[0031] Figure 4A A block diagram is shown illustrating processing of image data in a LIDAR-based system in accordance with some embodiments of the present disclosure.
[0032] Figure 4B A block diagram is shown illustrating processing of image data in a camera-based system in accordance with some embodiments of the present disclosure.
[0033] Figure 5 A block diagram is shown illustrating weakly supervised learning in a camera-based system from a LIDAR-based system in accordance with some embodiments of the present disclosure.
[0034] Figure 6 A block diagram is shown illustrating ground truth reinforcement learning in a camera-based system from a LIDAR-based system in accordance with some embodiments of the present disclosure.
[0035] Figure 7A A block diagram is shown illustrating intermediate camera features from a LIDAR-based system in accordance with some embodiments of the present disclosure.
[0036] Figure 7B A block diagram is shown illustrating correlating a portion of intermediate camera features from a LIDAR-based system in accordance with some embodiments of the present disclosure.
[0037] Figure 8A A flowchart is shown illustrating an example method for processing image data to train a second modality image system using first modality image system data in accordance with some embodiments of the present disclosure.
[0038] Figure 8B A flowchart is shown illustrating an example method for processing image data in a second modality image system trained using first modality image system data in accordance with some embodiments of the present disclosure.
[0039] Figure 9 is a block diagram illustrating an example processor configuration for image data processing in an image capture device in accordance with one or more embodiments of the present disclosure.
[0040] Like reference numerals and names in the various figures indicate like elements. DETAILED DESCRIPTION
[0041] The following detailed description in conjunction with the accompanying drawings is intended as a description of various configurations and is not intended to limit the scope of the present disclosure. On the contrary, the detailed description includes specific details for providing a thorough understanding of the subject matter of the present invention. It will be apparent to those skilled in the art that these specific details are not required in every case, and in some instances, well-known structures and components are shown in block diagram form for the sake of clarity of presentation.
[0042] The present disclosure provides systems, devices, methods, and computer-readable media that support image processing (including techniques for object detection) and may be particularly advantageous in intelligent vehicle applications. The provided techniques relate to training an image processing system of a first modality based on a first output of an image processing system of the first modality (e.g., a camera-based image processing system) and a second output of an image processing system of a second modality (e.g., a LiDAR-based image processing system). Specifically, a model associated with the image processing system of the first modality is trained based on the first output and the second output.
[0043] Certain specific implementations of the subject matter described in the present disclosure may be implemented to achieve one or more of the following potential advantages or benefits. In some aspects, the present disclosure provides techniques for improving object detection in camera-based systems by using training information from different modalities of a sensing system. Improved object detection increases the accuracy of downstream perception tasks that utilize these features to provide vehicle assistance services. In particular, these techniques may enable more accurate tracking of vehicles, pedestrians, obstacles, road signs, road markings, and the like.
[0044] One benefit of improved tracking is that it allows a vehicle control system to more accurately navigate a vehicle around obstacles. This can be particularly useful in situations where unexpected obstacles or road conditions may pose a danger to the driver. Additionally, improved tracking can help increase overall safety on the road by reducing vehicle collisions. With better tracking capabilities, a vehicle can respond more quickly to nearby obstacles and can more effectively bypass detected obstacles. These improvements can also extend to driver assistance systems, which can benefit from increased monitoring capabilities. By expanding the number, type, and variety of surrounding objects that can be detected, these systems can provide more accurate warnings and assistance to the driver when necessary, without generating unnecessary notifications or distractions.
[0045] Example devices for capturing image frames using one or more image sensors, such as smart phones, may include a configuration of one, two, three, four, or more cameras on the rear side of the device (e.g., the side opposite the main user display) and / or the front side of the device (e.g., the side the same as the main user display). These devices may include one or more image signal processors (ISPs), computer vision processors (CVPs) (e.g., AI engines), or other suitable circuitry for processing the images captured by the image sensors. The one or more image signal processors (ISPs) may store the output image frames in a memory and / or otherwise provide the output image frames to the processing circuitry (such as via a bus). The processing circuitry may perform further processing, such as encoding, storing, transmitting, or other manipulation of the output image frames.
[0046] As used herein, an image sensor may refer to the image sensor itself and any particular other components coupled to the image sensor for generating an image frame for processing by an image signal processor or other logic circuitry or storing in a memory (whether a short-term buffer or a long-term non-volatile memory). For example, an image sensor may include other components of a camera, including a shutter, a buffer, or other readout circuitry for accessing the individual pixels of the image sensor. An image sensor may also refer to an analog front end or other circuitry for converting an analog signal to a digital representation of the image frame, which is provided to digital circuitry coupled to the image sensor.
[0047] In the description of the embodiments herein, numerous specific details (such as examples of specific components, circuits, and processes) are set forth to provide a thorough understanding of the present disclosure. As used herein, the term "coupled" means directly connected or connected through one or more intermediate components or circuits. Additionally, in the following description and for purposes of explanation, specific terms are set forth to provide a thorough understanding of the present disclosure. However, those skilled in the art will appreciate that implementing the teachings disclosed herein may not require these specific details. In other instances, well-known circuits and devices are shown in block diagram form to avoid obscuring the teachings of the present disclosure.
[0048] Certain portions of the detailed description that follow are presented in terms of procedures, logic blocks, processing, and other symbolic representations of operations on data bits within a computer memory. In the present disclosure, the procedures, logic blocks, processes, etc. are conceived as self-consistent sequences of steps or instructions leading to a desired result. These steps are those requiring physical operations on physical quantities. Although not necessarily, typically, these physical quantities take the form of electrical or magnetic signals capable of being stored, transferred, combined, compared, and otherwise manipulated in a computer system.
[0049] In the figures, a single box may be described as performing one or more functions. The one or more functions performed by the box may be performed in a single component or across multiple components and / or may be performed using hardware, software, or a combination of hardware and software. To clearly illustrate such interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are described below in terms of their functionality. Implementing such functionality as hardware or software depends on the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure. Also, example devices may include components other than those shown, including well-known components such as processors, memories, and the like.
[0050] Aspects of the present disclosure are applicable to any electronic device that includes, is coupled to, or otherwise processes data from one, two, or more image sensors capable of capturing image frames (or “frames”). The terms “output image frame” and “corrected image frame” may refer to an image frame that has been processed by any of the techniques discussed herein. Additionally, aspects of the present disclosure may be implemented in image sensors or devices coupled to the image sensors having the same or different capabilities and characteristics, such as resolution, shutter speed, sensor type, and the like. Further, aspects of the present disclosure may be implemented in devices for processing image frames, whether or not the device includes or is coupled to an image sensor, such as a processing device that can retrieve stored images for processing, including processing devices present in cloud computing systems.
[0051] Unless otherwise specifically stated, it should be understood that throughout this application, discussions using terms such as “access,” “receive,” “transmit,” “use,” “select,” “determine,” “normalize,” “multiply,” “average,” “monitor,” “compare,” “apply,” “update,” “measure,” “derive,” “set,” “generate,” etc., refer to actions and processes of a computer system or similar electronic computing device that manipulate and transform data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the registers, memories, or other such information storage, transmission, or display devices of the computer system.
[0052] The terms "device" and "apparatus" are not limited to one or a specific number of physical objects (e.g., a smart phone, a camera controller, a processing system, etc.). As used herein, a device can be any electronic device having one or more parts that can implement at least some parts of the present disclosure. Although the description and examples herein use the term "device" to describe various aspects of the present disclosure, the term "device" is not limited to a specific configuration, type, or number of objects. As used herein, an apparatus can include a device or a part of a device for performing the described operations.
[0053] Certain components in a device or apparatus described as "component for accessing", "component for receiving", "component for transmitting", "component for using", "component for selecting", "component for determining", "component for normalizing", "component for multiplying", or other similarly named terms referring to one or more operations on data (such as image data) can refer to a processing circuit (e.g., an application specific integrated circuit (ASIC), a digital signal processor (DSP), a graphics processing unit (GPU), a central processing unit (CPU)) configured to perform the functions by a combination of hardware, software, or hardware configured by software.
[0054] Figure 1A is a perspective view of a motor vehicle having a driver monitoring system according to an embodiment of the present disclosure. The vehicle 160 can include a forward-facing camera 172 mounted in the cab for observing through the windshield 162. The vehicle can also include a cab-facing camera 174 mounted in the cab and directed toward the occupants of the vehicle 160 and in particular the driver of the vehicle 160. Although a set of mounting positions for the cameras 172 and 174 are shown for the vehicle 160, other mounting positions can also be used for the cameras 172 and 174. For example, one or more cameras can be mounted on one of the driver or passenger B-pillars 186 or on one of the driver or passenger C-pillars 188, such as near the top of the pillar 186 or 188. As another example, one or more cameras can be mounted at the front of the vehicle 160, such as behind the radiator grille 190 or integrated with the bumper 192. As a further example, one or more cameras can be mounted as part of the driver or passenger side mirror assembly 194.
[0055] The camera 172 can be oriented such that the field of view of the camera 172 captures a scene in front of the vehicle 160 in the direction in which the vehicle 160 is moving when in a drive mode or a forward direction. In some embodiments, additional cameras can be located at the rear of the vehicle 160 and oriented such that the field of view of the additional cameras captures a scene behind the vehicle 160 in the direction in which the vehicle 160 is moving when in a reverse direction. Although embodiments of the present disclosure may be described with reference to a “forward-facing” camera (reference camera 172), aspects of the present disclosure can be similarly applied to a “rear-facing” camera that faces the reverse direction of the vehicle 160. Thus, the benefits obtained when an operator drives the vehicle 160 in a forward direction can be similarly obtained when the operator drives the vehicle 160 in a reverse direction.
[0056] In addition, although embodiments of the present disclosure may be described with reference to a “forward-facing” camera (reference camera 172), aspects of the present disclosure can be similarly applied to inputs received from a camera array mounted around the vehicle 160 to provide a larger field of view, which can be as large as approximately 360 degrees parallel to the ground and / or as large as approximately 360 degrees in a vertical direction approximately perpendicular to the ground. For example, additional cameras can be mounted around the outside of the vehicle 160, such as on or integrated into doors, on or integrated into wheels, on or integrated into bumpers, on or integrated into hoods, and / or on or integrated into roofs.
[0057] The camera 174 can be oriented such that the field of view of the camera 174 captures a scene in the cab of the vehicle and includes the user operator of the vehicle, and in particular, the face of the user operator of the vehicle, with sufficient detail to resolve the direction of the user operator's gaze.
[0058] Each of the cameras 172 and 174 can include one, two, or more image sensors, such as including a first image sensor. When there are multiple image sensors, the first image sensor can have a larger field of view (FOV) than the second image sensor, or the first image sensor can have a different sensitivity or a different dynamic range than the second image sensor. In one example, the first image sensor can be a wide-angle image sensor, and the second image sensor can be a telephoto image sensor. In another example, the first sensor is configured to obtain an image through a first lens having a first optical axis, and the second sensor is configured to obtain an image through a second lens having a second optical axis different from the first optical axis. Additionally or alternatively, the first lens can have a first magnification, and the second lens can have a second magnification different from the first magnification. Such a configuration can occur in a camera module having a lens group, where multiple image sensors and associated lenses are located at offset positions within the camera module. Additional image sensors with larger, smaller, or the same field of view can be included.
[0059] Each image sensor may include components for capturing data representative of a scene, such as an image sensor (including a charge-coupled device (CCD), a Bayer filter sensor, an infrared (IR) detector, an ultraviolet (UV) detector, a complementary metal oxide semiconductor (CMOS) sensor) and / or a time-of-flight detector. The apparatus may also include one or more components for focusing and / or concentrating light onto one or more of the image sensors (including a simple lens, a compound lens, a spherical lens, and an aspherical lens). These components may be controlled to capture a first image frame, a second image frame, and / or more image frames. The image frames may be processed to form a single output image frame (such as, by a fusion operation), and the output image frame may be further processed in accordance with the aspects described herein.
[0060] As used herein, an image sensor may refer to the image sensor itself and any particular other components coupled to the image sensor for generating image frames for processing by an image signal processor or other logic circuitry or for storage in a memory (whether a short-term buffer or a long-term non-volatile memory). For example, an image sensor may include other components of a camera, including a shutter, a buffer, or other readout circuitry for accessing individual pixels of the image sensor. An image sensor may also refer to an analog front end or other circuitry for converting an analog signal to a digital representation of an image frame, which digital representation is provided to digital circuitry coupled to the image sensor.
[0061] Figure 1B A block diagram of an example device 100 (e.g., vehicle 160) for performing image capture from one or more image sensors is shown. Device 100 may include or otherwise be coupled to an image signal processor 112 for processing image frames from one or more image sensors, such as a first image sensor 101, a second image sensor 102, and a depth sensor 140. In some particular implementations, device 100 also includes or is coupled to a processor 104 and a memory 106 storing instructions 108. Device 100 may also include or be coupled to a display 114 and input / output (I / O) components 116. The I / O components 116 may be used to interact with a user, such as a touchscreen interface and / or physical buttons.
[0062] The I / O component 116 may also include a network interface for communicating with other devices including a wide area network (WAN) adapter 152, a local area network (LAN) adapter 153, and / or a personal area network (PAN) adapter 154. An example WAN adapter is a 4G LTE or 5G NR wireless network adapter. An example LAN adapter 153 is an IEEE 802.11 WiFi wireless network adapter. An example PAN adapter 154 is a Bluetooth wireless network adapter. Each of the adapters 152, 153, and / or 154 may be coupled to an antenna, which includes multiple antennas configured for main set reception and diversity reception and / or configured for receiving a specific frequency band.
[0063] The device 100 may also include or be coupled to a power source 118 for the device 100, such as a battery or a component that couples the device 100 to an energy source. The device 100 may also include or be coupled to Figure 1B additional feature portions or components not shown. In one example, a wireless interface that may include multiple transceivers and a baseband processor may be coupled to or included in the WAN adapter 152 for a wireless communication device. In another example, an analog front end (AFE) for converting analog image frame data to digital image frame data may be coupled between the image sensors 101 and 102 and the image signal processor 112.
[0064] The device may include or be coupled to a sensor hub 150, which is used to interface with sensors to receive data about the movement of the device 100, data about the environment around the device 100, and / or other non-camera sensor data. An example non-camera sensor is a gyroscope, i.e., a device configured to measure rotation, orientation, and / or angular velocity to generate motion data. Another example non-camera sensor is an accelerometer, i.e., a device configured to measure acceleration, which can also be used to determine speed and the distance traveled by appropriately integrating the measured acceleration, and one or more of acceleration, speed, and / or distance may be included in the generated motion data. In some aspects, the gyroscope in an electronic image stabilization system (EIS) may be coupled to the sensor hub or directly coupled to the image signal processor 112. In other examples, the non-camera sensor may be a global positioning system (GPS) receiver, a light detection and ranging (LiDAR) system, a radio detection and ranging (RADAR) system, or other ranging systems. For example, the sensor hub 150 may interface with a vehicle bus to transmit configuration commands and / or receive information from vehicle sensors 156, such as a distance (e.g., ranging) sensor or a vehicle-to-vehicle (V2V) sensor (e.g., a sensor for receiving information from nearby vehicles).
[0065] The image signal processor 112 may receive image data such as for forming an image frame. In one embodiment, a local bus connection couples the image signal processor 112 to the image sensors 101 and 102 of the first camera 103 and the second camera 105, respectively. In another embodiment, a wire interface couples the image signal processor 112 to an external image sensor. In yet another embodiment, a wireless interface couples the image signal processor 112 to the image sensors 101, 102.
[0066] The first camera 103 may include a first image sensor 101 and a corresponding first lens 131. The second camera may include a second image sensor 102 and a corresponding second lens 132. Each of the lenses 131 and 132 may be controlled by an associated autofocus (AF) algorithm 133 executed in the ISP 112, which adjusts the lenses 131 and 132 to focus on a specific focal plane at a certain scene depth from the image sensors 101 and 102. The AF algorithm 133 may be assisted by a depth sensor 140.
[0067] The first image sensor 101 and the second image sensor 102 are configured to capture one or more image frames. The lenses 131 and 132 focus light at the image sensors 101 and 102, respectively, through one or more apertures for receiving light, one or more shutters for blocking light when outside an exposure window, one or more color filter arrays (CFAs) for filtering light outside a specific frequency range, one or more analog front ends for converting analog measurements to digital information, and / or other suitable components for imaging. The first lens 131 and the second lens 132 may have different fields of view to capture different representations of a scene. For example, the first lens 131 may be an ultra-wide (UW) lens, and the second lens 132 may be a wide (W) lens. The plurality of image sensors may include a combination of ultra-wide (high field of view (FOV)) sensors, wide sensors, tele sensors, and ultra-tele (low FOV) sensors.
[0068] That is, each image sensor can be configured via hardware configuration and / or software settings to obtain different but overlapping fields of view. In one configuration, the image sensor is configured with different lenses having different magnifications, which results in different fields of view. The sensors can be configured such that the UW sensor has a larger FOV than the W sensor, the W sensor has a larger FOV than the T sensor, and the T sensor has a larger FOV than the UT sensor. For example, a sensor configured for a wide FOV can capture a field of view in the range of 64 degrees to 84 degrees, a sensor configured for an ultra-side FOV can capture a field of view in the range of 100 degrees to 140 degrees, a sensor configured for a remote FOV can capture a field of view in the range of 10 degrees to 30 degrees, and a sensor configured for an ultra-remote FOV can capture a field of view in the range of 1 degree to 8 degrees.
[0069] Camera 103 can have a variable aperture (VA) camera, where the aperture can be controlled to a specific size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, etc. Larger aperture values correspond to smaller aperture sizes, and smaller aperture values correspond to larger aperture sizes. Camera 103 can have different characteristics based on the current aperture size, such as different depths of field (DOF) at different aperture sizes.
[0070] Image signal processor 112 processes the image frames captured by image sensors 101 and 102. Although Figure 1B Device 100 is illustrated as including two image sensors 101 and 102 coupled to image signal processor 112, any number (e.g., one, two, three, four, five, six, etc.) of image sensors can be coupled to image signal processor 112. In some aspects, a depth sensor such as depth sensor 140 can be coupled to image signal processor 112 and process the output from the depth sensor in a manner similar to that of image sensors 101 and 102. Example depth sensors include active sensors, including one or more of indirect time of flight (iToF), direct time of flight (dToF), light detection and ranging (LiDAR), mmWave, radio detection and ranging (radar), and / or hybrid depth sensors (such as structured light). In embodiments without depth sensor 140, similar information about the depth or depth map of an object can be generated passively from the disparity between two image sensors (e.g., using disparity depth measurement or stereo depth measurement), phase detection autofocus (PDAF) sensors, etc. Additionally, there can be any number of additional image sensors or image signal processors for device 100.
[0071] In some embodiments, the image signal processor 112 may execute instructions from a memory, such as instructions 108 from memory 106, instructions stored in a separate memory coupled to or included in the image signal processor 112, or instructions provided by the processor 104. Additionally or alternatively, the image signal processor 112 may include specific hardware (such as one or more integrated circuits (ICs)) configured to perform one or more operations described in the present disclosure. For example, the image signal processor 112 may include one or more image front ends (IFE) 135, one or more image post - processing engines 136 (IPE), one or more automatic exposure compensation (AEC) 134 engines, and / or one or more video analysis engines (EVA). AF 133, AEC 134, IFE 135, IPE 136, and EVA 137 may each include dedicated circuitry and may be embodied as software code executed by the ISP 112 and / or a combination of hardware and software code executed on the ISP 112.
[0072] In some particular implementations, the memory 106 may include a non - transitory or non - transient computer - readable medium storing computer - executable instructions 108 to perform all or a part of one or more operations described in the present disclosure. In some particular implementations, the instructions 108 include a camera application (or other suitable application) for generating an image or video to be executed by the device 100. The instructions 108 may also include other applications or programs to be executed by the device 100, such as an operating system and specific applications other than those for image or video generation. A camera application, such as executed by the processor 104, may cause the device 100 to generate an image using the image sensors 101 and 102 and the image signal processor 112. The memory 106 may also be accessed by the image signal processor 112 to store processed frames or may be accessed by the processor 104 to obtain the processed frames. In some embodiments, the device 100 does not include the memory 106. For example, the device 100 may be a circuit including the image signal processor 112, and the memory may be external to the device 100. The device 100 may be coupled to an external memory and be configured to access the memory to write output frames for display or long - term storage. In some embodiments, the device 100 is a system - on - chip (SoC) that combines the image signal processor 112, the processor 104, the sensor hub 150, the memory 106, and the input / output component 116 into a single package.
[0073] In some embodiments, at least one of image signal processor 112 or processor 104 executes instructions to perform the various operations described herein, including object detection and model training operations. For example, execution of the instructions may direct image signal processor 112 to begin or end capturing an image frame or sequence of image frames, where the capture includes a scene with an object as described in the embodiments herein. In some embodiments, processor 104 may include one or more general-purpose processor cores 104A capable of executing a script or instructions of one or more software programs (such as instructions 108 stored in memory 106). For example, processor 104 may include one or more application processors configured to execute a camera application (or other suitable application for generating images or video) stored in memory 106.
[0074] When executing the camera application, processor 104 may be configured to direct image signal processor 112 to perform one or more operations with reference to image sensors 101 or 102. For example, a camera application executed on processor 104 may receive a user command to start a video preview display, and upon receipt of the user command, capture and process a video including a sequence of image frames from one or more of image sensors 101 or 102 via image signal processor 112. Image processing such as to produce an "output" or "corrected" image frame according to the techniques described herein may be applied to one or more of the image frames in the sequence. Execution of instructions 108 by processor 104 external to the camera application may also cause device 100 to perform any number of functions or operations. In some embodiments, processor 104 may include an IC or other hardware (such as artificial intelligence (AI) engine 124 or other coprocessor) to offload certain tasks from core 104A. AI engine 124 may be used to offload tasks related to, for example, face detection and / or object recognition. In some other embodiments, device 100 does not include processor 104, such as when all of the described functionality is configured in image signal processor 112.
[0075] In some embodiments, the display 114 may include one or more suitable displays or screens that allow user interaction and / or present items to the user, such as previews of image frames captured by the image sensors 101 and 102. In some embodiments, the display 114 is a touch-sensitive display. The I / O component 116 may be or include any suitable mechanism, interface, or device to receive input (such as commands) from the user and provide output to the user via the display 114. For example, the I / O component 116 may include (but is not limited to) a graphical user interface (GUI), a keyboard, a mouse, a microphone, a speaker, a squeezable bezel, one or more buttons (such as a power button), a slider, a switch, etc. In some embodiments related to autonomous driving, the I / O component 216 may include an interface to a vehicle bus for providing commands and information to and receiving information from the vehicle system 158, which includes propulsion (e.g., commands to increase or decrease speed or apply brakes) and steering systems (e.g., commands to turn the wheels, change routes, or change a final destination). According to embodiments of the present disclosure, the accuracy of outputting commands to the vehicle system 158 can be improved by training a camera-based system using input data from the LIDAR system to improve object detection that may affect the commands transmitted to the vehicle system 158.
[0076] Although shown as being coupled to each other via the processor 104, components such as the processor 104, the memory 106, the image signal processor 112, the display 114, and the I / O component 116 may be coupled to each other in various other arrangements, such as being coupled to each other via one or more local buses, which are not shown for simplicity. Although the image signal processor 112 is illustrated as being separate from the processor 104, the image signal processor 112 may be the core of the processor 104, which is an application processor unit (APU), included in a system-on-chip (SoC), or otherwise included in the processor 104. Although aspects of the present disclosure are described herein with reference to the device 100 for execution, some device components may not be Figure 1B shown in order to prevent obscuring aspects of the present disclosure. Additionally, other components, the number of components, or combinations of components may be included in a suitable device for executing aspects of the present disclosure. Accordingly, the present disclosure is not limited to a particular device or component configuration, including the device 100.
[0077] Operable Figure 1B exemplary image capture devices to obtain improved images by better detecting objects in the scenes of the images. Figure 2 An example method of operating one or more cameras (such as the camera 103) is shown and described below.
[0078] Figure 2FIG. 0 is a block diagram illustrating an example data flow path for image data processing in an image capture device according to one or more embodiments of the present disclosure. A processor 104 of the system 200 may communicate with an image signal processor (ISP) 112 via a bidirectional bus and / or separate control and data lines. The processor 104 may control the camera 103 via a camera control 210, such as to configure the camera 103 via a driver executed on the processor 104. The camera control 210 may be managed by a camera application 204 executed on the processor 104, which provides user-accessible settings such that a user may specify individual camera settings or select a profile with corresponding camera settings. The camera control 210 communicates with the camera 103 to configure the camera 103 according to commands received from the camera application 204. The camera application 204 may be, for example, a photography application, a document scanning application, a messaging application, or other application that processes image data obtained from the camera 103.
[0079] Camera configurations may be parameters that specify, for example, frame rate, image resolution, readout duration, exposure level, aspect ratio, aperture size, and the like. The camera 103 may obtain image data based on the camera configuration. For example, the processor 104 may execute the camera application 204 to instruct the camera control 210 to instruct the camera 103 to set a first camera configuration of the camera 103, obtain first image data from the camera 103 operating in the first camera configuration, instruct the camera 103 to set a second camera configuration of the camera 103, and obtain second image data from the camera 103 operating in the second camera configuration.
[0080] In some embodiments in which the camera 103 is a variable aperture (VA) camera system, the processor 104 may execute the camera application 204 to instruct the camera 103 to be configured to a first aperture size, obtain first image data from the camera 103, instruct the camera 103 to be configured to a second aperture size, and obtain second image data from the camera 103. Reconfiguration of the aperture and the obtaining of the first and second image data may occur when there is little or no change in the scene that may be captured at the first aperture size and at the second aperture size. Example aperture sizes are f / 2.0, f / 2.8, f / 3.2, f / 8.0, and the like. Larger aperture values correspond to smaller aperture sizes, and smaller aperture values correspond to larger aperture sizes. That is, f / 2.0 is a larger aperture size than f / 8.0.
[0081] Image data received from camera 103 can be processed in one or more blocks of ISP 112 to form image frame 230 stored in memory 106 and / or provided to processor 104. Processor 104 can further process the image data to apply effects to image frame 230. The effects can include Bokeh, lighting, color cast, and / or high dynamic range (HDR) merging. In some embodiments, the functionality can be embedded in different components, such as ISP 112, DSP, ASIC, or other custom logic circuitry for performing additional image processing.
[0082] Device 100 can communicate as a user equipment (UE) within wireless network 300, such as via WAN adapter 252, as Figure 3 shown. Figure 3 is a block diagram illustrating details of an example wireless communication system in accordance with one or more aspects. Wireless network 300 can include, for example, a 5G wireless network. As recognized by those skilled in the art, Figure 3 the components present in are likely to have related corresponding components in other network arrangements, including, for example, cellular-style network arrangements and non-cellular-style network arrangements (e.g., device-to-device or peer-to-peer or ad-hoc network arrangements).
[0083] Figure 3 The illustrated wireless network 300 includes base stations 305 and other network entities. A base station can be a station that communicates with a UE and can also be referred to as an evolved Node B (eNB), a next-generation eNB (gNB), an access point, etc. Each base station 305 can provide communication coverage for a particular geographic area. In 3GPP, the term "cell" can refer to the particular geographic coverage area of a base station or the base station subsystem serving that coverage area, depending on the context in which the term is used. In the specific implementation of wireless network 300 herein, base stations 305 can be associated with the same operator or different operators (e.g., wireless network 300 can include multiple operator wireless networks). Additionally, in the specific implementation of wireless network 300 herein, base stations 305 can use one or more frequencies in the same frequency as adjacent cells (e.g., one or more frequency bands in licensed spectrum, unlicensed spectrum, or a combination thereof) to provide wireless communication. In some examples, a separate base station 305 or UE 315 can be operated by more than one network operating entity. In some other examples, each base station 305 and UE 315 can be operated by a single network operating entity.
[0084] A base station can provide communication coverage for a macro cell or a small cell (e.g., a pico cell or a femto cell) or other types of cells. A macro cell generally covers a relatively large geographical area (e.g., with a radius of several kilometers) and can allow unrestricted access by UEs having a service subscription with the network provider. A small cell (such as a pico cell) generally covers a relatively small geographical area and can allow unrestricted access by UEs having a service subscription with the network provider. A small cell (such as a femto cell) generally also covers a relatively small geographical area (e.g., a home), and in addition to unrestricted access, can also provide restricted access by UEs associated with the femto cell (e.g., UEs in a closed subscriber group (CSG), UEs of users in a home, etc.). A base station for a macro cell can be referred to as a macro base station. A base station for a small cell can be referred to as a small cell base station, a pico base station, a femto base station, or a home base station. In Figure 3 In the example shown, base stations 305d and 305e are conventional macro base stations, while base stations 305a - 305c are macro base stations implemented using one of three-dimensional (3D), full-dimensional (FD), or massive MIMO. Base stations 305a - 305c utilize their higher-dimensional MIMO capabilities to employ 3D beamforming in elevation and azimuth beamforming to increase coverage and capacity. Base station 305f is a small cell base station, which can be a home node or a portable access point. A base station can support one or more (e.g., two, three, four, etc.) cells.
[0085] The wireless network 300 can support synchronous or asynchronous operation. For synchronous operation, base stations can have similar frame timings, and transmissions from different base stations can be approximately aligned in time. For asynchronous operation, base stations can have different frame timings, and transmissions from different base stations may not be aligned in time. In some cases, the network can be enabled or configured to handle dynamic switching between synchronous and asynchronous operations.
[0086] UEs 315 are scattered throughout the wireless network 300, and each UE can be stationary or mobile. It should be understood that although in the standards and specifications promulgated by 3GPP, a mobile device is generally referred to as a UE, such a device can additionally or otherwise be referred to by those skilled in the art as a mobile station (MS), subscriber station, mobile unit, subscriber unit, radio unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal (AT), mobile terminal, wireless terminal, remote terminal, cell phone, terminal, user agent, mobile client, client, gaming device, augmented reality device, vehicle component, vehicle device, or vehicle module or some other suitable term.
[0087] Some non-limiting examples of mobile devices, such as specific implementations that may include one or more of UEs 315, include mobile phones, cellular (cell) phones, smart phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, laptop computers, personal computers (PCs), notebooks, netbooks, smartbooks, tablet computers, personal digital assistants (PDAs), and vehicles. Although UEs 315a-j are specifically shown as vehicles, vehicles may adopt the communication configurations described with reference to any one of UEs 315a-315k.
[0088] In one aspect, a UE may be a device that includes a Universal Integrated Circuit Card (UICC). In another aspect, a UE may be a device that does not include a UICC. In some aspects, a UE that does not include a UICC may also be referred to as an IoE device. Figure 3 The UEs 315a-315d in the specific implementations illustrated are examples of mobile smart phone type devices that access the wireless network 300. A UE may also be a machine specifically configured to enable connected communication, which includes Machine Type Communication (MTC), Enhanced MTC (eMTC), Narrowband IoT (NB-IoT), and the like. Figure 3 The UEs 315e-315k illustrated are examples of various machines configured for communication that access the wireless network 300.
[0089] A mobile device (such as UE 315) may be capable of communicating with any type of base station, whether it is a macro base station, a pico base station, a femto base station, a relay station, etc. In Figure 3 this, the communication link (represented as a lightning bolt) indicates a wireless transmission between the UE and the serving base station (the serving base station is the base station designated to serve the UE on the downlink or uplink), a desired transmission between base stations, and a backhaul transmission between base stations. A UE may operate as a base station or other network node in some scenarios. The backhaul communication between the base stations of the wireless network 300 may be performed using a wired or wireless communication link.
[0090] In operation, at the wireless network 300, base stations 305a-305c use 3D beamforming and cooperative spatial techniques (such as Coordinated Multi-Point (CoMP) or multi-connectivity) to serve UEs 315a and 315b. Macro base station 305d performs backhaul communication with base stations 305a-305c and the small cell (base station 305f). Macro base station 305d also transmits multicast services subscribed to and received by UEs 315c and 315d. Such multicast services may include mobile TV or streaming video, or may include other services for providing community information, such as weather emergencies or alerts, such as Amber alerts or Gray alerts.
[0091] The specifically implemented wireless network 300 supports communication with ultra-reliable and redundant links for such devices. The redundant communication links with the UE 315e include links from the macro base stations 305d and 305e and the small cell base station 305f. Other machine type devices such as the UE 315f (thermometer), UE 315g (smart meter), and UE 315h (wearable device) can communicate directly with base stations such as the small cell base station 305f and the macro base station 305e through the wireless network 300, or communicate in a multi-hop configuration by communicating with another user equipment that relays its information to the network. For example, the UE 315f communicates temperature measurement information to the smart meter UE 315g, and then reports it to the network through the small cell base station 305f. The wireless network 300 can also provide additional network efficiency through dynamic, low-latency TDD communication or low-latency FDD communication (e.g., in a vehicle-to-vehicle (V2V) mesh network between the UEs 315i - 315k communicating with the macro base station 305e).
[0092] Figure 4A A block diagram is shown illustrating the processing of image data in a LiDAR-based system according to some embodiments of the present disclosure. Figure 4A The data flow of illustrates an example of a LiDAR image processing system. When receiving point cloud data as input (e.g., from the depth sensor 140), the LiDAR-based system includes an encoder that extracts intermediate 3D features from the point cloud data. Then, the LiDAR-based system flattens the intermediate 3D features in two dimensions to determine intermediate BEV features. The decoder of the LiDAR-based system predicts 3D bounding boxes from the intermediate BEV features. In various aspects, the LiDAR-based system includes a model for performing the above actions of the LiDAR-based system. For example, the model of the LiDAR-based system can be implemented as one or more machine learning models, including supervised learning models, unsupervised learning models, other types of machine learning models, and / or other types of prediction models. For example, the model can be implemented as one or more of a neural network, a transformer model, a decision tree model, a support vector machine, a Bayesian network, a classifier model, a regression model, etc. In at least some aspects, the model of the LiDAR-based system is trained using ground truth annotations and then fixed for knowledge transfer to the camera model.
[0093] Figure 4B A block diagram is shown illustrating the processing of image data in a camera-based system according to some embodiments of the present disclosure.
[0094] Figure 4BThe data stream of [description] illustrates an example of a camera image processing system. When receiving camera image data as input (e.g., from a first image sensor 101 and a second image sensor 102), a camera-based system includes an encoder that extracts intermediate perspective camera image features from the camera image data. The camera-based system then transforms the intermediate perspective camera image features into the BEV plane to determine intermediate BEV features. A decoder of the camera-based system predicts 3D bounding boxes from the intermediate BEV features. In various aspects, the camera-based system includes a model for performing the above actions of the camera-based system. For example, the model of the camera-based system can be implemented as one or more machine learning models, including supervised learning models, unsupervised learning models, other types of machine learning models, and / or other types of prediction models. For example, the model can be implemented as one or more of a neural network, a transformer model, a decision tree model, a support vector machine, a Bayesian network, a classifier model, a regression model, etc.
[0095] Figure 5 A block diagram is shown that illustrates weak supervision learning in a camera-based system from a LIDAR-based system according to some embodiments of the present disclosure. The output 504 of the camera-based decoder 502 includes a set of 3D bounding boxes within the scene. The output 514 of the LiDAR-based decoder 512 includes a set of 3D bounding boxes within the scene. The outputs 504 and 514 can be compared during a weak supervision learning process, in which the output 504 can be improved based on the output 514. The camera-based system is thus improved based on the output 514 to improve the detection of objects in the scene.
[0096] Figure 6 A block diagram is shown that illustrates ground truth augmentation learning in a camera-based system from a LIDAR-based system according to some embodiments of the present disclosure. Figure 6 The data processing in [description] is similar to Figure 5 The data processing in [description]. A comparison process 602 can compare the output 514 of the LIDAR-based system with the output 504 of the camera-based system. Ground truth bounding boxes 604 (which can indicate known and / or confirmed objects, such as objects selected by a human) can be used to improve the camera-based model, in addition to the output 514.
[0097] Figure 7AA block diagram is shown that illustrates a comparison of camera features and LIDAR features according to some embodiments of the present disclosure. The comparison process 702 may compare the output 514 of the LIDAR-based system with the output 504 of the camera-based system, which may be used to update the object recognition model of the camera-based system. Additionally, LiDAR feature maps may be directly learned in the camera-based system after equalization. For example, as part of process 704, intermediate features of the camera-based model may be compared with intermediate features of the LiDAR-based model, and this comparison may be used to update the camera-based model.
[0098] Figure 7B A block diagram is shown that illustrates pseudo-LiDAR camera features learned from a LIDAR-based system according to some embodiments of the present disclosure. Figure 7B The data stream of... illustrates the generation of feature distillation. In Figure 7B the processing of..., a "pseudo-Lidar" feature map may be directly learned in the camera-based system after equalization, which is different from only camera features. That is, in at least some aspects, only the intermediate features that match between the intermediate LiDAR features and the intermediate camera features are used to train the camera-based system. Based on this training of the camera-based system, in various aspects, the LIDAR-based system is not utilized during the inference phase. In other words, in these aspects, the processing performed by the camera-based system during the inference phase remains the same as Figure 4B that shown in..., although the output 504 generated by the camera-based system is improved based on using the LIDAR-based system to train the camera-based system.
[0099] In some embodiments, the methods herein may include training an additional adversarial network on top of the 3D prediction to predict whether the learned output is from "annotated GT" or "pseudo GT".
[0100] Figure 2 The system 200 of... may be configured to execute the reference Figure 3 ..., Figure 4, Figure 5 ..., Figure 6 ..., Figure 7A and Figure 7B the data flow diagrams of... or Figure 8A ..., Figure 8B or Figure 9The operations described in the method flow chart. This process can be performed based on the received image data. For example, when the image sensor is configured with a camera configuration, the first image data is received from the image sensor. For example, the first image data can be received at the ISP 112, processed by the image front end (IFE) and / or the image post-processing engine (IPE) of the ISP 112, and stored in a memory (such as, memory 106). In some embodiments, the capture of the image data can be initiated by a camera application executed on the processor 104, which causes the camera control 210 to activate the camera 103 to capture the image data and provides the image data to a processor, such as the processor 104 or the ISP 112.
[0101] Figure 8A A method of performing image processing according to the above embodiments is shown. Figure 8A FIG. is a flow chart illustrating an example method 800 for processing image data to train a second modality image system using first modality image system data. Method 800 includes training a first modality imaging system (e.g., a LiDAR-based system) at block 802.
[0102] At block 804, time-synchronized first input data samples and second input data samples are received from the first modality image system and the second modality image system (e.g., a camera-based system), respectively. In other words, the first input data sample (e.g., first point cloud data) is received from the LiDAR-based system and the second input data sample (e.g., first image data) is received from the camera-based system, and the first point cloud data is time-synchronized with the first image data. In various aspects, training the LiDAR-based system includes receiving a third input data sample (e.g., second point cloud data) from the LiDAR-based system and determining a model for the LiDAR-based system based on the first point cloud data and a first ground truth corresponding to the first point cloud data.
[0103] At block 806, the first point cloud data is processed in the LiDAR-based system to generate a first output (e.g., output 514). In various aspects, processing the point cloud data in the LiDAR-based system includes determining intermediate 3D point cloud features based on the first point cloud data, and processing the first image data in the camera-based system includes determining intermediate camera features based on the first image data. In these aspects, method 800 may further include training the camera-based system based on the intermediate 3D point cloud features. The intermediate 3D point cloud features can be extracted from the first point cloud data by a first encoder (e.g., see Figure 4A ). The intermediate camera features can be extracted from the first image data by a second encoder (e.g., see Figure 4B ).
[0104] At block 808, image data is processed in a camera-based system to generate a second output (e.g., output 504). In various aspects, output 514 includes a first plurality of bounding boxes corresponding to objects in a scene, and output 504 includes a second plurality of bounding boxes corresponding to objects in the scene. In these aspects, training the camera-based system includes training the camera-based system using a subset of objects in both the first plurality of bounding boxes and the second plurality of bounding boxes in the scene.
[0105] At block 810, the camera-based system is trained using both output 504 and output 514. In other words, the camera-based system is trained based on the output from the camera-based system and the output from the LiDAR-based system.
[0106] Figure 8B A method of performing image processing using a trained camera-based system according to the above embodiments is shown. Figure 8B It is a flowchart of an example method 820 for processing image data in a second modality image system (e.g., a camera-based system), which is trained using data output by a first modality image system (e.g., a LiDAR-based system). It should be understood that one or more blocks of method 800 may be combined with one or more blocks of method 820. Method 820 includes receiving a third input data sample (e.g., second image data) from the camera-based system at block 822.
[0107] At block 824, the second image data is processed in the camera-based system to generate a third output based on a model for the camera-based system, which is trained based on the output (e.g., output 504) of the first modality imaging system (e.g., a LiDAR-based system). Based on the training from the LiDAR-based system, the third output can be more accurate than the output (e.g., output 504) generated by the camera-based system before training from the LiDAR-based system. In various aspects, the third output includes at least one bounding box corresponding to an object detected in the second image data. In some aspects, method 820 may include operating a vehicle based on the at least one bounding box.
[0108] Figure 9FIG. 0 is a block diagram illustrating an example configuration of a processor 104 for image data processing in an image capture device according to one or more embodiments of the present disclosure. In this example, point cloud data and first image data are received as inputs at the processor 104. The processor 104 performs LiDAR-based detection 904A on the input point cloud data to extract intermediate features of the input point cloud data, determine a 3D bounding box based on the input point cloud data, or both. The processor 104 also performs camera-based detection 904B on the input first image data to extract intermediate features of the input first image data, determine a 3D bounding box based on the input first image data, or both. The processor 104 also performs camera model training 904C to train a model associated with the camera based on LiDAR intermediate features and camera intermediate features, LiDAR 3D bounding boxes and camera 3D bounding boxes, or both. Additionally, the processor 104 may perform LiDAR model training 904D to train a model associated with the LiDAR sensor based on LiDAR 3D bounding boxes and a ground truth set of 3D bounding boxes. Using the trained camera model, the processor 104 may receive second image data as an input and output a 3D bounding box based on the trained camera model and the second image data.
[0109] In one or more aspects, techniques for supporting image processing (which may support vehicle operations) may include additional aspects, such as any individual aspect or any combination of aspects described below or in connection with one or more other processes or devices described elsewhere herein. In a first aspect, an apparatus is configured to perform operations including: training a first modality imaging system; receiving time-synchronized first input data samples and second input data samples from the first modality imaging system and a second modality imaging system, respectively; processing the first input data samples in the first modality imaging system to generate a first output; processing the second input data samples in the second modality imaging system to generate a second output; and training the second modality imaging system based on the first output and the second output. In some implementations, the apparatus includes a wireless device, such as a UE. In some implementations, the apparatus may include at least one processor and a memory coupled to the processor. The processor may be configured to perform the operations described herein for the apparatus. In some other implementations, the apparatus may include a non-transitory computer-readable medium having program code recorded thereon, and the program code may be executable by a computer to cause the computer to perform the operations described herein with reference to the apparatus. In some implementations, the apparatus may include one or more components configured to perform the operations described herein. In some implementations, a method of wireless communication may include one or more operations described herein with reference to the apparatus.
[0110] In a second aspect in combination with the first aspect, training the first modality imaging system includes: receiving a third input data sample from the first modality imaging system; and determining a model for the first modality imaging system based on the first input data sample and a first ground truth corresponding to the first input data sample.
[0111] In a third aspect in combination with one or more of the first aspect or the second aspect, the operation further includes receiving a third input data sample from the second modality imaging system; and processing the third input data sample in the second modality imaging system to generate a third output based on a model for the second modality imaging system, the model being trained based on the first output of the first modality imaging system.
[0112] In a fourth aspect in combination with the third aspect, the third output includes at least one bounding box corresponding to an object detected in the third input data sample.
[0113] In a fifth aspect in combination with the fourth aspect, the operation further includes operating a vehicle based on the at least one bounding box.
[0114] In a sixth aspect in combination with one or more of the first aspect to the fifth aspect, the first modality imaging system includes a LIDAR-based system, and the second modality imaging system includes a camera-based system.
[0115] In a seventh aspect in combination with one or more of the first aspect to the sixth aspect, processing the first input data sample in the first modality imaging system includes determining intermediate 3D point cloud features based on the first input data sample, and processing the second input data sample in the second modality imaging system includes determining intermediate camera features based on the second input data sample. In the seventh aspect, the operation further includes training the second modality imaging system based on the intermediate 3D point cloud features.
[0116] In an eighth aspect in combination with one or more of the first aspect to the seventh aspect, the first output includes a first plurality of bounding boxes corresponding to a first object in a scene, and the second output includes a second plurality of bounding boxes corresponding to a second object in the scene, and training the second modality imaging system includes training the second modality imaging system using a subset of the first object and the second object, the subset of objects being in both the first plurality of bounding boxes and the second plurality of bounding boxes.
[0117] In a ninth aspect, in combination with one or more of the second to eighth aspects, a non-transitory computer-readable medium stores instructions that, when executed by an image signal processor, cause the processor to perform operations including: training a first modality imaging system; receiving time-synchronized first input data samples and second input data samples from the first modality imaging system and a second modality imaging system, respectively; processing the first input data samples in the first modality imaging system to generate a first output; processing the second input data samples in the second modality imaging system to generate a second output; and training the second modality imaging system based on the first output and the second output.
[0118] In a tenth aspect, in combination with one or more of the second to eighth aspects, a method includes training a first modality imaging system; receiving time-synchronized first input data samples and second input data samples from the first modality imaging system and a second modality imaging system, respectively; processing the first input data samples in the first modality imaging system to generate a first output; processing the second input data samples in the second modality imaging system to generate a second output; and training the second modality imaging system based on the first output and the second output.
[0119] Those skilled in the art should understand that any of a variety of different technologies and techniques can be used to represent information and signals. For example, data, instructions, commands, information, signals, bits, symbols, and chips that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.
[0120] As used herein with respect to FIGS. 1 to Figure 9 The components, functional blocks, and modules described herein include processors, electronic devices, hardware devices, electronic components, logic circuits, memories, software code, firmware code, etc., or any combination thereof. Software should be broadly construed to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executables, execution threads, procedures, and / or functions, etc., regardless of whether it is referred to as software, firmware, middleware, microcode, hardware description language, or other terms. Additionally, the features discussed herein may be implemented via dedicated processor circuitry, via executable instructions, or a combination thereof.
[0121] Those skilled in the art should understand that: one or more blocks (or operations) described with reference to FIGS. 4 and Figure 5 may be combined with one or more blocks (or operations) described with reference to another figure in the figures. For example, one or more blocks (or operations) of FIG. 4 may be combined with FIGS. 1 to Figure 3combine one or more of the boxes (or operations). As another example, one or more boxes associated with Figure 5 may be combined with one or more of the boxes (or operations) associated with Figures 5 to 9 .
[0122] Those of ordinary skill in the art should also recognize that: various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the disclosure herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in different ways for each particular application, but such specific implementation decisions should not be construed as causing a departure from the scope of the present disclosure. Those skilled in the art will also readily recognize that the order or combination of components, methods, or interactions described herein are merely examples, and the components, methods, or interactions of various aspects of the present disclosure can be combined or performed in ways other than those illustrated and described herein.
[0123] The various illustrative logical components, logical blocks, modules, circuits, and algorithmic processes described in connection with the specific implementations disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. The interchangeability of hardware and software has been generally described in terms of functionality and illustrated in the various illustrative components, blocks, modules, circuits, and processes above. Whether such functionality is implemented as hardware or software depends upon the particular application and the design constraints imposed on the overall system.
[0124] The hardware and data processing means for implementing or performing the various illustrative logics, logical blocks, modules, and circuits described in connection with the aspects disclosed herein can be realized using a general-purpose single-chip or multi-chip processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, or any conventional processor, controller, microcontroller, or state machine. In some specific implementations, the processor can be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. In some specific implementations, specific processes and methods can be performed by circuitry specific to a given function.
[0125] In one or more aspects, the described functionality may be implemented in hardware, digital electronic circuitry, computer software, firmware, including the structures disclosed in this specification and structural equivalents thereof, or any combination thereof. Specific implementations of the subject matter described in this specification may also be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus.
[0126] If implemented in software, the functionality may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The processes of the methods or algorithms disclosed herein may be implemented in a processor-executable software module that may reside on a computer-readable medium. Computer-readable media includes both computer storage media and communication media including any medium that can be implemented to transfer a computer program from one place to another. The storage media may be any available media that is accessible by a computer. By way of example and not limitation, such computer-readable media may comprise random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection may be properly termed a computer-readable medium. As used herein, disks and optical disks include compact disk (CD), laser disk, optical disk, digital versatile disk (DVD), floppy disk, and Blu-ray disk where disks usually reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media. Additionally, the operations of a method or algorithm may be embodied as a code and instruction set, or any combination of a code and instruction set, residing on a machine-readable medium and a computer-readable medium, which may be incorporated into a computer program product.
[0127] Various modifications to the specific implementations described in this disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to some other specific implementations without departing from the spirit or scope of this disclosure. Thus, the claims are not intended to be limited to the specific implementations shown herein but are to be accorded the widest scope consistent with this disclosure, the principles disclosed herein, and the novel features.
[0128] Additionally, those of ordinary skill in the art will readily recognize that, for ease of description of the drawings, opposing terms such as "upper" and "lower" or "front" and "back" or "top" and "bottom" or "forward" and "backward" are sometimes used and indicate relative positions corresponding to the orientation of the drawing on the correctly oriented page, and may not reflect the correct orientation of any device as implemented.
[0129] Certain features that are described in the context of separate embodiments in this specification can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub - combination in multiple embodiments. Additionally, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from the claimed combination can in some cases be removed from the combination, and the claimed combination can be directed to a sub - combination or variations of the sub - combination.
[0130] Similarly, although operations are depicted in the figures in a particular order, this should not be construed as requiring that such operations be performed in the particular order shown or in a sequential order, or that all of the illustrated operations be performed to achieve the desired result. Additionally, the figures may schematically depict one or more example processes in the form of a flowchart. However, other operations not depicted can be incorporated into the example processes schematically illustrated. For example, one or more additional operations can be performed before, after, simultaneously with, or between any of the illustrated operations. In certain environments, multitasking and parallel processing are advantageous. Additionally, the separation of various system components in the embodiments described above should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Additionally, some other embodiments also fall within the scope of the appended claims. In some cases, the acts recited in the claims can be performed in a different order and still achieve the desired result.
[0131] As used herein (including in the claims), the term "or" as used in a list of two or more items means that any one of the listed items can be employed alone or any combination of two or more of the listed items can be employed. For example, if a composition is described as containing components A, B, or C, the composition can contain A alone; B alone; C alone; a combination of A and B; a combination of A and C; a combination of B and C; or a combination of A, B, and C. Further, as used herein (including in the claims), "or" as used in a list of items beginning with "at least one of" indicates a disjunctive list such that, for example, a list of "at least one of A, B, or C" means A or B or C or AB or AC or BC or ABC (i.e., A and B and C) or any combination of any of these items.
[0132] The term "substantially" is defined as being largely but not necessarily wholly that which is specified (and includes that which is specified; e.g., substantially 90 degrees includes 90 degrees and substantially parallel includes parallel), as understood by one of ordinary skill in the art. In any of the disclosed specific embodiments, the term "substantially" can be replaced with "[percentage] within" that which is specified, where the percentage includes 0.1%, 1%, 5%, or 10%.
[0133] The foregoing description of the disclosure has been provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method, the method comprises: training a first modality imaging system; receiving a first input data sample and a second input data sample from the first modality imaging system and a second modality imaging system respectively, the first input data sample and the second input data sample being time - synchronized; processing the first input data sample in the first modality imaging system to generate a first output; processing the second input data sample in the second modality imaging system to generate a second output; and training the second modality imaging system based on the first output and the second output.
2. The method according to claim 1, wherein training the first modality imaging system comprises: receiving a third input data sample from the first modality imaging system; and determining a model for the first modality imaging system based on the first input data sample and a first ground truth corresponding to the first input data sample.
3. The method according to claim 1, the method further comprises: receiving a third input data sample from the second modality imaging system; and processing the third input data sample in the second modality imaging system to generate a third output based on a model for the second modality imaging system, the model being trained based on the first output of the first modality imaging system.
4. The method according to claim 3, wherein the third output comprises at least one bounding box corresponding to an object detected in the third input data sample.
5. The method according to claim 4, the method further comprises operating a vehicle based on the at least one bounding box.
6. The method according to claim 1, wherein the first modality imaging system comprises a LIDAR - based system, and the second modality imaging system comprises a camera - based system.
7. The method according to claim 1, wherein: processing the first input data sample in the first modality imaging system comprises determining intermediate 3D point cloud features based on the first input data sample, and processing the second input data sample in the second modality imaging system comprises determining intermediate camera features based on the second input data sample, the method further comprises: training the second modality imaging system based on the intermediate 3D point cloud features.
8. The method according to claim 1, wherein: the first output comprises a first plurality of bounding boxes corresponding to a first object in a scene, and the second output comprises a second plurality of bounding boxes corresponding to a second object in the scene, and training the second modality imaging system comprises training the second modality imaging system using a subset of the first object and the second object, the subset of objects being in both the first plurality of bounding boxes and the second plurality of bounding boxes.
9. A non - transitory computer - readable medium storing instructions, the instructions when executed by a processor cause the processor to perform operations comprising the following: training a first modality imaging system; Receiving a first input data sample and a second input data sample from the first modality imaging system and the second modality imaging system respectively, the first input data sample and the second input data sample being time-synchronized; Processing the first input data sample in the first modality imaging system to generate a first output; Processing the second input data sample in the second modality imaging system to generate a second output; And Training the second modality imaging system based on the first output and the second output.
10. The non-transitory computer-readable medium according to claim 9, wherein training the first modality imaging system Comprises: Receiving a third input data sample from the first modality imaging system; And determining a model for the first modality imaging system based on the first input data sample and a first ground truth corresponding to the first input data sample.
11. The non-transitory computer-readable medium according to claim 9, wherein the operations further Comprise: Receiving a third input data sample from the second modality imaging system; And Processing the third input data sample in the second modality imaging system to generate a third output based on a model for the second modality imaging system, the model being trained based on the first output of the first modality imaging system.
12. The non-transitory computer-readable medium according to claim 11, wherein the third output comprises at least one bounding box corresponding to an object detected in the third input data sample.
13. The non-transitory computer-readable medium according to claim 12, the operations further comprising operating a vehicle based on the at least one bounding box.
14. The non-transitory computer-readable medium according to claim 9, wherein the first modality imaging system comprises a LIDAR-based system and the second modality imaging system comprises a camera-based system.
15. The non-transitory computer-readable medium according to claim 9, Wherein: Processing the first input data sample in the first modality imaging system comprises determining intermediate 3D point cloud features based on the first input data sample, and Processing the second input data sample in the second modality imaging system comprises determining intermediate camera features based on the second input data sample, The operations further comprise: Training the second modality imaging system based on the intermediate 3D point cloud features.
16. The non-transitory computer-readable medium according to claim 9, Wherein: The first output comprises a first plurality of bounding boxes corresponding to a first object in a scene, and The second output comprises a second plurality of bounding boxes corresponding to a second object in the scene, and Training the second modality imaging system comprises training the second modality imaging system using a subset of the first object and the second object, the subset of objects being in both the first plurality of bounding boxes and the second plurality of bounding boxes.
17. An image capture device, the image capture device Comprises: An image sensor; A memory storing processor-readable code; And At least one processor, the at least one processor being coupled to the memory and the image sensor, the at least one processor being configured to execute the processor-readable code to cause the at least one processor to perform operations including the following: Train a first modality imaging system; Receive a first input data sample and a second input data sample from the first modality imaging system and a second modality imaging system respectively, the first input data sample and the second input data sample being temporally synchronized; Process the first input data sample in the first modality imaging system to generate a first output; Process the second input data sample in the second modality imaging system to generate a second output; And Train the second modality imaging system based on the first output and the second output.
18. The image capture device according to claim 17, wherein training the first modality imaging system Comprises: Receive a third input data sample from the first modality imaging system; And determine a model for the first modality imaging system based on the first input data sample and a first ground truth corresponding to the first input data sample.
19. The image capture device according to claim 17, wherein the operations further Comprise: Receive a third input data sample from the second modality imaging system; And Process the third input data sample in the second modality imaging system to generate a third output based on a model for the second modality imaging system, the model being trained based on the first output of the first modality imaging system.
20. The image capture device according to claim 19, wherein the third output comprises at least one bounding box corresponding to an object detected in the third input data sample.
21. The image capture device according to claim 20, the operations further comprising operating a vehicle based on the at least one bounding box.
22. The image capture device according to claim 17, wherein the first modality imaging system comprises a LIDAR-based system and the second modality imaging system comprises a camera-based system.
23. The image capture device according to claim 17, Wherein: Processing the first input data sample in the first modality imaging system comprises determining intermediate 3D point cloud features based on the first input data sample, and Processing the second input data sample in the second modality imaging system comprises determining intermediate camera features based on the second input data sample, The operations further comprise: Train the second modality imaging system based on the intermediate 3D point cloud features.
24. The image capture device according to claim 17, Wherein: The first output comprises a first plurality of bounding boxes corresponding to a first object in a scene, and The second output comprises a second plurality of bounding boxes corresponding to a second object in the scene, and Training the second modality imaging system comprises training the second modality imaging system using a subset of the first object and the second object, the subset of objects being in both the first plurality of bounding boxes and the second plurality of bounding boxes.