High performance and low complexity adaptive video image defogging
By generating brightness distribution maps and using adaptive smoothing techniques based on Bezier curves, the problems of uneven brightness and high computational complexity in video image dehazing are solved, achieving a high-performance, low-complexity adaptive dehazing effect suitable for multi-camera systems in vehicles.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AMBARELLA INT LP
- Filing Date
- 2024-11-28
- Publication Date
- 2026-05-29
AI Technical Summary
Existing video image dehazing technologies struggle to achieve high-performance, low-complexity adaptive dehazing when faced with uneven fog conditions, resulting in uneven output video brightness and high computational complexity, making them unsuitable for real-time applications.
By configuring the interface and processor, a brightness distribution map is generated, multiple dehazing intensity weights are determined, adaptive smoothing is performed, and high-quality dehazed video frames are generated. A smooth transition is achieved using Bezier curves to adapt to the non-uniformity and changes of fog and reduce computational complexity.
It achieves adaptive control of defogging intensity in real-time environment, generates high-quality output images, reduces computational complexity and hardware resource requirements, and is suitable for multi-camera systems on vehicles.
Smart Images

Figure CN122115255A_ABST
Abstract
Description
Technical Field
[0001] This invention relates generally to video processing, and more specifically to methods and / or apparatus for achieving high-performance and low-complexity adaptive video image defogging. Background Technology
[0002] Video image processing is a rapidly evolving field. Various types of video and image processing can improve video quality, enhance details, and improve low-resolution video. With the continuous development of video image processing technology, expectations for video image quality in special environments are also gradually increasing. Particularly in vehicle applications, video image processing can provide enhanced driver assistance features. The image quality of dashcam lenses, vehicle surround-view cameras, and vehicle rearview mirrors is related to driving safety. However, in extreme weather conditions (e.g., foggy weather), image quality is affected. Fog and other distortions can cause large areas of blurring in video images, making it impossible to clearly see image details. Unclear images impair the driver's ability to observe road conditions. Foggy conditions can be one of the most dangerous driving situations. Therefore, video image defogging control has practical significance.
[0003] Dehazing video images is a challenging problem. Fog is typically uneven. Conventional dehazing techniques can result in uneven brightness in the output video. Conventional dehazing techniques cannot adapt to changing fog conditions. Conventional dehazing techniques are complex and computationally expensive, which can make them difficult to implement in real-time applications.
[0004] The goal is to achieve high-performance and low-complexity adaptive video image dehazing. Summary of the Invention
[0005] This invention relates to an apparatus including an interface and a processor. The interface can be configured to receive pixel data from the environment. The processor can be configured to: process the pixel data arranged as video frames; generate a luminance distribution map of the video frames in response to a low-pass filtering operation; determine a plurality of dehazing intensity weights for the luminance distribution map; perform adaptive smoothing on each of the plurality of dehazing intensity weights; and generate a dehazed video frame in response to the video frame and the plurality of dehazing intensity weights with adaptive smoothing. The plurality of dehazing intensity weights can each correspond to a luminance interval among a plurality of luminance intervals in the luminance distribution map. The adaptive smoothing can be configured to prevent luminance differences in the dehazed video frames. Attached Figure Description
[0006] Embodiments of the invention will become apparent from the following detailed description and the appended claims and drawings.
[0007] Figure 1 This is a diagram illustrating an example of a camera that can achieve high-performance and low-complexity adaptive video image dehazing according to an exemplary embodiment of the present invention.
[0008] Figure 2 This is a diagram showing an example edge device camera.
[0009] Figure 3 This is a diagram illustrating an example embodiment of the invention configured to provide a full-around view of a vehicle.
[0010] Figure 4 This is a block diagram showing the camera system.
[0011] Figure 5 This is a block diagram illustrating the operation for high-performance and low-complexity adaptive video image dehazing.
[0012] Figure 6 This is a diagram showing an example input video frame in a foggy environment.
[0013] Figure 7 This is a diagram showing the region of the low-frequency layer of the input video frame used for the luminance distribution map.
[0014] Figure 8 This is a graph showing the brightness values used in the brightness distribution map.
[0015] Figure 9 This is a diagram showing an example high-frequency layer of the input video frame.
[0016] Figure 10 This is a diagram showing an example output of a dehazed video frame.
[0017] Figure 11 This is a flowchart illustrating a method for providing high-performance and low-complexity adaptive video image dehazing.
[0018] Figure 12 This is a flowchart illustrating a method for determining the smoothing control strength value for a brightness interval.
[0019] Figure 13 This is a flowchart illustrating a method for setting the defogging strength.
[0020] Figure 14 This is a flowchart illustrating a method for generating dehazed video frames. Detailed Implementation
[0021] Embodiments of the present invention include providing high-performance and low-complexity adaptive video image dehazing, which can: (i) adapt to the non-uniformity of fog in an image, (ii) provide adaptive control of the dehazing intensity in a real-time environment, (iii) generate a high-quality output image, (iv) implement dehazing in real time with low complexity, (v) adjust the amount of regional dehazing of the image based on a brightness distribution map, (vi) prevent dehazing from causing large changes in brightness, (vii) remove blur caused by fog, (viii) provide a smooth transition for dehazing based on a Bezier curve, and / or (ix) be implemented as one or more integrated circuits.
[0022] Embodiments of the present invention can be configured to perform a dehazing operation on an input video frame. The dehazing operation can be configured to generate dehazed video frames. Dehazed video frames can be generated in real time to reduce the amount of fog present in the video frames. The reduction of fog can be adaptively controlled. For example, the amount of fog reduction in each video frame can be regionally determined based on its position and brightness distribution within the video frame. The adaptive control of dehazing can be configured to respond to the non-uniformity of fog in the real-time environment and / or changing fog conditions. For example, the dehazing operation can control the dehazing intensity based on changes in fog in the real-time environment to generate a high-quality dehazed output image.
[0023] Dehazing operations can be configured to provide adaptive video image dehazing with high performance and low complexity. High performance enables real-time execution of the dehazing operation. Low complexity enables the dehazing operation to be performed without expensive computation. Low complexity allows the camera to implement the dehazing operation in hardware, which may be inexpensive and / or limit power consumption and heat generation. For example, a vehicle can implement multiple cameras to provide a full-around view. Low complexity allows multiple cameras on a vehicle to perform dehazing operations independently. With high performance and low complexity, the use of hardware resources may be limited when performing dehazing operations with adaptive smoothing control.
[0024] Dehazing operations can include: processing input image data, extracting the brightness distribution at different locations in the image through filtering, and / or performing zoned dehazing control on the obtained brightness distribution at different locations. Different dehazing intensities can be set independently in different brightness regions to efficiently remove image blur caused by fog. Zoned dehazing control can be based on curve fitting. In one example, curve fitting could be a Bézier curve. Curve fitting enables smooth transitions and adaptive control of dehazing in different brightness regions to generate high-definition, high-quality video images.
[0025] The dehazing operation can be configured to divide the input image data into a high-frequency detail layer and a low-frequency brightness layer. A low-pass filtering operation can be performed on the input image data to generate the low-frequency layer. Based on the image location distribution of brightness in the low-frequency layer, a brightness distribution map can be determined. The brightness distribution map accurately represents the brightness distribution of the input image data. The brightness distribution map can be based on multiple rectangular regions of the low-frequency layer. For example, output statistics corresponding to a specific rectangular region can constitute a brightness distribution map.
[0026] Based on the brightness distribution map, the brightness distribution can be divided into multiple brightness intervals (e.g., a total of N intervals). These N brightness intervals can be ordered from darkest to brightest. Control points for defogging intensity can be set for each brightness interval. For example, the input parameter values for these control points can be used as weights for the defogging intensity.
[0027] In response to the selection of control weights for N dehazing intensity control points, an adaptive smoothing control can be determined for the dehazing operation. This adaptive smoothing control can implement a fitting control. For example, the implemented fitting control can be a Bézier curve. A Bézier curve can be a type of smoothing commonly used in graphics software. The smoothing performance of the Bézier curve can be set as the intensity weight of the dehazing control points. The fitting control can be implemented to ensure the smoothness of the brightness distribution of the image after dehazing. Ensuring the smoothness of the brightness distribution avoids large jumps in image brightness differences (e.g., preventing high contrast differences in local and / or adjacent regions in the output video frame). Generally, the higher the order of the implemented Bézier curve, the smoother the fitted curve. However, higher-order Bézier curves may increase the complexity of implementing the dehazing operation (e.g., using more hardware resources compared to lower-order Bézier curves). Embodiments of the present invention can balance the desired smoothness of the output with the hardware resource requirements.
[0028] Dehazing can be performed on the entire image brightness distribution based on N smoothing dehazing intensity control weights. Within the corresponding image brightness interval, the control weights can be fitted based on the brightness distribution map. The fitting of the control weights can be achieved in real time to obtain adaptive video image dehazing control based on image position and brightness distribution.
[0029] refer to Figure 1 The diagram illustrates an example of a camera capable of high-performance and low-complexity adaptive video image dehazing according to an exemplary embodiment of the present invention. A top view of zone 50 is shown. In the example shown, zone 50 may be an outdoor location. Streets, vehicles, and buildings are shown.
[0030] Devices 100a-100n are shown at various locations within zone 50. Devices 100a-100n can each implement an edge device. Edge devices 100a-100n may include intelligent IP cameras (e.g., camera systems). Edge devices 100a-100n may include low-power technologies designed for deployment in embedded platforms at the network edge (e.g., microprocessors running on sensors, cameras, or other battery-powered devices), where power consumption is a critical issue. In the examples, edge devices 100a-100n may include various traffic cameras and Intelligent Transportation System (ITS) solutions.
[0031] Edge devices 100a-100n can be implemented for a variety of applications. In the illustrated example, edge devices 100a-100n may include an Automatic License Plate Recognition (ANPR) camera 100a, a traffic camera 100b, a vehicle camera 100c, an access control camera 100d, an ATM camera 100e, a bullet camera 100f, a dome camera 100n, etc. In the example, edge devices 100a-100n can be implemented as a traffic camera and Intelligent Transportation System (ITS) solution designed to enhance road safety by utilizing a combination of people and vehicle detection, vehicle brand / model recognition, and Automatic License Plate Recognition (ANPR) functions.
[0032] In the illustrated example, zone 50 may be an outdoor location. In some embodiments, edge devices 100a-100n may be implemented at various indoor locations. In the example, edge devices 100a-100n may incorporate convolutional neural networks for use in security (surveillance) applications and / or access control applications. In the example, edge devices 100a-100n implemented as security cameras and access control applications may include battery-powered cameras, doorbell cameras, outdoor cameras, indoor cameras, etc. According to embodiments of the invention, security cameras and access control applications can realize performance benefits from the application of convolutional neural networks. In the example, edge devices utilizing convolutional neural networks according to embodiments of the invention can acquire large amounts of image data and perform on-device inference to obtain useful information (e.g., multiple temporal instances of images performed by each network), thereby reducing bandwidth and / or power consumption. In another example, security (surveillance) applications and / or location monitoring applications (e.g., tracking cameras) may benefit from large optical zoom. The design, type, and / or application performed by edge devices 100a-100n may vary depending on the design criteria of a particular implementation.
[0033] Camera systems 100a-100n can capture video in foggy environments within outdoor location area 50. For example, visibility may vary as the weather changes in outdoor location area 50. Visibility in outdoor location area 50 can change in real time. Even cameras in indoor locations can capture foggy conditions (e.g., a wet ice hockey rink may appear foggy). Each of camera systems 100a-100n can be configured to achieve high-performance and low-complexity adaptive video image dehazing.
[0034] refer to Figure 2 The diagram illustrates an example edge device camera. Camera systems 100a-100n are shown. Each camera device 100a-100n can have a different style and / or use case. For example, camera 100a can be an action camera, camera 100b can be a ceiling-mounted security camera, camera 100n can be a network camera, etc. Other types of cameras can be implemented (e.g., home security cameras, battery-powered cameras, doorbell cameras, stereo cameras, etc.). In some embodiments, camera systems 100a-100n can be fixed cameras (e.g., mounted and / or fixed in a single location). In some embodiments, camera systems 100a-100n can be handheld cameras. In some embodiments, camera systems 100a-100n can be configured for panning across areas, can be attached to mounts, gimbals, camera stands, etc. The design / style of cameras 100a-100n can vary depending on the design criteria of a particular implementation.
[0035] Each of the camera systems 100a-100n may include block (or circuit) 102, block (or circuit) 104, and / or block (or circuit) 106. Circuit 102 may implement a processor. Circuit 104 may implement a capture device. Circuit 106 may implement an inertial measurement unit (IMU). Camera systems 100a-100n may include other components (not shown). They can be used with... Figure 4 The components of the cameras 100a-100n are described in detail.
[0036] Processor 102 can be configured to implement an artificial neural network (ANN). In an example, the ANN may include a convolutional neural network (CNN). Processor 102 can be configured to implement a video encoder. Processor 102 can be configured to process pixel data arranged as video frames. Capture device 104 can be configured to capture pixel data that can be used by processor 102 to generate video frames. IMU 106 can be configured to generate motion data (e.g., vibration information, camera jitter, translation direction, etc.). In some embodiments, a structured light projector can be implemented to project a speckle pattern onto the environment. Capture device 104 can capture pixel data including a background image (e.g., the environment) with a speckle pattern. Although each of the cameras 100a-100n is shown as not implementing a structured light projector, some of the cameras 100a-100n may utilize a structured light projector (e.g., a camera implementing a sensor for capturing IR light).
[0037] Cameras 100a-100n can be edge devices. A processor 102 implemented by each of the cameras 100a-100n enables the cameras 100a-100n to perform various functions internally (e.g., at the local level). For example, processor 102 can be configured to perform object / event detection (e.g., computer vision operations), 3D reconstruction, presence detection, depth map generation, video encoding, electronic image stabilization, and / or video transcoding on the device. For example, processor 102 can even perform advanced processes such as computer vision and 3D reconstruction without uploading video data to a cloud service, thereby offloading computationally intensive functions (e.g., computer vision, video encoding, video transcoding, etc.).
[0038] In some embodiments, multiple camera systems may be implemented (e.g., camera systems 100a-100n may operate independently of each other). For example, each camera in 100a-100n may individually analyze the captured pixel data and perform event / object detection locally. In some embodiments, cameras 100a-100n may be configured as a camera network (e.g., security cameras that send video data to a central source such as a network-attached storage device and / or cloud service). The location and / or configuration of cameras 100a-100n may vary according to the design criteria of a particular implementation.
[0039] The capture device 104 of each camera in the camera systems 100a-100n may include a single lens (e.g., a monocular camera). The processor 102 may be configured to accelerate the preprocessing of speckle structured light for monocular 3D reconstruction. Monocular 3D reconstruction can be performed to generate depth maps and / or disparity maps without using a stereo camera.
[0040] refer to Figure 3 The diagram illustrates an example embodiment of the invention configured to provide a full-around view of a vehicle. An external environment 70 with vehicle 80 is shown. In the example shown, vehicle 80 may be a personal vehicle. In one example, vehicle 80 may be a commercial vehicle (e.g., a parcel delivery vehicle, service vehicle, public transport vehicle, etc.). In some embodiments, vehicle 80 may be a commercial truck (e.g., a semi-trailer truck). In some embodiments, vehicle 80 may be a pickup truck (e.g., a light vehicle, medium vehicle, heavy vehicle, etc.). In some embodiments, vehicle 80 may be a commuter vehicle and / or a family vehicle (e.g., a family vehicle such as a sedan, minivan, SUV, crossover, etc.). Vehicle 80 may be an internal combustion engine (ICE) vehicle, a diesel vehicle, a hybrid electric vehicle, a battery electric vehicle, etc. The type of vehicle 80 implemented may vary depending on the design criteria of a particular implementation.
[0041] External side mirrors 82a-82b are shown on vehicle 80. Side mirror 82a may be a driver-side side mirror of vehicle 80. Side mirror 82b may be a passenger-side side mirror of vehicle 80. Driver 90 is shown inside vehicle 80. Vehicle 80 may include devices 100a-100n. Devices 100a-100n may be camera systems. Camera systems 100a-100b are shown as integrated as part of vehicle 80. Camera system 100a is shown on the passenger side of vehicle 80. Camera system 100a is shown below passenger side mirror 82b. Camera system 100b is shown on the front grille of vehicle 80. In the view of the vehicle 80 shown, the three camera systems 100a-100b and 100e may be visible. However, one of the camera systems 100a-100n can be implemented horizontally below the driver's side mirror 82 (not visible from the perspective of the external view shown). Other camera systems 100a-100n can be positioned throughout the exterior and / or interior of the vehicle 80. The camera systems 100a-100n can be configured to capture a full-around view of the environment 70 surrounding the vehicle 80.
[0042] Dashed lines 92a-92e are shown. In the example shown, dashed line 92a is shown extending from camera system 100a, and dashed line 92b is shown extending from camera system 100b toward the exterior of the vehicle. Dashed lines 92c-92d can similarly extend from the corresponding camera systems 100c-100d (not visible from the shown viewpoint). Dashed lines 92a-92d can provide an illustrative representation of the field of view captured by each of the camera systems 100a-100d. Together, the fields of view 92a-92d can provide a full-around view of the environment surrounding vehicle 80.
[0043] All-around views 92a-92d are shown. In the example, all-around views 92a-92d can implement an all-around view (AVM) system. The AVM system may include four cameras (e.g., each camera may include one of camera systems 100a-100n and / or a combination of a pair of stereo lenses implemented by camera systems 100a-100n). In the viewpoint shown in environment 70, camera systems 100a and 100b may each be one of the four cameras, and the other two cameras may be invisible. In the example, camera system 100b may be a camera located on the front grille of vehicle 80, one of the cameras may be rearward (e.g., above the license plate), camera system 100a may be located below the passenger-side side mirror 82b, and one of the cameras may be located below the driver-side side mirror 82a. The arrangement of the cameras may vary depending on the design criteria of a particular implementation.
[0044] The dashed line 92e is shown extending from the camera system 100e toward the interior of the vehicle 80. The camera system 100e may be a passenger compartment surveillance camera system. The camera system 100e may be configured to capture a field of view 92e of the passenger compartment of the vehicle 80. The field of view 92e may be directed toward the driver 90. In some embodiments, the field of view 92e may be directed toward the driver 90 and / or other occupants of the vehicle 80.
[0045] In some embodiments, each of the camera systems 100a-100e may be configured to capture pixel data arranged as video frames. In some embodiments, each of the camera systems 100a-100d providing 360-degree views 92a-92d and / or the camera system 100e providing a cabin view may implement a fisheye lens (e.g., a 180-degree angular aperture may be used to capture video frames). The 360-degree views 92a-92d are shown as providing coverage of the field of view surrounding the vehicle 80. For example, a portion of 360-degree views 92a may provide coverage of the passenger side of the vehicle 80, a portion of 360-degree views 92b may provide coverage of the front of the vehicle 80, a portion of 360-degree views 92c may provide coverage of the driver side of the vehicle 80, and a portion of 360-degree views 92d may provide coverage of the rear of the vehicle 80. Each portion of the 360-degree views 92a-92d may be a field of view of a camera mounted to the vehicle 80. Each segment of the panoramic views 92a-92d can be de-distorted and stitched together by a video processor to provide enhanced video frames representing a top-down view of the vicinity of vehicle 80. Camera systems 100a-100d can be configured to implement a bird's-eye view transformer network (e.g., a deep learning model designed to generate BEV representations from multi-camera images). In this example, the panoramic views 92a-92d can be used to provide a bird's-eye view representation of vehicle 80.
[0046] Camera systems 100a-100e can provide representative examples of mechanisms for image acquisition. In one example, camera systems 100a-100e can be implemented as a monocular camera. In another example, camera systems 100a-100e can be implemented as stereo cameras (e.g., two capture devices implemented in a stereo pair). In some embodiments, the stereo cameras can be horizontally oriented. In some embodiments, the stereo cameras can be vertically oriented. In one example, four stereo cameras (e.g., eight capture devices) can be implemented, with one stereo camera on each side of vehicle 80. In some embodiments, camera systems 100a-100n can be installed as aftermarket products. For example, vehicle 80 can be sold without cameras, and one or more of camera systems 100n can be installed on vehicle 80. The implementation and / or location of camera systems 100a-100e on vehicle 80 and / or the orientation of camera systems 100a-100e can vary depending on the design criteria of a particular implementation.
[0047] Camera systems 100a-100d can capture foggy conditions in the external environment 70. For example, vehicle 80 may drive through weather conditions that may vary with visibility. Each of camera systems 100a-100e can be configured to achieve high-performance, low-complexity adaptive video image defogging. For cameras located outside vehicle 80 (e.g., camera systems 100a-100d), fog may affect the driver's visibility 90. Fog may affect the quality of images captured by external camera systems 100a-100d for use by various driver assistance systems.
[0048] refer to Figure 4 The diagram illustrates a block diagram of a camera system. The camera system (or device) 100 can be... Figure 2 The cameras 100a-100n shown in association and / or with Figure 3 Representative examples of cameras 100a-100e are shown in association. Camera system 100 may include processor / SoC 102, capture device 104, and IMU 106.
[0049] The camera system 100 may further include blocks (or circuits) 150, 152, 154, 156, 158, 160, 164, and / or 166. Circuit 150 may implement a memory. Circuit 152 may implement a battery. Circuit 154 may implement a communication device. Circuit 156 may implement a wireless interface. Circuit 158 may implement a general-purpose processor. Block 160 may implement an optical lens. Circuit 164 may implement one or more sensors. Circuit 166 may implement a human-machine interface (HID) device. In some embodiments, the camera system 100 may include a processor / SoC 102, a capture device 104, an IMU 106, a memory 150, a lens 160, a sensor 164, a battery 152, a communication module 154, a wireless interface 156, and a processor 158. In another example, camera system 100 may include a processor / SoC 102, a capture device 104, an IMU 106, a processor 158, a lens 160, and a sensor 164 as a single device, while memory 150, battery 152, communication module 154, and wireless interface 156 may be components of separate devices. Camera system 100 may include other components (not shown). The number, type, and / or arrangement of components in camera system 100 may vary depending on the design criteria of a particular implementation.
[0050] In some embodiments, processor 102 may be implemented as a video processor. In an example, processor 102 may be configured to receive three-sensor video input using a high-speed SLVS / MIPI-CSI / LVCMOS interface. In some embodiments, processor 102 may be configured to perform depth sensing in addition to generating video frames. In an example, depth sensing may be performed in response to depth information and / or vector light data captured in the video frames. In some embodiments, processor 102 may be implemented as a data stream vector processor. In an example, processor 102 may include a highly parallel architecture configured to perform image / video processing and / or radar signal processing.
[0051] Memory 150 can store data. Memory 150 can be implemented in various types of memory, including but not limited to cache, flash memory, memory card, random access memory (RAM), dynamic RAM (DRAM), etc. The type and / or size of memory 150 can vary according to the design criteria of a particular implementation. The data stored in memory 150 may correspond to video files, motion information (e.g., readings from sensor 164), video fusion parameters, image stabilization parameters, user input, computer vision models, feature sets, radar data cubes, radar detection and / or metadata information. In some embodiments, memory 150 can store reference images. Reference images can be used for computer vision operations, 3D reconstruction, automatic exposure, etc. In some embodiments, reference images may include reference structured light images.
[0052] Processor / SoC 102 can be configured to execute computer-readable code and / or procedural information. In various embodiments, the computer-readable code can be stored within processor / SoC 102 (e.g., microcode, etc.) and / or in memory 150. In one example, processor / SoC 102 can be configured to execute one or more artificial neural network models (e.g., face recognition CNN, object detection CNN, object classification CNN, 3D reconstruction CNN, presence detection CNN, etc.) stored in memory 150. In another example, memory 150 can store one or more directed acyclic graphs (DAGs) and one or more sets of weights or biases defining one or more artificial neural network models. In yet another example, memory 150 can store instructions for performing transformation operations (e.g., discrete cosine transform, discrete Fourier transform, fast Fourier transform, etc.). Processor / SoC 102 can be configured to receive input from memory 150 and / or present output to memory 150. Processor / SoC 102 can be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of the processor / SoC 102 can vary depending on the design criteria of a particular implementation. The processor / SoC 102 can be configured for low-power (e.g., battery) operation.
[0053] Battery 152 can be configured to store and / or supply power to components of camera system 100. The dynamic actuator mechanism for the rolling shutter sensor 130 can be configured to conserve power. Reduced power consumption allows camera system 100 to operate for extended periods using battery 152 without recharging. Battery 152 can be rechargeable. Battery 152 can be built-in (e.g., non-replaceable) or replaceable. Battery 152 can have an input for connection to an external power source (e.g., for charging). In some embodiments, device 100 can be powered by an external power source (e.g., battery 152 may not be implemented or may be implemented as a backup power source). Battery 152 can be implemented using various battery technologies and / or chemistry. The type of battery 152 implemented can vary depending on the design criteria of a particular implementation.
[0054] The communication module 154 can be configured to implement one or more communication protocols. For example, the communication module 154 and the wireless interface 156 can be configured to implement one or more of the following: IEEE 102.11, IEEE 102.15, IEEE 102.15.1, IEEE 102.15.2, IEEE 102.15.3, IEEE 102.15.4, IEEE 102.15.5, IEEE 102.20, and / or In some embodiments, the communication module 154 may be a hardwired data port (e.g., a USB port, a mini-USB port, a USB-C connector, an HDMI port, an Ethernet port, a DisplayPort interface, a Lightning port, etc.). In some embodiments, the wireless interface 156 may also implement one or more protocols associated with a cellular communication network (e.g., GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, SMS, etc.). In embodiments where the camera system 100 is implemented as a wireless camera, the protocols implemented by the communication module 154 and the wireless interface 156 may be wireless communication protocols. The type of communication protocol implemented by the communication module 154 may vary depending on the design criteria of the specific implementation.
[0055] Communication module 154 and / or wireless interface 156 can be configured to generate broadcast signals as output from camera system 100. The broadcast signals can transmit video data, parallax data, and / or (multiple) control signals to external devices. For example, the broadcast signals can be sent to cloud storage services (e.g., storage services capable of on-demand scaling). In some embodiments, communication module 154 may not transmit data until the processor / SoC 102 has performed video analysis and / or radar signal processing to determine that an object is in the field of view of camera system 100.
[0056] In some embodiments, the communication module 154 can be configured to generate a manual control signal. The manual control signal can be generated in response to a signal received from a user by the communication module 154. The manual control signal can be configured to activate the processor / SoC 102. The processor / SoC 102 can be activated in response to the manual control signal, regardless of the power state of the camera system 100.
[0057] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive a feature set. The received feature set may be used to detect events and / or objects. For example, the feature set may be used to perform computer vision operations. The feature set information may include instructions for the processor 102 to determine which types of objects correspond to objects of interest and / or events.
[0058] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive user input. User input allows the user to adjust operating parameters for various features implemented by the processor 102. In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to interact with an application (e.g., an app) (e.g., using an application programming interface (API)). For example, the application may be implemented on a smartphone to allow the end user to adjust various settings and / or parameters for various features implemented by the processor 102 (e.g., setting video resolution, selecting frame rate, selecting output format, setting tolerance parameters for 3D reconstruction, etc.).
[0059] Processor 158 can be implemented using general-purpose processor circuitry. Processor 158 may be operable to interact with video processing circuitry 102 and memory 150 to perform various processing tasks. Processor 158 may be configured to execute computer-readable instructions. In one example, the computer-readable instructions may be stored in memory 150. In some embodiments, the computer-readable instructions may include controller operations. Typically, input from sensor 164 and / or human-machine interface device 166 is shown to be received by processor 102. In some embodiments, general-purpose processor 158 may be configured to receive and / or analyze data from sensor 164 and / or HID 166 and make decisions in response to input. In some embodiments, processor 158 may send data to and / or receive data from other components of camera system 100, such as battery 152, communication module 154, and / or wireless interface 156. In some embodiments, processor 158 may implement an integrated digital signal processor (IDSP). For example, IDSP 158 may be configured to implement a warp engine. Which functions of the camera system 100 are executed by processor 102 and general-purpose processor 158 can vary depending on the design criteria of a particular implementation.
[0060] Lens 160 may be attached to capture device 104. Capture device 104 may be configured to receive an input signal (e.g., LIN) via lens 160. The signal LIN may be an optical input (e.g., an analog image). Lens 160 may be implemented as an optical lens. Lens 160 may provide zoom and / or focus features. In one example, capture device 104 and / or lens 160 may be implemented as a single lens assembly. In another example, lens 160 may be implemented separately from capture device 104.
[0061] Capture device 104 can be configured to convert input light LIN into computer-readable data. Capture device 104 can capture data received through lens 160 to generate raw pixel data. In some embodiments, capture device 104 can capture data received through lens 160 to generate a bitstream (e.g., generate video frames). For example, capture device 104 can receive focused light from lens 160. Lens 160 can be oriented, tilted, panned, scaled, and / or rotated to provide a target view from camera system 100 (e.g., a view of video frames, a view of panoramic video frames captured using multiple camera systems 100a-100n, a target image and reference image view for stereo vision, etc.). Capture device 104 can generate a signal (e.g., VIDEO). The signal VIDEO can be pixel data (e.g., a sequence of pixels that can be used to generate video frames). In some embodiments, the signal VIDEO can be video data (e.g., a sequence of video frames). The signal VIDEO can be presented to one of the inputs of processor 102. In some embodiments, the pixel data generated by the capture device 104 may be uncompressed and / or raw data generated in response to focused light from the lens 160. In some embodiments, the output of the capture device 104 may be a digital video signal.
[0062] In this example, capture device 104 may include block (or circuitry) 180, block (or circuitry) 182, and block (or circuitry) 184. Circuitry 180 may be an image sensor. Circuitry 182 may be a processor and / or logic unit. Circuitry 184 may be memory circuitry (e.g., a frame buffer). Lens 160 (e.g., a camera lens) may be oriented to provide a view of the environment surrounding camera system 100. Lens 160 may be designed to capture environmental data (e.g., light input LIN). Lens 160 may be a wide-angle lens and / or a fisheye lens (e.g., a lens capable of capturing a wide field of view). Lens 160 may be configured to capture and / or focus light for capture device 104. Typically, image sensor 180 is located behind lens 160. Based on the light captured from lens 160, capture device 104 may generate bitstream and / or video data (e.g., a signal VIDEO).
[0063] Capture device 104 can be configured to capture video image data (e.g., light collected and focused by lens 160). Capture device 104 can capture data received through lens 160 to generate a video bitstream (e.g., pixel data for a sequence of video frames). In various embodiments, lens 160 can be implemented as a fixed-focus lens. Fixed-focus lenses are generally advantageous for smaller size and lower power consumption. In examples, fixed-focus lenses can be used in battery-powered, doorbell, and other low-power camera applications. In some embodiments, lens 160 can be oriented, tilted, panned, zoomed, and / or rotated to capture the environment around camera system 100 (e.g., capturing data from the field of view). In examples, professional camera models can utilize active lens systems for enhanced functionality, remote control, etc.
[0064] The capture device 104 can convert received light into a digital data stream. In some embodiments, the capture device 104 can perform analog-to-digital conversion. For example, the image sensor 180 can perform photoelectric conversion on the light received by the lens 160. The processor / logic unit 182 can convert the digital data stream into a video data stream (or bitstream), a video file, and / or multiple video frames. In this example, the capture device 104 can present the video data as a digital video signal (e.g., VIDEO). The digital video signal can include video frames (e.g., continuous digital images and / or audio). In some embodiments, the capture device 104 can include a microphone for capturing audio. In some embodiments, the microphone can be implemented as a separate component (e.g., one of the sensors 164).
[0065] Video data captured by capture device 104 can be represented as a signal / bitstream / data VIDEO (e.g., a digital video signal). Capture device 104 can present the signal VIDEO to processor / SoC 102. The signal VIDEO can represent video frames / video data. The signal VIDEO can be a video stream captured by capture device 104. In some embodiments, the signal VIDEO may include pixel data operable by processor 102 (e.g., a video processing pipeline, image signal processor (ISP), etc.). Processor 102 can generate video frames in response to the pixel data in the signal VIDEO.
[0066] The signal VIDEO may include pixel data arranged as video frames. In some embodiments, the signal VIDEO may be an image including a background (e.g., the captured environment and / or objects) and a speckle pattern generated by a structured light projector. The signal VIDEO may include a single-channel source image. A single-channel source image may be generated in response to capturing pixel data using a monocular lens 160.
[0067] Image sensor 180 can receive input light LIN from lens 160 and convert the light LIN into digital data (e.g., a bitstream). For example, image sensor 180 can perform photoelectric conversion on light from lens 160. In some embodiments, image sensor 180 may have an additional margin that is not used as part of the image output. In some embodiments, image sensor 180 may not have an additional margin. In various embodiments, image sensor 180 can be implemented as an RGB sensor, RGB-IR sensor, RCCB sensor, monocular image sensor, stereo image sensor, thermal sensor, event-based sensor, etc. For example, image sensor 180 can be any type of sensor configured to provide sufficient output for a computer vision operation (e.g., neural network-based detection) to be performed on the output data. In the context of the illustrated embodiments, image sensor 180 can be configured to generate an RGB-IR video signal. In a field of view illuminated only by infrared light, image sensor 180 can generate a monochrome (B / W) video signal. In a field of view illuminated by both IR and visible light, image sensor 180 can be configured to generate color information in addition to generating a monochrome video signal. In various embodiments, the image sensor 180 may be configured to generate video signals in response to visible light and / or infrared (IR) light.
[0068] In some embodiments, camera sensor 180 may include a rolling shutter sensor or a global shutter sensor. In an example, the rolling shutter sensor 180 may implement an RGB-IR sensor. In some embodiments, capture device 104 may include a rolling shutter IR sensor and an RGB sensor (e.g., implemented as separate components). In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR rolling shutter complementary metal-oxide-semiconductor (CMOS) image sensor. In one example, the rolling shutter sensor 180 may be configured to assert a signal indicating the exposure time of the first row. In one example, the rolling shutter sensor 180 may apply a mask to a monochrome sensor. In an example, the mask may include multiple cells containing a red pixel, a green pixel, a blue pixel, and an IR pixel. The IR pixel may contain red, green, and blue filter materials that efficiently absorb all light in the visible spectrum while allowing longer infrared wavelengths to pass through with minimal loss. In the case of a rolling shutter, as each row (or line) of the sensor begins exposure, all pixels in that row (or line) may begin exposure simultaneously.
[0069] Processor / logic unit 182 can convert the bitstream into human-readable content (e.g., video data that an average person can understand regardless of image quality, such as video frames and / or pixel data that can be converted into video frames by processor 102). For example, processor / logic unit 182 can receive raw (e.g., raw) data from image sensor 180 and generate (e.g., encode) video data (e.g., bitstream) based on the raw data. Capture device 104 may have memory 184 to store raw data and / or processed bitstreams. For example, capture device 104 may implement frame memory and / or buffer 184 to store (e.g., provide temporary storage and / or cache) one or more video frames (e.g., digital video signals). In some embodiments, processor / logic unit 182 can perform analysis and / or correction on the video frames stored in memory / buffer 184 of capture device 104. Processor / logic unit 182 can provide status information about the captured video frames.
[0070] IMU 106 can be configured to detect motion and / or movement of camera system 100. IMU 106 is shown as receiving signals (e.g., MTN). Signal MTN may include a combination of forces acting on camera system 100. Signal MTN may include movement, vibration, jitter, translation, abrupt changes, etc. Signal MTN may represent movement in three-dimensional space (e.g., movement in the X, Y, and Z directions). The type and / or amount of motion received by IMU 106 may vary depending on the design criteria of a particular implementation.
[0071] IMU 106 may include block (or circuitry) 186. Circuitry 186 may implement a motion sensor. In one example, motion sensor 186 may be a gyroscope. Gyroscope 186 may be configured to measure the amount of movement. For example, gyroscope 186 may be configured to detect the amount and / or direction of movement of signal MTN and convert that movement into electrical data. IMU 106 may be configured to determine the amount and / or direction of movement measured by gyroscope 186. IMU 106 may convert the electrical data from gyroscope 186 into a format readable by processor 102. IMU 106 may be configured to generate a signal (e.g., M_INFO). Signal M_INFO may include measurement information in a format readable by processor 102. IMU 106 may present signal M_INFO to processor 102. The number, type, and / or arrangement of components of IMU 106 and / or the number, type, and / or functionality of signals communicated by IMU 106 may vary according to design criteria for a particular implementation.
[0072] Sensor 164 can implement multiple sensors, including but not limited to motion sensors, ambient light sensors, proximity sensors (e.g., ultrasonic, radar, passive infrared, lidar, etc.), audio sensors (e.g., microphones), etc. In embodiments implementing a motion sensor, sensor 164 can be configured to detect motion anywhere (or some location outside the field of view) monitored by camera system 100. In various embodiments, motion detection can be used as a threshold for activating capture device 104. Sensor 164 can be implemented as an internal component of camera system 100 and / or an external component of camera system 100. In one example, sensor 164 can be implemented as a passive infrared (PIR) sensor. In another example, sensor 164 can be implemented as a smart motion sensor. In yet another example, sensor 164 can be implemented as a microphone. In embodiments implementing a smart motion sensor, sensor 164 may include a low-resolution image sensor configured to detect motion and / or people.
[0073] In various embodiments, sensor 164 may generate signals (e.g., SENS). The SENS may include various data (or information) collected by sensor 164. In an example, the SENS may include data collected in response to motion detected in a monitored field of view, ambient light levels in the monitored field of view, and / or sound picked up in the monitored field of view. However, other types of data may be collected and / or generated based on application-specific design criteria. The SENS may be presented to processor / SoC 102. In an example, sensor 164 may generate (assert) the SENS when motion is detected in a field of view monitored by camera system 100. In another example, sensor 164 may generate (assert) the SENS when audio is triggered in a field of view monitored by camera system 100. In yet another example, sensor 164 may be configured to provide directional information about motion and / or sound detected in the field of view. This directional information may also be communicated to processor / SoC 102 via the SENS.
[0074] HID 166 can implement an input device. For example, HID 166 can be configured to receive human input. In one example, HID 166 can be configured to receive password input from a user. In another example, HID 166 can be configured to receive user input to provide various parameters and / or settings to processor 102 and / or memory 150. In some embodiments, camera system 100 may include a keypad, touchpad (or screen), doorbell switch, and / or other human-machine interface device (HID) 166. In an example, sensor 164 can be configured to determine when an object approaches HID 166. In an example where camera system 100 is implemented as part of an access control application, capture device 104 can be activated to provide images for identifying people attempting access, and the lock zone and / or illumination for access touchpad 166 can be turned on. For example, a combination of input from HID 166 (e.g., a password or PIN code) can be combined with presence determination and / or depth analysis performed by processor 102 to achieve two-factor authentication. HID 166 can present a signal (e.g., USR) to processor 102. The signal USR may include input received by HID 166.
[0075] In an embodiment of the camera system 100 implementing a structured light projector, the structured light projector may include a structured light pattern lens and / or a structured light source. The structured light source may be configured to generate a structured light pattern signal (e.g., a speckle pattern) that can be projected onto an environment near the camera system 100. The structured light pattern may be captured by a capture device 104 as part of a light input LIN. The structured light pattern lens may be configured to emit structured light generated by the structured light source of the structured light projector while protecting the structured light source. The structured light pattern lens may be configured to decompose the laser pattern generated by the structured light source into a pattern array (e.g., a dense dot pattern array for a speckle pattern).
[0076] In the example, the structured light source can be implemented as an array of vertical-cavity surface-emitting lasers (VCSELs) and a lens. However, other types of structured light sources can be implemented to meet the design criteria of specific applications. In the example, the VCSEL array is generally configured to generate laser patterns (e.g., signal SLP). The lens is generally configured to decompose the laser pattern into an array of dense point patterns. In the example, the structured light source can be implemented as a near-infrared (NIR) source. In various embodiments, the light source of the structured light source can be configured to emit light at a wavelength of approximately 940 nanometers (nm), which is invisible to the human eye. However, other wavelengths can be utilized. In the example, wavelengths in the range of approximately 800 nm to 1000 nm can be utilized.
[0077] Processor / SoC 102 can receive signals VIDEO, M_INFO, SENS, and USR. Processor / SoC 102 can generate one or more video output signals (e.g., VIDOUT), one or more control signals (e.g., CTRL), one or more depth data signals (e.g., DIMAGES), and / or one or more twist table data signals (e.g., WT) based on signals VIDEO, M_INFO, SENS, USR, and / or other inputs. In some embodiments, signals VIDOUT, DIMAGES, WT, and CTRL can be generated based on analysis of signals VIDEO and / or objects detected in signals VIDEO. In some embodiments, signals VIDOUT, DIMAGES, WT, and CTRL can be generated based on analysis of signals VIDEO, motion information captured by IMU 106, and / or inherent properties of lens 160 and / or capture device 104.
[0078] In various embodiments, the processor / SoC 102 may be configured to perform one or more of the following: feature extraction, object detection, object tracking, electronic image stabilization, 3D reconstruction, presence detection, and object identification. For example, the processor / SoC 102 may determine motion information and / or depth information by analyzing frames from a signal VIDEO and comparing those frames with previous frames. The comparison may be used to perform digital motion estimation. In some embodiments, the processor / SoC 102 may be configured to generate a video output signal VIDOUT, a distortion table data signal WT, and / or a depth data signal DIMAGES, including a disparity map and a depth map, based on the signal VIDEO. The video output signal VIDOUT, the distortion table data signal WT, and / or the depth data signal DIMAGES may be presented to memory 150, communication module 154, and / or wireless interface 156. In some embodiments, the video signal VIDOUT, the distortion table data signal WT, and / or the depth data signal DIMAGES may be used internally by the processor 102 (e.g., not presented as output). In one example, the warped table data signal WT can be used by a warping engine implemented by a digital signal processor (e.g., processor 158).
[0079] The signal VIDOUT may be presented to the communication module 154 and / or the wireless interface 156. In some embodiments, the signal VIDOUT may include encoded video frames generated by the processor 102. In some embodiments, the encoded video frames may include a complete video stream (e.g., encoded video frames representing all video captured by the capture device 104). The encoded video frames may be encoded, cropped, stitched, stabilized, and / or enhanced versions of pixel data received from the signal VIDEO. In an example, the encoded video frames may be high-resolution, digital, encoded, de-distorted, stabilized, cropped, blended, stitched, and / or rolled shutter effect corrected versions of the signal VIDEO.
[0080] In some embodiments, the signal VIDOUT can be generated based on video analysis (e.g., computer vision operations) performed by processor 102 on generated video frames. Processor 102 can be configured to perform computer vision operations to detect objects and / or events in the video frames, and then convert the detected objects and / or events into statistics and / or parameters. In one example, the data determined by the computer vision operations can be converted by processor 102 into a human-readable format. The data from the computer vision operations can be used to detect objects and / or events. The computer vision operations can be performed locally by processor 102 (e.g., without needing to communicate with external devices to offload computational operations). Similarly, other video processing and / or encoding operations (e.g., stabilization, compression, stitching, cropping, rolling shutter effect correction, etc.) can be performed locally by processor 102. For example, locally executed computer vision operations allow the computer vision operations to be performed by processor 102 and avoid heavy video processing running on a backend server. Avoiding video processing running on a backend (e.g., remotely located) server can protect privacy.
[0081] In some embodiments, the signal VIDOUT can be data generated by processor 102 (e.g., video analysis results, audio / speech analysis results, stabilized video frames, etc.) that can be communicated to cloud computing services to aggregate information and / or provide training data for machine learning (e.g., to improve object detection, improve audio detection, improve presence detection, etc.). In some embodiments, the signal VIDOUT can be provided to cloud services for mass storage (e.g., to enable users to retrieve encoded video using smartphones and / or desktop computers). In some embodiments, the signal VIDOUT can include data extracted from video frames (e.g., computer vision results) and can communicate the results to another device (e.g., a remote server, cloud computing system, etc.) to offload the analysis of the results to another device (e.g., offloading the analysis of the results to a cloud computing service instead of performing all analysis locally). The type of information communicated by the signal VIDOUT can vary depending on the design criteria of a particular implementation.
[0082] The CTRL signal can be configured to provide a control signal. The CTRL signal can be generated in response to a decision made by the processor 102. In one example, the CTRL signal can be generated in response to a detected object and / or features extracted from a video frame. The CTRL signal can be configured to enable, disable, or change the operating mode of another device. In one example, the CTRL signal can be used to lock / unlock a door controlled by an electronic lock. In another example, the CTRL signal can be used to put a device into sleep mode (e.g., low power mode) and / or activate it from sleep mode. In yet another example, the CTRL signal can be used to generate an alarm and / or notification. The type of device controlled by the CTRL signal and / or the response performed by the device in response to the CTRL signal can vary depending on the design criteria of the specific implementation.
[0083] A signal CTRL can be generated based on data received by sensor 164 (e.g., temperature readings, motion sensor readings, etc.). A signal CTRL can be generated based on input from HID 166. A signal CTRL can be generated based on human behavior detected by processor 102 in a video frame. A signal CTRL can be generated based on the type of object detected (e.g., person, animal, vehicle, etc.). A signal CTRL can be generated in response to detecting a specific type of object at a specific location. A signal CTRL can be generated in response to user input to provide various parameters and / or settings to processor 102 and / or memory 150. Processor 102 can be configured to generate a signal CTRL in response to sensor fusion operations (e.g., aggregating information received from different sources). Processor 102 can be configured to generate a signal CTRL in response to the result of an presence detection performed by processor 102. The conditions used to generate the signal CTRL can vary depending on the design criteria of a particular implementation.
[0084] Signals DIMAGES may include one or more depth maps and / or disparity maps generated by processor 102. Signals DIMAGES may be generated in response to 3D reconstruction performed on a monocular single-channel image. Signals DIMAGES may be generated in response to analysis of captured video data and structured light patterns.
[0085] A multi-step approach can be implemented to activate and / or disable the capture device 104 based on the output of the motion sensor 164 and / or any other power consumption characteristics of the camera system 100, thereby reducing the power consumption of the camera system 100 and extending the lifespan of the battery 152. The motion sensor in sensor 164 may have low power consumption on the battery 152 (e.g., less than 10W). In the example, the motion sensor in sensor 164 may be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 102. Video analysis performed by the processor / SoC 102 may have relatively high power consumption on the battery 152 (e.g., greater than that of the motion sensor 164). In the example, the processor / SoC 102 may be in a low-power state (or powered down) until some motion is detected by the motion sensor in sensor 164.
[0086] Camera system 100 can be configured to operate using various power states. For example, in a power-down state (e.g., sleep state, low power state), the motion sensor in sensor 164 and processor / SoC 102 can be turned on, and other components of camera system 100 (e.g., image capture device 106, memory 150, communication module 154, etc.) can be turned off. In another example, camera system 100 can operate in an intermediate state. In the intermediate state, image capture device 106 can be turned on, and memory 150 and / or communication module 154 can be turned off. In yet another example, camera system 100 can operate in a power-on (or high-power) state. In the power-on state, sensor 164, processor / SoC 102, capture device 104, memory 150, and / or communication module 154 can be turned on. Camera system 100 can consume some power from battery 152 (e.g., a relatively small and / or minimal amount of power) in the power-down state. In the power-on state, camera system 100 can consume more power from battery 152. The number of components of the camera system 100 that are turned on when the camera system 100 is in power state and / or operating in each power state can vary according to the design criteria of a particular implementation.
[0087] In some embodiments, the camera system 100 may be implemented as a system-on-a-chip (SoC). For example, the camera system 100 may be implemented as a printed circuit board including one or more components. The camera system 100 may be configured to perform intelligent video analysis on video frames of a video. The camera system 100 may be configured to crop and / or enhance the video.
[0088] In some embodiments, the video frame may be a view (or a derivative of a view) captured by the capture device 104. Pixel data signals may be enhanced by the processor 102 (e.g., color conversion, noise filtering, automatic exposure, automatic white balance, automatic focus, etc.). In some embodiments, the video frame may provide a series of cropped and / or enhanced video frames that improve the view from the perspective of the camera system 100 (e.g., providing night vision, providing high dynamic range (HDR) imaging, providing more viewing area, highlighting detected objects, providing additional data (e.g., digital distance to the detected object), etc.) to enable the processor 102 to see the location better than a human can see with human vision.
[0089] Encoded video frames can be processed locally. In one example, the encoded video can be stored locally by memory 150, enabling processor 102 to facilitate computer vision analysis internally (e.g., without first uploading the video frames to a cloud service). Processor 102 can be configured to select video frames to be encapsulated into a video stream that can be transmitted over a network (e.g., a bandwidth-limited network).
[0090] In some embodiments, processor 102 may be configured to perform sensor fusion operations. The sensor fusion operations performed by processor 102 may be configured to analyze information from multiple sources (e.g., capture device 104, IMU 106, sensor 164, and HID 166). By analyzing various data from different sources, sensor fusion operations may be able to make inferences about the data that might not be possible from just one data source. For example, sensor fusion operations implemented by processor 102 may analyze video data (e.g., mouth movements) as well as speech patterns from directional audio. Different sources can be used to develop scene models to support decision-making. For example, processor 102 may be configured to compare the synchronization of detected speech patterns with mouth movements in video frames to determine which person is speaking in the video frame. Sensor fusion operations may also provide temporal correlation, spatial correlation, and / or reliability between the received data.
[0091] In some embodiments, processor 102 may implement convolutional neural network (CNN) capabilities. CNN capabilities can be implemented using deep learning techniques for computer vision. CNN capabilities can be configured to perform pattern and / or image recognition using a training process involving multi-layer feature detection. Computer vision and / or CNN capabilities can be executed locally by processor 102. In some embodiments, processor 102 may receive training data and / or feature set information from external sources. For example, external devices (e.g., cloud services) may access various data sources to provide training data that camera system 100 may not be able to obtain. However, computer vision operations performed using feature sets can be performed using the computational resources of processor 102 within camera system 100.
[0092] The video pipeline of processor 102 can be configured to perform local dewarping, cropping, enhancement, rolling shutter correction, stabilization, downsizing, encapsulation, compression, conversion, mixing, synchronization, and / or other video operations. The video pipeline of processor 102 can support multiple streams (e.g., generating multiple bitstreams in parallel, each including a different bitrate). In the example, the video pipeline of processor 102 can implement an image signal processor (ISP) with an input pixel rate of 320 Mbps. The architecture of the video pipeline of processor 102 enables real-time and / or near-real-time video operations on high-resolution video and / or high-bitrate video data. The video pipeline of processor 102 can perform computer vision processing, stereo vision processing, object detection, 3D noise reduction, fisheye lens correction (e.g., real-time 360-degree dewarping and lens distortion correction), oversampling, and / or high dynamic range processing on 4K resolution video data. In one example, the architecture of the video pipeline can achieve 4K ultra-high resolution with H.264 encoding at twice the real-time speed (e.g., 60fps) and 4K ultra-high resolution with H.265 / HEVC and / or 4K AVC encoding at 30fps (e.g., 4KP30 AVC and HEVC encoding with multi-stream support). The type of video operation and / or the type of video data operated on by the processor 102 can vary according to the design criteria of the specific implementation.
[0093] The camera sensor 180 can be a high-resolution sensor. Using the high-resolution sensor 180, the processor 102 can combine oversampling of the image sensor 180 with digital scaling within the cropped area. Oversampling and digital scaling can each be one of the video operations performed by the processor 102. Oversampling and digital scaling can be implemented to provide a higher resolution image within the total size limit of the cropped area.
[0094] In some embodiments, lens 160 may be a fisheye lens. One of the video operations implemented by processor 102 may be a dedistortion operation. Processor 102 may be configured to dedistort generated video frames. Dedistortion may be configured to reduce and / or remove severe distortion caused by fisheye lens and / or other lens characteristics. For example, dedistortion may reduce and / or eliminate bulging effects to provide a linear image.
[0095] Processor 102 can be configured to crop (e.g., trim) a region of interest from a full video frame (e.g., generate a region of interest video frame). Processor 102 can generate video frames and select regions. In the example, cropping the region of interest can generate a second image. The cropped image (e.g., the region of interest video frame) can be smaller than the original video frame (e.g., the cropped image can be a portion of the captured video).
[0096] The region of interest (ROI) can be dynamically adjusted based on the location of the audio source. For example, the detected audio source may be moving, and its location may shift as video frames are captured. Processor 102 can update the coordinates of the selected ROI and dynamically update the cropped portion (e.g., a directional microphone implemented as one or more sensors in sensor 164 can dynamically update its location based on captured directional audio). The cropped portion may correspond to the selected ROI. As the ROI changes, the cropped portion may change. For example, the selected coordinates of the ROI may change from frame to frame, and processor 102 may be configured to crop the selected region in each frame.
[0097] Processor 102 can be configured to oversample image sensor 180. Oversampling of image sensor 180 can produce a higher resolution image. Processor 102 can be configured to digitally upscale regions of video frames. For example, processor 102 can digitally upscale a cropped region of interest. For example, processor 102 can establish a region of interest based on directional audio, crop the region of interest, and then digitally upscale the cropped region of interest video frame.
[0098] The dedistortion operation performed by processor 102 can adjust the visual content of the video data. The adjustment performed by processor 102 can make the visual content appear natural (e.g., as if seen by a person viewing a position corresponding to the field of view of capture device 104). In the example, dedistortion can alter the video data to generate linear video frames (e.g., correcting artifacts caused by lens characteristics of lens 160). Dedistortion operations can be implemented to correct distortion caused by lens 160. Adjusted visual content can be generated to achieve more accurate and / or reliable object detection.
[0099] Various features (e.g., de-distortion, digital scaling, cropping, etc.) can be implemented as hardware modules in processor 102. Implementing hardware modules can increase the video processing speed of processor 102 (e.g., faster than software implementations). Hardware implementations enable video processing while reducing latency. The hardware components used can vary depending on the design standards of the specific implementation.
[0100] In some embodiments, processor 102 may implement one or more coprocessors, cores, and / or chiplets. For example, processor 102 may implement one coprocessor configured as a general-purpose processor and another coprocessor configured as a video processor. In some embodiments, processor 102 may be a dedicated hardware module designed to perform a specific task. In one example, processor 102 may implement an AI accelerator. In another example, processor 102 may implement a radar processor. In yet another example, processor 102 may implement a dataflow vector processor. In some embodiments, other processors implemented by device 100 may be general-purpose processors and / or video processors (e.g., coprocessors that are physically different from processor 102 in chipset and / or silicon). In one example, processor 102 may implement the x86-64 instruction set. In another example, processor 102 may implement the ARM instruction set. In yet another example, processor 102 may implement the RISC-V instruction set. The number of cores, coprocessors, design optimizations, and / or instruction sets implemented by processor 102 may vary depending on the design criteria of a particular implementation.
[0101] The processor 102 shown includes multiple blocks (or circuits) 190a-190n. Blocks 190a-190n can implement various hardware modules implemented by the processor 102. Hardware modules 190a-190n can be configured to provide various hardware components to implement video processing pipelines, radar signal processing pipelines, and / or AI processing pipelines. Circuits 190a-190n can be configured to receive pixel data VIDEO, generate video frames based on the pixel data, perform various operations on the video frames (e.g., de-distortion, rolling shutter correction, cropping, magnification, image stabilization, 3D reconstruction, presence detection, automatic exposure, etc.), prepare video frames for communication with external hardware (e.g., encoding, encapsulation, color correction, etc.), parse feature sets, and implement various computer vision operations (e.g., object detection, segmentation, classification, etc.). Hardware modules 190a-190n can be configured to implement various security features (e.g., secure boot, I / O virtualization, etc.). Various implementations of processor 102 may not necessarily utilize all the features of hardware modules 190a-190n. The features and / or functions of hardware modules 190a-190n may vary depending on the design criteria of a particular implementation. Details of hardware modules 190a-190n may be described in connection with U.S. Patent Application No. 16 / 831,549, filed April 16, 2020; U.S. Patent Application No. 16 / 288,922, filed February 28, 2019; U.S. Patent Application No. 15 / 593,493 (now U.S. Patent No. 10,437,600), filed May 12, 2017; U.S. Patent Application No. 15 / 931,942, filed May 14, 2020; U.S. Patent Application No. 16 / 991,344, filed August 12, 2020; and U.S. Patent Application No. 17 / 479,034, filed September 20, 2021, the appropriate portions of which are incorporated herein by reference in their entirety.
[0102] Hardware modules 190a-190n can be implemented as dedicated hardware modules. Compared to software implementation, using dedicated hardware modules 190a-190n to implement various functions of processor 102 allows processor 102 to be highly optimized and / or customized to limit power consumption, reduce heat generation, and / or increase processing speed. Hardware modules 190a-190n can be customizable and / or programmable to implement multiple types of operations. Implementing dedicated hardware modules 190a-190n allows the hardware used to perform each type of computation to be optimized for speed and / or efficiency. For example, hardware modules 190a-190n can implement several relatively simple operations frequently used in computer vision operations, which together enable computer vision operations to be performed in real time. The video pipeline can be configured to identify objects. Objects can be identified by interpreting numerical and / or symbolic information to determine the specific type and / or characteristics of the object represented by the visual data. For example, the number of pixels and / or the color of pixels in video data can be used to identify portions of video data as objects. Hardware modules 190a-190n enable computationally intensive operations (e.g., computer vision operations, video encoding, video transcoding, 3D reconstruction, depth map generation, presence detection, etc.) to be performed locally by the camera system 100.
[0103] One of the hardware modules 190a-190n (e.g., 190a) can implement a scheduler circuit. Scheduler circuit 190a can be configured to store a directed acyclic graph (DAG). In the example, scheduler circuit 190a can be configured to generate and store a DAG in response to received (e.g., loaded) feature set information. The DAG can define video operations to be performed to extract data from video frames. For example, the DAG can define various mathematical weights (e.g., neural network weights and / or biases) to be applied when performing computer vision operations to classify various groups of pixels into specific objects.
[0104] Scheduler circuit 190a can be configured to parse acyclic graphs to generate various operators. Operators can be scheduled by scheduler circuit 190a in one or more of the other hardware modules 190a-190n. For example, one or more hardware modules 190a-190n can implement a hardware engine configured to perform a specific task (e.g., a hardware engine designed to perform repetitive specific mathematical operations used for performing computer vision operations). Scheduler circuit 190a can schedule operators based on when they are ready to be processed by hardware engines 190a-190n.
[0105] Scheduler circuit 190a can time-multiplex tasks to hardware modules 190a-190n for execution based on the availability of hardware modules 190a-190n. Scheduler circuit 190a can parse a directed acyclic graph into one or more data streams. Each data stream can include one or more operators. Once the directed acyclic graph is parsed, scheduler circuit 190a can allocate data streams / operators to hardware engines 190a-190n and send relevant operator configuration information to initiate the operators.
[0106] The binary representation of each directed acyclic graph can be an ordered traversal of the directed acyclic graph, where descriptors and operators are interleaved based on data dependencies. Descriptors typically provide registers that link data buffers to specific operands in a dependent operator. In various embodiments, operators may not appear in the directed acyclic graph representation until all dependent descriptors have been declared for operands.
[0107] One of the hardware modules 190a-190n (e.g., 190b) can implement an Artificial Neural Network (ANN) module. The ANN module can be implemented as a fully connected neural network or a Convolutional Neural Network (CNN). In this example, the fully connected network is "structure-agnostic" because no special assumptions need to be made about the input. A fully connected neural network consists of a series of fully connected layers that connect each neuron in one layer to each neuron in another layer. In a fully connected layer, for n inputs and m outputs, there are n*m weights. There are also bias values for each output node, resulting in a total of (n+1)*m parameters. In a trained neural network, (n+1)*m parameters have been determined during the training process. A trained neural network typically includes an architectural specification and a set of parameters (weights and biases) determined during training. In another example, a CNN architecture may explicitly assume that the input is an image, enabling the encoding of specific attributes into the model architecture. A CNN architecture may include a sequence of layers, where each layer transforms one activation into another through a differentiable function.
[0108] In the example shown, the artificial neural network 190b can implement a convolutional neural network (CNN) module. The CNN module 190b can be configured to perform computer vision operations on video frames. The CNN module 190b can be configured to achieve object recognition through multi-layer feature detection. The CNN module 190b can be configured to compute descriptors based on the performed feature detections. The descriptors enable the processor 102 to determine the probability that a pixel in a video frame corresponds to a specific object (e.g., a specific manufacture / model / year of a vehicle, identifying a person as a specific individual, detecting the type of animal, detecting facial features, etc.).
[0109] CNN module 190b can be configured to implement convolutional neural network capabilities. CNN module 190b can be configured to use deep learning techniques to implement computer vision. CNN module 190b can be configured to use a training process through multi-layer feature detection to implement pattern and / or image recognition. CNN module 190b can be configured to perform inference against machine learning models.
[0110] The CNN module 190b can be configured to perform feature extraction and / or matching solely in hardware. Feature points typically represent regions of interest (e.g., corners, edges, etc.) in a video frame. By temporarily tracking feature points, estimates of the self-motion of the capture platform or motion models of objects observed in a scene can be generated. To track feature points, a matching operation is typically incorporated into the CNN module 190b by hardware to find the most probable correspondence between feature points in a reference video frame and a target video frame. During the matching of reference and target feature points, each feature point can be represented by a descriptor (e.g., image patch, SIFT, BRIEF, ORB, FREAK, etc.). Implementing the CNN module 190b using dedicated hardware circuitry allows for real-time computation of descriptor matching distances.
[0111] The CNN module 190b can be configured to perform face detection, face recognition, and / or presence determination. For example, face detection, face recognition, and / or presence determination can be performed based on a trained neural network implemented by the CNN module 190b. In some embodiments, the CNN module 190b can be configured to generate a depth image based on a structured light pattern. The CNN module 190b can be configured to perform various detection and / or recognition operations and / or perform 3D recognition operations.
[0112] CNN module 190b can be a dedicated hardware module configured to perform feature detection on video frames. Features detected by CNN module 190b can be used to compute descriptors. CNN module 190b can determine, in response to the descriptors, the probability that a pixel in a video frame belongs to a specific object and / or multiple objects. For example, using the descriptors, CNN module 190b can determine the probability that a pixel corresponds to a specific object (e.g., a person, furniture item, pet, vehicle, etc.) and / or characteristics of the object (e.g., the shape of the eyes, the distance between facial features, the hood of a vehicle, body parts, a vehicle's license plate, a person's face, clothing worn by a person, etc.). Implementing CNN module 190b as a dedicated hardware module of processor 102 allows device 100 to perform computer vision operations locally (e.g., on-chip) without relying on the processing power of remote devices (e.g., communicating data to a cloud computing service).
[0113] The computer vision operations performed by the CNN module 190b can be configured to perform feature detection on video frames to generate descriptors. The CNN module 190b can perform object detection to determine regions in the video frames with a high probability of matching a specific object. In one example, an open operand stack can be used to customize the type of the objects(s) to be matched (e.g., reference objects) (thus enabling the programmability of processor 102 to implement various artificial neural networks defined by directed acyclic graphs, each providing instructions for performing various types of object detection). The CNN module 190b can be configured to perform local masking on regions with a high probability of matching a specific object(s) to detect the object.
[0114] In some embodiments, the CNN module 190b can determine the location (e.g., 3D coordinates and / or positional coordinates) of various features (e.g., characteristics) of a detected object. In one example, 3D coordinates can be used to determine the position of a person's arms, legs, chest, and / or eyes. A positional coordinate for the vertical positioning of a body part in 3D space on a first axis and another coordinate for the horizontal positioning of a body part in 3D space on a second axis can be stored. In some embodiments, the distance from the lens 160 can represent a coordinate for the depth positioning of a body part in 3D space (e.g., a positional coordinate on a third axis). Using the positioning of the various body parts in 3D space, the processor 102 can determine the body position and / or the detected human body characteristics.
[0115] The CNN module 190b can be pre-trained (e.g., configured to perform computer vision to detect objects based on received training data). For example, the results of the training data (e.g., a machine learning model) can be pre-programmed and / or loaded into processor 102. The CNN module 190b can perform inference against the machine learning model (e.g., to perform object detection). Training may include determining weights for each layer of the neural network model. For example, weights can be determined for each layer used for feature extraction (e.g., convolutional layers) and / or for classification (e.g., fully connected layers). The weights learned by the CNN module 190b can vary depending on the design criteria of a particular implementation.
[0116] The CNN module 190b can perform feature extraction and / or object detection by performing convolution operations. These convolution operations can be hardware-accelerated for fast (e.g., real-time) computation that can be performed with low power consumption. In some embodiments, the convolution operations performed by the CNN module 190b can be used to perform computer vision operations. In some embodiments, the convolution operations performed by the CNN module 190b can be used for any function (e.g., 3D reconstruction) that may involve computing convolution operations and is performed by the processor 102.
[0117] Convolution operations can include sliding a feature detection window along a layer while performing computations (e.g., matrix operations). The feature detection window can apply filters to pixels and / or extract features associated with each layer. The feature detection window can be applied to a pixel and multiple surrounding pixels. In the example, a layer can be represented as a matrix representing the values of features from one of the pixels and / or layers, and the filters applied by the feature detection window can be represented as matrices. Convolution operations can apply matrix multiplication between regions of the current layer covered by the feature detection window. Convolution operations can cause the feature detection window to slide along a region of the layer to generate a result representing each region. The size of the regions, the type of operation applied by the filters, and / or the number of layers can vary depending on the design criteria of a particular implementation.
[0118] Using convolutional operations, the CNN module 190b can compute multiple features for pixels of the input image at each extraction step. For example, each layer in a layer can receive input from a set of features located in a small neighborhood (e.g., a region) of a previous layer (e.g., a local receptive field). Convolutional operations can extract basic visual features (e.g., oriented edges, endpoints, corners, etc.), which are then combined by higher layers. Since the feature extraction window operates on pixels and nearby pixels (or subpixels), the result of the operation can have location invariance. Layers can include convolutional layers, pooling layers, non-linear layers, and / or fully connected layers. In the example, convolutional operations can learn to detect edges from raw pixels (e.g., the first layer), then use features from previous layers (e.g., detected edges) to detect shapes in the next layer, and then use the shapes to detect higher-level features in higher layers (e.g., facial features, pets, vehicles, vehicle components, furniture, etc.), and the last layer can be a classifier using the higher-level features.
[0119] The CNN module 190b can perform data streams directed to feature extraction and matching, including two-stage detection, warping operators, component operators manipulating component lists (e.g., components can be regions of vectors sharing common attributes and can be combined with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, min-max reduction operators, pooling operators, non-minimum non-maximum suppression operators, scan-window-based non-maximum suppression operators, aggregation operators, scattering operators, statistical operators, classifier operators, integral image operators, comparison operators, indexing operators, pattern matching operators, feature extraction operators, feature detection operators, two-stage object detection operators, score generation operators, block reduction operators, and upsampling operators. The types of operations performed by the CNN module 190b to extract features from training data can vary depending on the design criteria of a particular implementation.
[0120] One or more of the hardware modules 190a-190n can be configured to implement other types of AI models. In one example, hardware modules 190a-190n can be configured to implement image-to-text AI models and / or video-to-text AI models. In another example, hardware modules 190a-190n can be configured to implement large language models (LLMs). Implementing AI models using hardware modules 190a-190n can provide AI acceleration, enabling the execution of complex AI tasks on edge devices such as edge devices 100a-100n.
[0121] One of the hardware modules 190a-190n can be configured to perform virtual aperture imaging. One of the hardware modules 190a-190n can be configured to perform transform operations (e.g., FFT, DCT, DFT, etc.). The number, type, and / or operations performed by the hardware modules 190a-190n can vary according to the design criteria of a particular implementation.
[0122] Each of the hardware modules 190a-190n can implement a processing resource (or hardware resource or hardware engine). Hardware engines 190a-190n can be operable to perform specific processing tasks. In some configurations, hardware engines 190a-190n can operate in parallel and independently of each other. In other configurations, hardware engines 190a-190n can operate collaboratively to perform assigned tasks. One or more of the hardware engines 190a-190n can be homogeneous processing resources (all circuits 190a-190n can have the same capabilities) or heterogeneous processing resources (two or more circuits 190a-190n can have different capabilities).
[0123] refer to Figure 5A block diagram illustrating operations for high-performance and low-complexity adaptive video image dehazing is shown. Block diagram 200 is shown. Block diagram 200 can implement a video dehazing module. The video dehazing module 200 can be implemented as one or more hardware modules 190a-190n of the processor 102. The video dehazing module 200 can be configured to implement high-performance and low-complexity adaptive video dehazing.
[0124] The video dehazing module 200 can be configured to receive a plurality of video frames 202a-202n. Video frames 202a-202n may be generated in response to a signal VIDEO. In an example, the video processing pipeline of processor 102 can be configured to process pixel data arranged as video frames 202a-202n. Pixel data may be generated by and / or received from image sensor 180. In some embodiments, the video processing pipeline can be configured to perform various preprocessing operations on the pixel data before (or after) a dehazing operation performed by the video dehazing module 200. In one example, the video dehazing module 200 may be a module implemented as part of a video processing pipeline. Video frames 202a-202n may be transmitted within processor 102 as signals (e.g., FRAMES). Signal FRAMES may include image input data. For example, video frames 202a-202n may include video frames that have not yet been dehazed (e.g., foggy input video frames).
[0125] The video dehazing module 200 may include blocks (or circuits) 204, 206, 208, 210, 212, 214, 216, and / or 218. Circuit 204 may implement a low-pass filter. Circuit 206 may implement a brightness interval control module. Block 208 may be a brightness distribution map. Circuit 210 may implement a smoothing control module. Circuit 212 may implement a multiplication module. Circuit 214 may implement a summation module. Block 216 may be a detail layer. Circuit 218 may implement a summation module. The video dehazing module 200 may include other components (not shown). The number, type, and / or arrangement of the components of the video dehazing module 200 may vary according to the design criteria of a specific implementation.
[0126] The video dehazing module 200 can be configured to divide input image data in the signal FRAMES into a high-frequency detail layer and a low-frequency brightness layer. The signal FRAMES can be presented to a low-pass filter 204, a summing module 214, and a summing module 218. The video dehazing module 200 can be configured to receive a signal (e.g., DFG-WGT). The signal DFG-WGT may include dehazing strength control points. The video dehazing module 200 can be configured to generate a signal (e.g., FRM-DFG). The signal FRM-DFG may be a dehazed video frame. The video dehazing module 200 can be configured to generate a dehazed video frame in the signal FRM-DFG in response to the signal FRAMES and the signal DFG-WGT.
[0127] Low-pass filter 204 can be configured to perform low-pass filtering. Low-pass filter 204 can be configured to generate a low-frequency layer in response to blocking high frequencies of the input and allowing low frequencies of the input to pass through. Low-pass filter 204 can be configured to filter signal FRAMES. Low-pass filter 204 can generate a signal (e.g., LFL). The signal LFL can transmit the low-frequency layer. For example, low-frequency data of input video frames 202a-202n may appear blurry and / or lack detail in the absence of high-frequency content. The low-frequency layer can provide a representation of various brightness positions of video frames 202a-202n. The signal LFL can be presented to the brightness interval control module 206, the summing point 214, and / or the multiplication module 212.
[0128] The low-pass filter 204 may have a cutoff frequency. For example, video data in video frames 202a-202n corresponding to frequencies above the cutoff frequency may be blocked, while frequencies below the cutoff frequency may pass through as a low-frequency layer. The value of the cutoff frequency may be an adjustable parameter. Typically, the value of the cutoff frequency may depend on the actual scene captured in video frames 202a-202n. For example, the cutoff frequency may be related to the image size and / or frame rate of video frames 202a-202n. The processor 102 may adjust the cutoff frequency in response to the size and / or frame rate of video frames 202a-202n generated by the image sensor 180. In one example, if video frames 202a-202n have an image size of 4k (e.g., 3840x2160) and a frame rate of 30fps, then the cutoff frequency setting for the low-pass filter 204 may be approximately 60MHz. In another example, if video frames 202a-202n have an image size of 1920x1080p and a frame rate of 30fps, then the cutoff frequency of the low-pass filter 204 can be approximately 15MHz. In some embodiments, the cutoff frequency can be a parameter that can be automatically adjusted by the processor 102 (e.g., in response to detecting image size and / or frame rate information and / or settings of the image sensor 180). In some embodiments, the cutoff frequency can be a parameter that is adjustable via user control (e.g., via input from the signal USR). The specific value of the cutoff frequency of the low-pass filter 204 and / or the method of adjusting the cutoff frequency parameter can vary depending on the design criteria of a particular implementation.
[0129] The luminance interval control module 206 can receive the signal LFL. The luminance interval control module 206 can be configured to generate a luminance distribution map 208 in response to a low-frequency layer. The luminance distribution map 208 can be generated in response to the luminance level in the low-frequency layer based on image location distribution (e.g., the location of luminance values in a specific video frame). Output statistics of luminance levels based on image location can include the luminance distribution map 208. The luminance interval control module 206 can set the luminance interval based on the actual application being implemented. The luminance interval control module 206 can provide the number of luminance intervals (e.g., N). The number of luminance intervals (e.g., N) can be an adjustable parameter. In one example, the processor 102 can automatically set the number of luminance intervals. In another example, the number of luminance intervals can be adjustable via user control (e.g., via input from the signal USR). The luminance distribution map 208 can be transmitted as a signal (e.g., LDM). The signal LDM can be generated in response to the signal LFL.
[0130] After the original (e.g., foggy) video frames 202a-202n are filtered by low-pass filter 204, the brightness distribution map 208 of the current image can be based on the low-frequency layer (LFL). The brightness values in the brightness distribution map 208 can be ordered from dark to bright, and each brightness value in the brightness distribution map 208 can correspond to one of the N brightness intervals.
[0131] The smoothing control module 210 can be configured to determine dehazing intensity weights for the brightness distribution map 208. The smoothing control module 210 can be configured to perform adaptive smoothing for each of the dehazing intensity weights. The smoothing control module 210 can be configured to receive the signal LDM and / or the signal DFG-WGT. The smoothing control module 210 can be configured to generate a signal (e.g., SMTH-DFG). The signal SMTH-DFG may include the dehazing intensity weights to which adaptive smoothing has been applied.
[0132] The smoothing control module 210 can be configured to sort the luminance distribution maps 208 of video frames 202a-202n from dark to bright. The smoothing control module 210 can be configured to divide the luminance distribution map 208 into multiple luminance intervals. For example, the luminance distribution map 208 can be divided into N intervals. The number of intervals used for the luminance distribution map 208 (e.g., N) can be an adjustable value based on available luminance values. For example, if the luminance values have a range of 0-255 (e.g., 0 for darkest and 255 for brightest), and the luminance interval is set to a value of 2, then the number N can be 128 (e.g., N = 256 / 2). For example, the luminance interval control module 206 can set the luminance intervals based on the implemented application. The luminance interval control module 206 can set the number N of luminance intervals. The specific luminance interval used can vary according to the design criteria of a particular implementation.
[0133] Control points for dehazing intensity can be set for each brightness interval. These control points for each interval can be adjustable parameters provided in the signal DFG-WGT. In some embodiments, the control points for dehazing intensity can be automatically set by the processor. In some embodiments, the control points for dehazing intensity can be user-adjustable values. The signal DFG-WGT can be input parameter values for the control points to provide weights for the dehazing intensity. For example, if a smooth dehazing control intensity is set for N brightness intervals, the video dehazing module 200 can adaptively determine the corresponding dehazing intensity based on the N brightness intervals for various different images (which may have different brightness distribution maps).
[0134] In some embodiments, the signal DFG-WGT can provide the number of intervals to be used and / or the defogging intensity for the intervals of the brightness distribution map 208. In one example, the signal DFG-WGT can be provided as user input via the signal USR. In another example, the signal DFG-WGT can be pre-programmed in memory 150 (e.g., based on engineering experience for providing accurate defogging). In yet another example, the signal DFG-WGT can be a learned value (e.g., based on an AI model and training data) for determining appropriate control points for specific environmental factors (e.g., the amount of light in the environment, the amount of fog in the environment, user preferences, etc.). The specific values selected for the control points in the signal DFG-WGT and / or the number of selected intervals can vary depending on the design criteria of the particular implementation.
[0135] The smoothing control module 210 can be configured to perform adaptive smoothing for control points based on curve fitting. Curve fitting can be applied to select the weights of the control points (e.g., N dehazing intensity control points). Curve fitting can be applied to the intensity of the dehazing control points. Smoothing of the dehazing control points can be achieved to ensure the smoothness of the image brightness distribution in the output dehazed video frames. For example, smoothing can prevent large changes in image brightness differences (e.g., avoiding high contrast differences in local areas). Any local area with high contrast differences may look unnatural and may distract the end user.
[0136] In one example, curve fitting can be performed based on a Bézier curve. Bézier curves typically provide curve smoothing in plotting software. Bézier curves can be reconfigured from plotting software to provide adaptive curve fitting for setting the intensity of control points. In some embodiments, the signal DFG-WGT can enable the selection of the order of the Bézier curve (or other curve fitting techniques that can be implemented, such as B-splines, Lanczos, Catmull-Rom splines, etc.). In some embodiments, memory 150 may include a lookup table that provides selection of the order of the Bézier curve based on the number of brightness intervals and / or a specific dehazing intensity control point. In some embodiments, processor 102 may implement an AI model trained to select the order of the Bézier curve for a specific application and / or based on a trade-off between the characteristics of video frames 202a-202n and / or output image quality and available computing resources. The method for selecting the order of the Bézier curve can vary depending on the design criteria of a particular implementation.
[0137] Typically, choosing a higher-order Bézier curve provides a smoother fit for the control point intensity. However, choosing a higher order can increase complexity, resulting in higher hardware resource consumption compared to lower-order Bézier curves. The smoothing control module 210 can be configured to balance the smoothness of the curve fit with the consumption of hardware resources to provide high-quality dehazing for the output video frame with limited complexity. Since the dehazing intensity value of the control point can be adjusted, the amount of residual fog in the output video frame can also be adjusted. Setting the dehazing intensity value to a high value may result in reduced brightness and / or detail in some areas of the dehazed video frame. For example, a certain amount of dehazing intensity can reduce brightness, which may result in reduced contrast of details in dark areas of the dehazed video frame.
[0138] The smoothing control module 210 can be configured to dehaze the entire brightness distribution based on N dehazing intensity control weights that have been smoothed by curve fitting. Within the corresponding image brightness interval, the control weights can be fitted based on the brightness distribution map 208. Fitting can be performed in real time to obtain adaptive video image dehazing control based on image position and brightness distribution. In response to the signal LDM, the signal DFG-WGT, and the applied curve fitting, the signal SMTH-DFG can be presented to the multiplication module 212.
[0139] Multiplication module 212 can be configured to perform multiplication operations. Multiplication module 212 can receive a signal SMTH-DFG and a signal LFL. For example, multiplication module 212 can be configured to perform a multiplication of adaptively smoothed dehazing intensity weights with a low-frequency layer. For example, the smoothed dehazing intensity weights (e.g., signal SMTH-DFG) can be based on a low-frequency layer (e.g., signal LFL). Each corresponding luminance interval in the low-frequency layer can be multiplied by the corresponding smoothed dehazing intensity weight in multiplication module 212. Multiplication module 212 can be configured to generate a signal (e.g., DEFOG). Signal DEFOG can include the result of applying the adaptively smoothed dehazing intensity weights. For example, signal DEFOG can provide a dehazing result. Signal DEFOG can include information corresponding to fog that can be removed in video frames 202a-202n. Signal DEFOG can be presented to summing module 218.
[0140] Summation module 214 can be configured to receive signals FRAMES and LFL. Summation module 214 can be configured to subtract the low-frequency layer from video frames 202a-202n. Summation module 214 can generate detail layer 216. Detail layer 216 can be a high-frequency detail layer. For example, detail layer 216 can be retained after subtracting the low-frequency layer from video frames 202a-202n. Detail layer 216 can be transmitted via a signal (e.g., DL). Detail layer 216 can include details of video frames 202a-202n corresponding to high-frequency information. For example, detail layer 216 can include fine details regarding texture, edges (e.g., abrupt changes between adjacent pixels), and / or other complex visual information. Detail layer 216 can define object boundaries and / or provide overall sharpness and definition of visual elements. Signal DL can be generated in response to signals FRAMES and LFL. Signal DL can be presented to summation module 218.
[0141] Summation point 218 can be configured to receive signals FRAMES, DL, and / or DEFOG. Summation module 218 can be configured to subtract the dehazing result generated in response to adaptively smoothed dehazing intensity weights from video frames 202a-202n and detail layer 216. Summation module 218 can generate a dehazed output video frame. The dehazed output video frame can be transmitted via signal FRM-DFG. The dehazed output video frame may include video data from video frames 202a-202n, where blur caused by fog in the environment is removed (or partially removed).
[0142] A high-frequency layer can be added to the final result to preserve the original image details after dehazing. The high-frequency layer can be added to the image processing result after applying adaptively smoothed dehazing control weights. For example, the high-frequency layer can be added to the dehazing result instead of directly to the dehazing weights (e.g., the smoothed dehazing weights are multiplied by the low-frequency layer). Video frames 202a-202n in the signal FRAMES can include both high-frequency and low-frequency data. The dehazing result can first be subtracted from video frames 202a-202n via summation point 218. After subtracting the dehazing result, the original high-frequency details may be lost (or partially lost, depending on the dehazing intensity). To recover the loss of high-frequency details based on the original high-frequency layer, a detail layer 216 can be added via summation point 218. The dehazed output video frames in the signal FRM-DFG can include video data with the dehazing result removed and the original high-frequency details recovered.
[0143] refer to Figure 6The diagram illustrates an example input video frame describing a foggy environment. Example video frame 250 is shown. Example video frame 250 can be one of video frames 202a-202n. For example, example video frame 250 can be an input video frame before fog removal (e.g., a foggy input video frame). The foggy input video frame 250 can be composed of... Figure 5 The video frame shown in association is one of the video frames processed by the video dehazing module 200.
[0144] The foggy input video frame 250 may include pixel data captured by the capture device 104. In one example, the foggy input video frame 250 may be provided to the processor 102 as a signal VIDEO. In another example, the foggy input video frame 250 may be generated by the processor 102 in response to the pixel data provided in the signal VIDEO. The pixel data may be received by the processor 102, and video processing operations may be performed by the video processing pipeline of the processor 102 to generate the foggy input video frame 250. In some embodiments, the foggy input video frame 250 may not be presented as human-visual video output to one or more video displays until a defogging operation has been performed by the video defogging module 200. In some embodiments, after the defogging operation has removed the fog, computer vision operations and / or video analysis operations may be performed within the processor 102 using the foggy input video frame 250. The foggy input video frame 250 may include pixel data arranged as video frames. The foggy input video frame 250 is shown as a visual representation (e.g., viewed by a person on a video output device (e.g., a monitor, touchscreen display, etc.)). Typically, processor 102 and / or video dehazing module 200 can perform operations on pixel data and / or pixel blocks.
[0145] Typically, a foggy input video frame 250 may include video images of vehicles traveling on a road, as well as trees and bushes on the side of the road. The environment in the foggy input video frame 250 may include foggy conditions. The foggy input video frame 250 may represent how the video output would look without applying defogging operations. For example, the view and / or details of the environment captured in the foggy input video frame 250 may be partially obscured by the foggy conditions.
[0146] The foggy input video frame 250 may include dashed lines 252 forming an irregular shape. Vertical dashed lines 254a-254n are shown within the irregular shape 252. The irregular shape 252 and the vertical dashed lines 254a-254n can represent a fog effect. For example, the irregular shape 252 can represent a fog boundary, and the vertical dashed lines 254a-254n can represent partial visual obstruction caused by fog. Partial visual obstruction caused by fog may manifest as a blurring effect.
[0147] A video frame portion 256 is shown on one side of the fog boundary 252. The video frame portion 256 may include a clear condition area. The clear condition area 256 may not be obscured by fog. For example, the clear condition area 256 may be outside the fog effect. Various objects, viewing distances, and / or visual details in the clear condition area 256 may appear clear.
[0148] A video frame portion 258 (e.g., opposite to the clear condition area 256) is shown on one side of the fog boundary 252. The video frame portion 258 may include a foggy condition area. The foggy condition area 258 may have some degree of visual obstruction caused by fog (or other types of humidity). For example, the foggy condition area 258 may be within a fog effect. Various objects, viewing distances, and / or visual details in the foggy condition area 258 may appear blurry, more difficult to see, and / or more difficult to distinguish compared to similar objects that may be in the clear condition area 256.
[0149] As an illustrative example, portions of a foggy input video frame 250 and / or objects located in the clear condition area 256 can be drawn with thicker lines than those used to draw objects located in the foggy condition area 258. The difference in line thickness between the clear condition area 256 and the foggy condition area 258 can provide a visual indication that objects, features, and / or characteristics may be blurred and / or difficult to interpret and / or viewed from a shorter distance in the foggy condition area 258. The amount and / or type of visual difference caused by fog can vary depending on the environmental conditions in the captured environment.
[0150] The foggy input video frame 250 may include a combination of low-frequency and high-frequency image content. This combination can cause the foggy input video frame 250 to appear natural (e.g., similar to what a person would see when viewing an environment captured in the foggy input video frame 250). The foggy input video frame 250 may include multiple visual details 260a-260j and multiple visual details 262a-262o. Visual details 260a-260j may represent low-frequency image content. Visual details 262a-262o may represent high-frequency image content. The low-frequency image content 260a-260j and the high-frequency image content 262a-262o can be shown as illustrative examples of different types of visual content in the foggy input video frame 250. For example, low-frequency image content 260a-260j may not represent all low-frequency image content in the foggy input video frame 250, and high-frequency image content 262a-262o may not represent all high-frequency image content. Typically, in video frames captured by device 100, low-frequency image content and high-frequency image content may have different visual characteristics compared to the representative examples shown in low-frequency image content 260a-260j and high-frequency image content 262a-262o.
[0151] Low-frequency image content 260a-260j can be in both the clear condition area 256 and the foggy condition area 258. In the example shown, low-frequency image content 260a can be a tree, low-frequency image content 260b can be a shrub, low-frequency image content 260c can be a tree, low-frequency image content 260d can be a vehicle (e.g., a sedan-style car), low-frequency image content 260e can be a wire fence with wooden poles, low-frequency image content 260f can be a roadside, low-frequency image content 260g can be a roadside, and low-frequency image content 260h-260j can be nearby vegetation. For example, nearby vegetation 260h can be shown in a clearer manner (e.g., thicker lines) in the clear condition area 256, and nearby vegetation 260i can be shown in a less clear manner (e.g., thinner lines) in the foggy condition area 258. In another example, tree 260a is shown partially in clear condition area 256 and partially in foggy condition area 258, and the portion of tree 260a in foggy condition area 258 may appear less clear (e.g., a thinner line) compared to the portion in clear condition area 256.
[0152] High-frequency image content 262a-262o can be in both the clear condition area 256 and the foggy condition area 258. In the example shown, high-frequency image content 262a can be the wood grain pattern of trees, high-frequency image content 262b can be the leaf detail of shrubs, high-frequency image content 262c can be the wood grain pattern of trees, high-frequency image content 262d can be the driver of a vehicle, high-frequency image content 262e can be a small visual feature of the vehicle (e.g., a side mirror), high-frequency image content 262f can be a design feature of the vehicle, high-frequency image content 262g-262h can be leaves on trees, high-frequency image content 262i can be the wood grain on a fence, high-frequency image content 262j can be a puddle, high-frequency image content 262k-262l can be a crack in the road, high-frequency image content 262m can be a road line, and high-frequency image content 262n-262o can be distant vegetation. For example, distant vegetation 262n may be shown in a clearer manner (e.g., thicker lines) in the clear condition area 256, and distant vegetation 262o may be shown in a less clear manner (e.g., thinner lines) in the foggy condition area 258. In another example, leaves 262g on a tree are shown partially in the clear condition area 256 and partially in the foggy condition area 258, and the portion of the leaves 262g on the tree in the foggy condition area 258 may appear less clear (e.g., thinner lines) compared to the portion in the clear condition area 256.
[0153] A foggy input video frame 250 can be presented to a low-pass filter 204, a summing module 214, and / or a summing module 218. The low-pass filter 204 can be configured to separate low-frequency image content 260a-260j from high-frequency image content 262a-262o. The video defogging module 200 can be configured to perform operations based on the low-frequency image content 260a-260j and the high-frequency image content 262a-262o to generate a defogging output video frame.
[0154] refer to Figure 7 This diagram illustrates the region of the low-frequency layer used for a luminance distribution map in an input video frame. An example low-frequency layer 300 is shown. The example low-frequency layer 300 can be transmitted in the signal LFL. The example low-frequency layer 300 can be generated based on one of video frames 202a-202n. In the example shown, it can be generated based on... Figure 5 The foggy input video frame 250 shown in association generates a low-frequency layer 300.
[0155] A foggy input video frame 250 can be presented to a low-pass filter 204. The low-pass filter 204 can block high-frequency image content and transmit a low-frequency layer in the signal LFL. For example, the foggy input video frame 250 can be filtered by the low-pass filter 204 at a cutoff frequency to generate a low-frequency layer 300. The low-frequency layer 300 can have similar visual content to the foggy input video frame 250, but some visual content is removed due to filtering. The low-frequency layer 300 can include the low-frequency image content 260a-260j of the foggy input video frame 250, but exclude the high-frequency image content 262a-262o. For example, the high-frequency image content 262a-262o can be filtered out from the foggy input video frame 250 to generate the low-frequency layer 300.
[0156] Low-frequency layer 300 may include data regarding the overall structure and / or extensive features of the input video frames 202a-202n. Typically, low-frequency layer 300 may include the overall shape and / or structure of various objects and / or visual features. For example, low-frequency layer 300 may provide the general outline and / or large-scale form of video frames 202a-202n. Low-frequency layer 300 may provide extensive color areas (e.g., large areas with similar colors and / or intensities). Low-frequency layer 300 may provide gradual transitions (e.g., slow changes in brightness and / or color relative to a position within video frames 202a-202n). The characteristics of low-frequency layer 300 may include coarse details and / or a blurred appearance (e.g., compared to the original video content in video frames 202a-202n). Low-frequency layer 300 may include the overall layout and composition of video frames 202a-202n, rather than fine details. Typically, the overall shape and / or structure provided by the low-frequency layer 300 is sufficient to identify large objects in video frames 202a-202n (e.g., using computer vision manipulation).
[0157] In the example shown, low-frequency layer 300 may provide low-frequency image content 260d of the vehicle, which may provide the overall shape of the vehicle (e.g., the shape of a sedan), but may not provide the fine details of the vehicle in the high-frequency image content 262e-262f (e.g., side mirrors and / or vehicle design details may be missing). Similarly, low-frequency layer 300 may provide low-frequency image content 260f-260g of the road shape, but may not provide the fine details of the high-frequency image content 262j-262m (e.g., cracks, lines, puddles, etc.). The amount of detail shown in low-frequency layer 300 may vary depending on the captured environment and / or the cutoff frequency of low-pass filter 204.
[0158] Fog effects 254a-254n can be found in low-frequency layer 300. For example, details of fog effects 254a-254n can be extracted from low-frequency layer 300 to determine control point strength in order to remove fog effects 254a-254n.
[0159] The low-frequency layer 300 can be presented to the brightness interval control module 206. The brightness interval control module 206 can be configured to generate a brightness distribution map 208 in response to the low-frequency layer 300. The brightness interval control module 206 can be configured to divide the low-frequency layer 300 into multiple rectangular regions to obtain brightness values.
[0160] The low-frequency layer 300 may include vertical lines 302a-302n and horizontal lines 304a-304m. The vertical lines 302a-302n and the horizontal lines 304a-304m may divide the low-frequency layer into multiple regions 306aa-306mn. Regions 306aa-306mn may be rectangular regions corresponding to specific image locations within the low-frequency layer 300. In the example shown, there may be more vertical lines 302a-302n than horizontal lines 304a-304m (e.g., the image size has more horizontal pixels than vertical pixels). In some embodiments, the low-frequency layer 300 may be divided into rectangular regions 306aa-306mn based on having more horizontal lines 304a-304m than vertical lines 302a-302n. The number of rectangular regions 306aa-306mn may be related to the image size. For example, the size of the rectangular region 306aa-306mn can be 16 pixels × 16 pixels. The block size can be a fixed value. In one example, if video frames 202a-202n are 4K images (e.g., 3840x2160p), the number of horizontal rectangular regions 306aa-306mn can be 240 (e.g., 3840 / 16), and the number of vertical rectangular regions 306aa-306mn can be 135 (e.g., 2160 / 16). The number of regions 306aa-306mn, the size of regions 306aa-306mn, and / or the aspect ratio of each of regions 306aa-306mn can vary according to the design criteria of a particular implementation.
[0161] The luminance interval control module 206 can be configured to extract information about the low-frequency layer 300 from rectangular regions 306aa-306mn. Each of regions 306aa-306mn can include position and / or luminance information about the foggy input video frame 250. For example, region 306aa can provide luminance information for the upper left position of the foggy input video frame 250. In another example, region 306mn can provide luminance information for the lower right position of the foggy input video frame 250. In yet another example, region 306ii (not specifically labeled) can provide luminance information about the approximate center position of the foggy input video frame 250. Output statistics from regions 306aa-306mn can be used to generate a luminance distribution map 208.
[0162] refer to Figure 8 A graph illustrating luminance values used for a luminance distribution map is shown. An example distribution map 350 is shown. Distribution map 350 can be an illustrative example of luminance distribution map 208. For example, distribution map 350 can be generated by luminance interval control module 206 in response to low-frequency layer 300.
[0163] Distribution map 350 may include multiple vertical lines 352a-352n and / or multiple horizontal lines 354a-354m. The vertical lines 352a-352n and the horizontal lines 354a-354m may divide distribution map 350 into multiple regions 356aa-356mn. Regions 356aa-356mn may each include a luminance value L. In the example shown, region 356aa may be a luminance value L00, region 356ab may be a luminance value L01, region 356an may be a luminance value L0n, region 356ba may be a luminance value L10, region 356ma may be a luminance value L0, region 356mn may be a luminance value Lmn, etc. In one example, each of the luminance values L00-Lmn may include values in cd / m². 2 A luminance value measured in units. For example, luminance values can range from 0.1 cd / m². 2 Up to 500 cd / m 2 The range and / or from 10 -5 cd / m 2 Up to 10 8 cd / m 2 The range. In another example, each of the luminance values L00-Lmn can be an encoded value. For example, luminance values can be encoded as values between 0 and 255. Specific luminance values can vary depending on the design criteria of a particular implementation.
[0164] The vertical lines 352a-352n of the distribution map 350 can correspond to the vertical lines 302a-302n of the low-frequency layer 300. The horizontal lines 354a-354m of the distribution map 350 can correspond to the horizontal lines 304a-304m of the low-frequency layer 300. The brightness value regions 356aa-356mn can correspond to the image location regions 306aa-306mn of the low-frequency layer 300. Each of the brightness values L00-Lmn in the distribution map 350 can represent the brightness value at the corresponding image location region 306aa-306mn of the low-frequency layer 300. For example, the brightness distribution map 208 can be determined based on multiple rectangular regions 306aa-306mn to obtain the corresponding brightness values L00-Lmn. The output statistics of the brightness level can be based on the brightness values L00-Lmn of the image location region 306aa-306mn of the low-frequency layer 300.
[0165] The smoothing control module 210 can receive the brightness values L00-Lmn from the brightness distribution map 208. The smoothing control module 210 can sort the brightness values L00-Lmn from dark to bright and divide them into multiple brightness intervals (e.g., N brightness intervals). For each of the N brightness intervals, the smoothing control module 210 can set a control point for the amount of defogging intensity. For example, the signal DFG-WGT can be an input parameter providing the control point weights for the defogging intensity. The control point can be a defogging intensity weight.
[0166] The smoothing control module 210 can provide adaptive smoothing for the intensity of N dehazing strength weights. The smoothing control module 210 can perform fitting control on the dehazing strength weights. Fitting control can be based on Bézier curve smoothing. The smoothing control module 210 can apply Bézier curve smoothing to the settings for the dehazing strength weights. Bézier curve smoothing can ensure the smoothness of the image brightness distribution after dehazing. For example, adaptive smoothing performed using Bézier curves can prevent large jumps (e.g., differences) in image brightness. Dehazing strength weights with adaptive smoothing can be generated for the SMTH-DFG signal.
[0167] Video frames 202a-202n can be dehazed by applying adaptively smoothed dehazing intensity weights based on N brightness distribution intervals. For the corresponding image brightness interval, the dehazing intensity weights can be fitted based on the brightness distribution map 208 to provide adaptive smoothing. The generation of adaptively smoothed dehazing intensity weights can be performed in real time to provide adaptive video image dehazing control based on image position and brightness distribution. Multiplication operations performed by multiplication module 212 can be performed between the adaptively smoothed dehazing intensity weights and the low-frequency layer 300. For example, the adaptively smoothed dehazing intensity weights in signal SMTH-DFG can be multiplied by the low-frequency layer 300 in signal LFL to generate signal DEFOG. Signal DEFOG can provide the amount of dehazing for each position of each video frame in the input video frames 202a-202n in real time.
[0168] refer to Figure 9 A diagram illustrating an example high-frequency layer of an input video frame is shown. A high-frequency detail layer 380 is shown. The example high-frequency detail layer 380 can be transmitted in the signal DL. The example high-frequency detail layer 380 can be generated based on one of the video frames 202a-202n. In the example shown, it can be based on... Figure 5 The foggy input video frame 250 shown in association generates a high-frequency detail layer 380.
[0169] A foggy input video frame 250 can be presented to a low-pass filter 204, a summing module 214, and a summing module 218. The low-pass filter 204 can block high-frequency image content and transmit a low-frequency layer in the signal LFL. For example, a low-frequency layer 300 can be generated by the low-pass filter 204. The low-frequency layer 300 can be subtracted from the foggy input video frame 250 (e.g., a corresponding one of input video frames 202a-202n) by the summing module 214 to generate a high-frequency detail layer 380. The high-frequency detail layer 380 can be... Figure 5 A representative example of detail layer 216 is shown in association.
[0170] The high-frequency detail layer 380 can have visual content similar to the foggy input video frame 250, but some visual content is removed due to the removal of low-frequency content from the low-frequency layer 300. The high-frequency detail layer 380 can include the high-frequency image content 262a-262o of the foggy input video frame 250, but exclude the low-frequency image content 260a-260j. For example, the high-frequency detail layer 380 can be generated by subtracting the low-frequency image content 260a-260j from the foggy input video frame 250.
[0171] The high-frequency detail layer 380 may include data on fine details and / or sharp transitions of the input video frames 202a-202n. Typically, the high-frequency detail layer 380 may include detail and edge data. The high-frequency detail layer 380 may be visually represented as a grayscale image (e.g., primarily without color data). The high-frequency detail layer 380 may correspond to areas of the image where pixel values change rapidly over short distances. For example, pixel values may change rapidly in portions of the input image such as: fine textures and complex patterns, sharp edges and boundaries between objects, small features and minute details, etc. The high-frequency detail layer 380 may provide visual characteristics. The high-frequency detail layer 380 may provide details corresponding to image sharpness (e.g., the high frequencies may include data for image crispness and clarity). The high-frequency detail layer 380 may provide contrast (e.g., sudden changes in brightness and / or color may be captured in the high-frequency data). The high-frequency detail layer 380 may include noise (e.g., random variations and / or graininess in the image may be present in the high-frequency components). The high-frequency detail layer 380 can represent rapid transitions between pixels and / or regions, where intensity and / or color values fluctuate rapidly across small areas in the spatial representation. Typically, the high-frequency detail layer 380 may include data with lower amplitudes compared to the data in the low-frequency layer 300 (e.g., high-frequency data may contribute less to the overall image).
[0172] In the example shown, high-frequency detail layer 380 may provide high-frequency image content 262a and 262g, which may correspond to fine details of trees (e.g., wood grain patterns), but not the overall shape and structure of the trees (e.g., provided in low-frequency image content 260a). Similarly, fine details in high-frequency image content 262d-262f (e.g., the driver in a vehicle, the vehicle's side mirrors, and design features of the vehicle) may be visible in high-frequency detail layer 380, but not the overall shape and / or structure of the vehicle (e.g., provided in low-frequency image content 260d). Similarly, fine details and / or sharpness of high-frequency image content 262j-262m (e.g., puddles, cracks, and lines on the road) may be visible in high-frequency detail layer 380, but not the overall structure of the road (e.g., provided in low-frequency image content 260f-260g). The amount of detail shown in high-frequency detail layer 380 may vary depending on the captured environment and / or the cutoff frequency of low-pass filter 204.
[0173] The fog effects 254a-254n may not be visible in the high-frequency detail layer 380. For example, the details of the fog effects 254a-254n may be in the low-frequency layer 300, which can be subtracted from the foggy input image 250 to generate the high-frequency detail layer 380.
[0174] The high-frequency detail layer 380 can be presented to the summing module 218. The high-frequency detail layer 380 can be used to generate dehazed output video frames. The high-frequency detail layer 380 can be used to recover the high-frequency details of the dehazed output video frames after subtracting the dehazed results from video frames 202a-202n.
[0175] refer to Figure 10 A diagram illustrating an example dehazed video frame output is shown. Example video frame 400 is shown. Example video frame 400 can be one of the dehazed output video frames in the signal FRM-DFG. For example, example dehazed output video frame 400 can be an output video frame after dehazing using dehazing intensity weights with adaptive smoothing. Example dehazed output video frame 400 can be generated in response to a dehazing operation performed by video dehazing module 200 in response to one of the input video frames 202a-202n. In the example shown, it can be based on... Figure 5 The foggy input video frame 250 shown in association generates a dehazed output video frame 400. For example, the dehazed output video frame 400 may include a combination of low-frequency image content 260a-260j that has been dehazed and the recovered high-frequency image content 262a-262o.
[0176] The dehazed output video frame 400 may include pixel data captured by the capture device 104 after a dehazing operation has been performed by the video dehazing module 200. In one example, the dehazed output video frame 400 may be provided as the output of the processor 102 as the signal VIDOUT. In another example, the processor 102 may internally use the dehazed output video frame 400 for various other operations (e.g., computer vision operations, video-to-text AI operations, sensor fusion operations utilizing radar data, etc.).
[0177] Typically, the dehazed output video frame 400 may include video content similar to the foggy input video frame 250 (e.g., video images of vehicles traveling on a road and trees and bushes on the side of the road). The dehazed output video frame 400 can provide similar content to the foggy input video frame 250, but with greater visual clarity due to the reduction in fog. For example, a view and / or detail of the environment partially obscured by foggy conditions, captured in the foggy input video frame 250, can be shown with greater clarity in the dehazed output video frame 400.
[0178] The dehazed output video frame 400 may include dashed lines 402 forming an irregular shape. Vertical dashed lines 404a-404m are shown within the irregular shape 402. The irregular shape 402 and the vertical dashed lines 404a-404m may represent a reduced fog effect. For example, the irregular shape 402 may represent a reduced fog boundary, and the vertical dashed lines 404a-404m may represent a reduced visual obstruction caused by fog. As a result of the dehazing operation performed by the video dehazing module 200, the reduced fog obstruction 404a-404m shown in the dehazed output video frame 400 may be less than the fog obstruction 254a-254n shown in the foggy input video frame 250.
[0179] A video frame portion 406 is shown on one side of the boundary 402 of the reduced fog. The video frame portion 406 may include an increased clarity area. The increased clarity area 406 may not be obscured by fog. For example, the increased clarity area 406 may be outside the reduced fog effect. Various objects, viewing distances, and / or visual details in the increased clarity area 406 may appear sharp. Due to the reduction of fog caused by the defogging operation, the portion of the defogging output video frame 400 including the increased clarity area 406 may be larger than the portion including the clarity area 256 in the foggy input video frame 250.
[0180] A video frame portion 408 (e.g., opposite to the increased clarity area 406) is shown on one side of the reduced fog boundary 402. The video frame portion 408 may include the reduced foggy area. The reduced foggy area 408 may have some degree of visual obstruction caused by fog (or other types of humidity). For example, the reduced foggy area 408 may be within a fog effect. Various objects, viewing distances, and / or visual details within the reduced foggy area 408 may appear blurry, may be more difficult to see, and / or may be more difficult to distinguish. Due to the reduction of fog caused by the defogging operation, the portion of the defogging output video frame 400 including the reduced foggy area 408 may be smaller than the portion including the foggy area 258 in the foggy input video frame 250.
[0181] In the example shown, because the reduced foggy area 408 in the dehazed output video frame 400 is smaller than the foggy area 258 in the foggy input video frame 250, more of the various objects and / or details can be visible without being visually obscured by the fog. For example, in the dehazed output video frame 400, puddles (e.g., high-frequency image content 262j) and nearby vegetation (e.g., low-frequency image content 260i) can be in the increased clarity area 406 after the fog is reduced, instead of being in the foggy area 258 as shown in the foggy input video frame 250.
[0182] Because of the adjustable dehazing strength value, the intensity of fog reduction can be adjusted. Typically, if the dehazing strength is strong, some bright areas and / or small details in the image may be reduced. To balance the potential loss of brightness and / or small details due to dehazing with the intensity of fog removal, the dehazed output video frame 400 may include some residual fog. Reduced fog barriers 404a-404m and reduced foggy areas 408 can represent the residual fog in the dehazed output video frame 400. Although residual fog may be present in the dehazed output video frame 400, its impact on the visual quality and / or detail of objects in the dehazed output video frame 400 may be less than the impact of fog in the foggy input video frame 250. For example, even if some objects / details may still be partially obscured by residual fog, these objects / details may be more visible after fog reduction (e.g., even in the reduced foggy area 408, the amount of visual occlusion and / or blurring effect due to fog may be less than in the foggy area 258).
[0183] As an illustrative example, portions of the dehazed output video frame 400 and / or objects located in the increased clarity area 406 can be drawn using lines thicker than those used to draw objects located in the reduced foggy area 408. However, the thickness of the lines in the reduced foggy area 408 can be shown as thicker than the thinnest lines used for the foggy area 258 in the foggy input video frame 250. The difference in line thickness between the reduced foggy area 408 and the foggy area 258 can provide a visual indication that, due to the fog removal operation, objects, features, and / or characteristics that may be blurry and / or difficult to interpret in the reduced foggy area 408 may be less blurry and / or less difficult to interpret and / or have a less short viewing distance compared to the foggy area 258. The amount and / or type of visual difference caused by fog reduction can vary depending on the environmental conditions in the capture environment.
[0184] The dehazed output video frame 400 may include a combination of low-frequency and high-frequency image content. The summing module 218 may receive a signal FRAME including the foggy input video frame 250, a signal DL including the high-frequency detail layer 380, and a signal DEFOG including the amount of dehazing for each location. For example, the high-frequency detail layer 280 may be added to the foggy input video frame 250, and the amount of dehazing for each location may be subtracted to generate the dehazed output video frame 400. The amount of dehazing for each location may be determined based on a dehazing intensity weight with adaptive smoothing. Adaptive smoothing ensures that the dehazed output video frame 400 provides a reduced foggy condition region 408 with reduced fog obstacles 404a-404m, while maintaining a gradual transition in brightness between adjacent regions in the dehazed output video frame 400. For example, fog reduction can be achieved without adding visually distracting artifacts. The dehazed output video frame 400 may be output as a signal FRM-DFG.
[0185] Dehazed output video frames 400 can be generated to provide clarity for driver assistance features (e.g., dehazing can provide a better view for backup cameras, rearview mirror cameras, in-vehicle cameras, vehicle surround views, etc.). For example, dehazed output video frames 400 can provide visual benefits when a person may be viewing the video output on a monitor. Dehazed output video frames 400 can be further generated to provide more detail for additional video processing operations such as computer vision operations. For example, when performing computer vision operations and / or video-to-text AI operations, the reduction of fog in the dehazed output video frames 400 can achieve accurate results and / or prevent uncertain (e.g., low confidence) results. Using dehazed output video frames 400, various objects can be detected in response to animal detection, household object detection, interior object detection, human detection, vehicle detection, road detection, sky area detection, obstacle detection, and / or external object detection (e.g., one or more of the neural network 190b and / or video-to-text AI models may include libraries configured to detect people, vehicles, objects, animals, etc.). In some embodiments, the reduction in blur caused by fog can help detect debris that may accumulate on lens 160.
[0186] Computer vision manipulation, debris analysis, and / or sensor fusion into text manipulation can be configured to detect characteristics of detected objects, behavior of detected objects, direction of movement of detected objects, context of detected objects, and / or liveness of detected objects. Object characteristics may include height, length, width, slope, arc length, color, color temperature, emitted light intensity, detected text on the object, movement path, movement speed, direction of movement, proximity to other objects, etc. Detected object characteristics may include the object's state (e.g., open, closed, on, off, etc.). Detected object characteristics may include distance measurements from lens 160 to the detected object. Behavior and / or liveness can be determined in response to the object's type and / or the characteristics of the detected object. In some embodiments, object behavior, direction of movement, and / or liveness can be determined by analyzing a sequence of dehazed output video frames captured over time in the signal FRM-DFG. For example, movement path and / or speed characteristics can be used to determine whether an object classified as a person is walking or running. The type of detected characteristics and / or behaviors can vary depending on the design criteria of a particular implementation.
[0187] Processor 102, CNN module 190b, and / or video-to-text AI model can be configured to implement region, animal, lens occlusion, object, and / or face detection techniques. In some embodiments, other types of subjects (e.g., vehicles, passengers, pedestrians, street signs, etc.) can be detected as objects of interest. Computer vision techniques and / or video-to-text techniques can generally be configured to detect regions of interest (ROIs) of detected objects and / or generate contextual information about detected objects and / or scenes. Computer vision techniques can be looped (e.g., to perform object / subject detection iteratively across the entire dehazed video frame) to determine whether any object of interest (e.g., as defined by a feature set) is within the field of view of lens 160 and / or image sensor 180.
[0188] Computer vision operations and / or video-to-text operations performed by processor 102, CNN module 190b, and / or video-to-text AI models can be configured to detect background objects and / or other types of objects. Background objects can be detected for other computer vision purposes (e.g., training data, labeling, depth detection, etc.). The type of subject identified as an object of interest can vary depending on the design criteria of a particular implementation. Details of computer vision, video-to-text operations, and / or sensor-fusion-to-text operations are described in conjunction with the following patent applications: U.S. Patent Application No. 18 / 583,298, filed February 11, 2024; U.S. Patent Application No. 18 / 621,504, filed March 29, 2024; U.S. Patent Application No. 18 / 657,588, filed May 7, 2024; and / or U.S. Patent Application No. 18 / 657,492, filed May 7, 2024, the appropriate portions of which are incorporated herein by reference.
[0189] refer to Figure 11 The diagram illustrates method (or process) 500. Method 500 can provide high-performance and low-complexity adaptive video image dehazing. Method 500 typically includes steps (or states) 502, 504, 506, decision step (or state) 508, 510, 512, 514, 516, 518, 520, and 522.
[0190] Step 502 can initiate method 500. In step 504, processor 102 can receive pixel data. For example, image sensor 180 can generate a signal VIDEO including pixel data in response to light input LIN captured by capture device 104. Next, in step 506, processor 102 can process the pixel data arranged as video frames. For example, processor 102 can perform various operations on the pixel data arranged as video frames (e.g., perform computer vision operations, calculate depth data, determine white balance, etc.). Video dehazing module 200 can receive video frames 202a-202n (e.g., as combined with...) Figure 5 (As shown). Next, method 500 can move to decision step 508.
[0191] In decision step 508, processor 102 may determine whether to adjust the cutoff frequency of low-pass filter 204. For example, video dehazing module 200 may adjust the cutoff frequency based on the resolution and / or frame rate of video frames 202a-202n. If it is determined that the cutoff frequency needs to be adjusted, method 500 may move to step 510. In step 510, the cutoff frequency of low-pass filter 204 may be set based on the application scenario. Next, method 500 may move to step 512. In decision step 508, if the cutoff frequency does not need to be adjusted, then method 500 may move to step 512. In step 512, low-pass filter 204 may perform a low-pass filtering operation on the current video frames in video frames 202a-202n to generate a low-frequency layer. Next, method 500 may move to step 514.
[0192] In step 514, the luminance interval control module 206 can generate a luminance distribution map 208. The luminance distribution map 208 can be generated in response to the low-frequency layer 300. Next, in step 516, the smoothing control module 210 can determine dehazing intensity weights for the luminance distribution map 208. The dehazing intensity weights can correspond to the luminance intervals of the luminance distribution map 208. For example, the luminance interval control module 206 can set the number of intervals for the luminance distribution map 208 to N. In step 518, the smoothing control module 210 can perform adaptive smoothing on each of the dehazing intensity weights to prevent luminance differences between regions of the current video frames in video frames 202a-202n. Next, in step 520, the video dehazing module 200 can generate dehazed video frames in response to the input video frames 202a-202n and the smoothed dehazing intensity weights (e.g., signal SMTH-DFG). For example, the dehazed video frames can be presented in the signal FRM-DFG. Next, method 500 can move to step 522. Step 522 can end method 500.
[0193] refer to Figure 12The diagram illustrates method (or process) 550. Method 550 can determine a smoothing control strength value for a brightness interval. Method 550 typically includes steps (or states) 552, 554, decision steps (or states) 556, 558, 560, 562, 564, 566, 568, and 570.
[0194] Step 522 can initiate method 550. In step 524, low-pass filter 204 can generate a low-frequency layer 300 based on the current video frame (e.g., foggy input video frame 250) in video frames 202a-202n. Next, method 550 can move to decision step 556. In decision step 556, luminance interval control module 206 can determine whether to adjust the number of luminance intervals. For example, the number of luminance intervals can be determined based on luminance interval values and / or a range of luminance values. If it is determined that the luminance intervals should be adjusted, method 550 can move to step 558. In step 558, luminance interval control module 206 can set the luminance intervals based on the application and / or the scene in video frames 202a-202n. Next, method 550 can move to step 560. In decision step 556, if the number of luminance intervals is not adjusted, then method 550 can move to step 560. In step 560, luminance interval control module 206 can generate a luminance distribution map 208 based on the low-frequency layer 300. Next, method 550 can move to decision step 562.
[0195] In decision step 562, the smoothing control module 210 can determine whether there are more brightness values L00-Lmn in the brightness distribution map 208. If there are more brightness values L00-Lmn, then method 550 can move to step 564. In step 564, the smoothing control module 210 can sort the next brightness value L00-Lmn from darkest to brightest. Next, method 550 can return to decision step 562. In decision step 562, if there are no more brightness values L00-Lmn, then method 550 can move to step 566. In step 566, the smoothing control module 210 can set each of the sorted brightness values L00-Lmn to an interval according to the brightness interval. Next, in step 568, the smoothing control module 210 can set the smoothing dehazing control intensity for multiple brightness intervals. The smoothing dehazing control intensity can be determined based on the signal DFG-WGT. Next, method 550 can move to step 570. Step 570 can end method 550.
[0196] refer to Figure 13The diagram illustrates method (or process) 600. Method 600 can set the defogging intensity. Method 600 typically includes steps (or states) 602, 604, 606, a decision step (or state) 608, 610, 612, 614, 616, and 618.
[0197] Step 602 can initiate method 600. In step 604, the smoothing control module 210 can receive a luminance distribution map 208 with sorted luminance intervals. In some embodiments, the luminance interval control module 206 can perform the sorting of the intervals of the luminance distribution map 208 from darkest to brightest. Next, in step 606, the smoothing control module 210 can receive a dehazing intensity control point. The dehazing intensity control point can be provided by a signal DFG-WGT. In this example, the signal DFG-WGT can be an input parameter for the video dehazing module 200. Next, method 600 can move to decision step 608.
[0198] In decision step 608, the smoothing control module 210 determines whether the dehazing intensity control point increases or decreases the dehazing intensity. If the dehazing intensity control point increases the dehazing intensity, method 600 moves to step 610. In step 610, it is determined that dehazing is used to remove more of the blurring effect caused by fog and reduce the brightness of the dehazed area. Next, method 600 moves to step 614. In decision step 608, if the dehazing intensity control point decreases the dehazing intensity, method 600 moves to step 612. In step 612, it is determined that dehazing is used to remove less of the blurring effect caused by fog and increase the brightness of the dehazed area. Next, method 600 moves to step 614.
[0199] In step 614, the smoothing control module 210 can use Bezier curve fitting control to perform adaptive smoothing of the dehazing strength control points to avoid brightness differences in regions 306aa-306mn. Next, in step 616, the video dehazing module 200 can generate dehazed video frames in the signal FRM-DFG, wherein there is a smooth transition between each region in the output video frame's region. Next, method 600 can move to step 618. Step 618 can end method 600.
[0200] refer to Figure 14The diagram illustrates method (or process) 650. Method 650 can generate dehazed video frames. Method 650 typically includes step (or state) 652, decision step (or state) 654, step (or state) 656, step (or state) 658, step (or state) 660, step (or state) 662, step (or state) 664, step (or state) 666, step (or state) 668, step (or state) 670, and step (or state) 672.
[0201] Step 652 can initiate method 650. Next, in decision step 654, video dehazing module 200 can determine whether any more video frames 202a-202n exist. Video frames 202a-202n can be provided by the signal FRAMES. If more video frames 202a-202n exist, method 650 can move to step 656. In step 656, video dehazing module 200 can receive the next input video frame 202a-202n. In step 658, low-pass filter 204 can perform low-pass filtering on the current video frame in video frames 202a-202n to generate low-frequency layer 300. Next, method 650 can move to steps 660 and 662. For example, steps 662-666 and step 660 can be performed in parallel and / or substantially in parallel.
[0202] In step 660, detail layer 216 can be generated. Detail layer 216 can be generated in response to removing low-frequency layer 300 from the current video frame in video frames 202a-202n (e.g., foggy input video frame 250). Summation module 214 can perform a subtraction operation to remove low-frequency layer 300 from video frames 202a-202n. Next, method 650 can move to step 668.
[0203] In step 662, the smoothing control module 210 can generate adaptively smoothed dehazing intensity weights based on the luminance distribution map 208. For example, the luminance interval control module 206 can generate the luminance distribution map 208, and the smoothing control module 210 can generate adaptively smoothed dehazing intensity weights based on the luminance interval. Next, in step 664, the video dehazing module 200 can determine the dehazing result (e.g., the signal DEFOG) based on the adaptively smoothed dehazing intensity weights and the low-frequency layer 300. For example, the multiplication module 212 can perform a multiplication operation between the adaptively smoothed dehazing weights in the SMTH-DFG signal and the low-frequency layer in the LFL signal. Next, method 650 can move to step 666.
[0204] In step 666, summing module 218 may remove the dehazing result from the current video frame (e.g., the foggy input video frame 250) in video frames 202a-202n. For example, the dehazing result may be provided in the DEFOG signal. Summing module 218 may be configured to subtract the dehazing result from the foggy input video frame. Next, in step 668, summing module 218 may use detail layer 216 to recover lost high-frequency details. For example, removing the dehazing result may result in some high-detail loss from video frames 202a-202n. Summing module 218 may be configured to perform an addition operation between the current video frame in video frames 202a-202n that has had its dehazing result removed and detail layer 216. In step 670, video dehazing module 200 may output a dehazed video frame based on the current video frame in video frames 202a-202n. Next, method 650 may return to decision step 654. In decision step 654, if more foggy input video frames exist, method 650 can repeat steps 656-670. If no more video frames 202a-202n exist, method 650 can move to step 672. Step 672 can end method 650.
[0205] Depend on Figure 1-14 The functions executed by the diagram can be implemented using one or more of a conventional general-purpose processor, digital computer, microprocessor, microcontroller, RISC (Reduced Instruction Set Computer) processor, CISC (Complex Instruction Set Computer) processor, SIMD (Single Instruction Multiple Data) processor, signal processor, central processing unit (CPU), arithmetic logic unit (ALU), video digital signal processor (VDSP), and / or similar computing machines programmed according to the teachings of this specification, as will be apparent to those skilled in the art. A skilled programmer can readily prepare appropriate software, firmware, code, routines, instructions, opcodes, microcode, and / or program modules based on the teachings of this disclosure, again as will be apparent to those skilled in the art. The software is typically executed from one or more media by one or more processors in a machine-implemented manner.
[0206] The present invention can also be implemented by preparing an ASIC (Application-Specific Integrated Circuit), a platform ASIC, an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), a CPLD (Complex Programmable Logic Device), a sea-of-gate, an RFIC (Radio Frequency Integrated Circuit), an ASSP (Application-Specific Standard Product), one or more monolithic integrated circuits, one or more chips or dies arranged as flip-chip modules and / or multi-chip modules, or by interconnecting conventional component circuits through a suitable network, as described herein, modifications of which will be apparent to those skilled in the art.
[0207] Therefore, the present invention may also include a computer product, which may be one or more storage media and / or one or more transmission media, including instructions that can be used to program a machine to perform one or more processes or methods according to the present invention. The execution of the instructions contained in the computer product by the machine, and the operation of the surrounding circuitry, can convert input data into one or more files on the storage medium and / or one or more output signals representing physical objects or substances, such as audio and / or visual depictions. The execution of the instructions contained in the computer product by the machine can be performed on data stored on the storage medium and / or user input and / or in combination with values generated by a random number generator implemented by the computer product. The storage medium may include, but is not limited to, any type of disk, including floppy disks, hard disks, magnetic disks, optical disks, CD-ROMs, DVDs, and magneto-optical disks, as well as circuits such as ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), UVPROM (Ultraviolet Erasable Programmable ROM), flash memory, magnetic cards, optical cards, and / or any type of medium suitable for storing electronic instructions.
[0208] The elements of this invention can form part or all of one or more devices, units, components, systems, machines, and / or apparatuses. Devices may include, but are not limited to, servers, workstations, storage array controllers, storage systems, personal computers, laptop computers, notebook computers, handheld computers, cloud servers, personal digital assistants, portable electronic devices, battery-powered devices, set-top boxes, encoders, decoders, transcoders, compressors, decompressors, preprocessors, post-processors, transmitters, receivers, transceivers, cryptographic circuits, cellular phones, digital cameras, positioning and / or navigation systems, medical devices, head-up displays, wireless devices, audio recording, audio storage and / or audio playback devices, video recording, video storage and / or video playback devices, gaming platforms, peripheral devices, and / or multi-chip modules. Those skilled in the art will understand that the elements of this invention can be implemented in other types of devices to meet specific application standards.
[0209] When used herein in conjunction with "is" and verbs, the terms "may" and "usually" are intended to convey the intention that the description is exemplary and is considered broad enough to encompass both the specific examples presented in this disclosure and the alternative examples that may be derived based on this disclosure. As used herein, the terms "may" and "usually" should not be construed as necessarily implying the desirability or possibility of omitting the corresponding element.
[0210] When the notations “a”-“n” are used herein for various components, modules, and / or circuits, a single component, module, and / or circuit or multiple such components, modules, and / or circuits are disclosed, wherein the notation “n” is applied to represent any particular integer. Each different component, module, and / or circuit having an instance (or event) notated as “a”-“n” may indicate that the different component, module, and / or circuit may have a matching number of instances or a different number of instances. An instance designated as “a” may represent the first of multiple instances, and an instance “n” may refer to the last of multiple instances without implying a specific number of instances.
[0211] Although the invention has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention.
Claims
1. An apparatus comprising: An interface configured to receive pixel data from the environment; as well as A processor configured to: (i) process the pixel data arranged as video frames; (ii) Generating a luminance distribution map of the video frame in response to a low-pass filtering operation; (iii) Determining a plurality of dehazing intensity weights for the luminance distribution map; (iv) Performing adaptive smoothing on each of the plurality of dehazing intensity weights; and (v) Generating a dehazed video frame in response to (a) the video frame and (b) the plurality of dehazing intensity weights having the adaptive smoothing, wherein, (a) Each of the plurality of dehazing intensity weights corresponds to one of the plurality of brightness intervals in the brightness distribution map, and (b) The adaptive smoothing is configured to prevent brightness differences in the dehazed video frames.
2. The apparatus according to claim 1, wherein, The adaptive smoothing includes Bézier curve fitting control.
3. The apparatus according to claim 2, wherein, The order of the Bézier curve fitting control is selected to provide a trade-off between eliminating the brightness difference and the complexity of the operation used to perform the Bézier curve fitting control.
4. The apparatus according to claim 1, wherein, The processor is configured to determine the plurality of dehazing intensity weights in response to: (i) dividing the brightness distribution map into a plurality of regions; and (ii) extracting a corresponding brightness value from each of the plurality of regions. (iii) Sort each of the corresponding brightness values from darkest to brightest to determine the plurality of brightness intervals.
5. The apparatus according to claim 4, wherein, The brightness difference is prevented in response to a smooth transition of defogging between each of the plurality of regions.
6. The apparatus according to claim 5, wherein, The brightness difference is prevented to avoid high contrast in local areas.
7. The apparatus according to claim 4, wherein, The adaptive smoothing is configured to respond to the non-uniformity of fog in the video frame based on: (i) the image location of the plurality of regions; and (ii) the brightness distribution.
8. The apparatus according to claim 4, wherein, (i) the plurality of regions of the luminance distribution map includes rectangular regions of the video frame, and (ii) the luminance value is determined for each rectangular region of the rectangular regions.
9. The apparatus according to claim 4, wherein, The number of the plurality of brightness intervals is determined in response to the following: the range of each brightness value among the corresponding brightness values from each of the plurality of regions divided by the brightness interval value.
10. The apparatus according to claim 4, wherein, Each of the plurality of regions is a 16x16 pixel rectangle of the video frame.
11. The apparatus according to claim 1, wherein, Determining the multiple defogging intensity weights and performing the adaptive smoothing enables the defogging intensity to be controlled according to real-time changes in fog conditions.
12. The apparatus according to claim 11, wherein, The plurality of defogging intensity weights are configured to control the amount of defogging intensity.
13. The apparatus according to claim 11, wherein, The amount of defogging intensity can be adjusted based on a value selected for the weight of the defogging intensity.
14. The apparatus according to claim 13, wherein, (i) Adjusting the amount of the dehazing intensity determines the amount of fog remaining in the dehazed video frame, and (ii) increasing the amount of the dehazing intensity reduces the brightness in the dehazed video frame.
15. The apparatus according to claim 14, wherein, Reducing the brightness results in a decrease in the contrast of details in the dark areas of the dehazed video frame.
16. The apparatus according to claim 1, wherein, The multiple dehazing intensity weights are configured to remove the image blurring effect caused by fog captured in the video frame.
17. The apparatus according to claim 1, wherein, The dehazed video frames are generated in response to a high-performance and low-complexity image dehazing technique based on image location and brightness distribution.
18. The apparatus according to claim 1, wherein, (i) the low-pass filtering operation is configured to separate the high-frequency layer from each video frame in the video frame, (ii) image processing is performed using the plurality of dehazing intensity weights to generate a dehazing result, and (iii) the high-frequency layer is added to the dehazing result to generate the dehazed video frame.
19. The apparatus according to claim 18, wherein, (i) The image processing includes multiplying the plurality of dehazing intensity weights with a low-frequency layer to generate the dehazing result, wherein the plurality of dehazing intensity weights have the adaptive smoothing, and the low-frequency layer is generated by the low-pass filtering operation; (ii) subtracting the dehazing result from the video frame resulting in a loss of detail; and (iii) adding the high-frequency layer to recover the loss of detail.
20. The apparatus according to claim 1, wherein, Each of the following items is an adjustable parameter used to generate the dehazed video frame: (i) the cutoff frequency for the low-pass filtering operation, (ii) the strength of the plurality of dehazing intensity weights, and (iii) the number of the plurality of luminance intervals.