Smart auto exposure control for rgb-ir sensors
Pixel data is generated by a structured light projector and an image sensor. The processor analyzes the IR and RGB images, adjusts the sensor control signal, solves the problem of inconsistent brightness of the RGB-IR sensor under different lighting conditions, realizes intelligent automatic exposure control, and improves image quality.
Patent Information
- Application Number
- CN202210373846.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-07
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2042-04-07
AI Technical Summary
RGB-IR sensors exhibit inconsistent brightness between RGB and IR images under different lighting conditions, leading to underexposure or overexposure. Existing automatic exposure control methods cannot effectively adjust this, thus affecting image quality.
The system uses a structured light projector and an image sensor to generate pixel data. The processor analyzes the IR and RGB images and adjusts the sensor control signal to achieve intelligent automatic exposure control. It selects either digital gain priority or shutter priority operation mode based on the amount of IR interference and coordinates the exposure adjustment of the RGB and IR channels.
It achieves brightness consistency between RGB and IR images under different lighting conditions, prevents overexposure, improves image quality, overcomes sensor physical limitations, and provides intelligent automatic exposure control.
Smart Images

Figure CN116962886B_ABST
Abstract
Description
Technical Field
[0001] In general, the present invention relates to video and image capture, and more specifically, to a method and / or apparatus for implementing intelligent automatic exposure control for RGB-IR sensors. Background Technology
[0002] Machine vision, optical technology, and artificial intelligence have developed rapidly. Due to advancements in robotics and deep learning, 3D reconstruction has become a significant branch of machine vision. One method for active 3D reconstruction is to generate a depth map using a monocular speckle structured light system. RGB-IR sensors can be used to capture infrared data, color data, and monocular speckle structured light.
[0003] An RGB-IR sensor is a single camera sensor that can be exposed to both visible (RGB) and infrared (IR) light. RGB and IR images are captured by a single sensor. Due to physical limitations, the camera's sensor controls (i.e., shutter speed and gain) are shared by the RGB and IR channels. In various environments, the illumination conditions for visible and IR light differ, resulting in inconsistent brightness in the simultaneously captured RGB and IR images. Fixed auto-exposure (AE) adjustment methods have drawbacks when handling different illumination conditions on different RGB and IR channels.
[0004] Furthermore, the IR channel is highly sensitive to IR interference. If AE control is inadequate, even a small amount of infrared light (i.e., from sunlight) can easily cause overexposure of the IR image. Without proper AE control, the RGB channel may be underexposed when the shutter speed is too low. Balancing the brightness of the RGB and IR image channels under various lighting conditions becomes a challenge for RGB-IR sensors.
[0005] Achieving intelligent automatic exposure control for RGB-IR sensors would be desirable. Summary of the Invention
[0006] This invention relates to an apparatus comprising a structured light projector, an image sensor, and a processor. The structured light projector can be configured to switch a structured light pattern in response to a timing signal. The image sensor can be configured to generate pixel data. The processor can be configured to process the pixel data arranged as video frames; extract an IR image having the structured light pattern, an IR image without the structured light pattern, and an RGB image from the video frames; generate a sensor control signal in response to the IR image having the structured light pattern and the RGB image; calculate an IR interference measurement in response to the IR image without the structured light pattern and the RGB image; and select an operating mode for IR channel control and RGB channel control in response to the IR interference measurement. The sensor control signal can be configured to adjust the exposure of the image sensor for both the IR image having the structured light pattern and the RGB image. The IR channel control can be configured to adjust the exposure of the IR image having the structured light pattern. The RGB channel control can be configured to adjust the exposure of both the RGB image and the IR image without the structured light pattern. Attached Figure Description
[0007] Embodiments of the invention will become apparent from the following detailed description, as well as the appended claims and drawings.
[0008] Figure 1 This is a diagram illustrating an example of an edge device that can utilize a processor configured to implement intelligent automatic exposure control for an RGB-IR sensor, according to an exemplary embodiment of the present invention.
[0009] Figure 2 This is a diagram illustrating an example camera that implements an example embodiment of the present invention.
[0010] Figure 3 This is a block diagram showing the camera system.
[0011] Figure 4 This is a diagram showing the processing circuitry of a camera system configured to perform 3D reconstruction using a convolutional neural network.
[0012] Figure 5 This is a block diagram illustrating an intelligent automatic exposure control system.
[0013] Figure 6 This diagram illustrates a processor configured to extract IR video frames with structured light patterns, IR video frames without structured light patterns, and RGB video frames.
[0014] Figure 7This is a block diagram illustrating an IR interference module configured to determine the amount of IR interference in an environment.
[0015] Figure 8 This is a block diagram showing how IR channel control and RGB channel control are adjusted in response to statistics measured from video frames.
[0016] Figure 9 This is a graph showing the exposure results of a video frame.
[0017] Figure 10 This is a flowchart illustrating a method for performing intelligent automatic exposure control for RGB-IR sensors.
[0018] Figure 11 This is a flowchart illustrating a method for performing IR interference measurements.
[0019] Figure 12 This is a flowchart illustrating a method for automatic exposure control in shutter priority operation mode.
[0020] Figure 13 This is a flowchart illustrating a method for automatic exposure control in digital gain-priority operating mode.
[0021] Specific implementation
[0022] Embodiments of the present invention include providing intelligent automatic exposure control for an RGB-IR sensor that can (i) automatically adjust sensor control, the IR channel, and the RGB channel; (ii) generate statistics on the IR image channel and the RGB image channel; (iii) provide consistent brightness for the IR image and the RGB image across various lighting conditions; (iv) enable parameter tuning based on user input; (v) prevent overexposure of the IR image due to IR interference; (vi) select different automatic exposure operating modes in response to the amount of detected IR interference; (vii) compare the IR image without a structured light pattern with the RGB image to determine the amount of IR interference in the environment; (viii) overcome the physical limitations of the RGB-IR sensor and / or (ix) be implemented as one or more integrated circuits.
[0023] Embodiments of the present invention can be configured to implement an intelligent and dynamic automatic exposure (AE) adjustment system for RGB-IR sensors. Instead of performing automatic exposure separately for each channel, the implemented AE adjustment can automatically adjust various exposure settings in response to RGB and IR statistics. The implemented AE adjustment can be configured to simultaneously coordinate sensor control, RGB channel adjustment, and IR channel adjustment.
[0024] Embodiments of the present invention can be implemented by dividing the AE adjustment into sensor control, IR channel control, and RGB channel control. The RGB and IR channels can share the same sensor control logic. Alternatively, the RGB and IR channels can each have separate channel control logic. The channel control logic can be configured to automatically adjust exposure parameters (e.g., digital gain, tone curve adjustment, other post-processing parameters, etc.). By implementing separate channel control logic for the RGB and IR channels, RGB and IR images can be tuned to similar brightness levels even under conditions of different visible and IR light.
[0025] Embodiments of the present invention can be configured to operate in multiple different operating modes to avoid IR interference sensitivity issues. The amount of IR interference can be detected. In response to the detected amount of IR interference, AE adjustment can be configured to determine whether to operate in a digital gain-priority operating mode or a shutter-priority operating mode. An aggressiveness value can be received as user input for additional manual tuning to determine how the operating mode is applied and / or takes effect.
[0026] In the example, under bright sunlight conditions (e.g., IR interference from sunlight might cause overexposure on an RGB-IR sensor without intelligent AE adjustment), the AE adjustment system can select a digital gain-priority operating mode. This mode allows the shutter speed to be adjusted to a low value (e.g., depending on the advance value) and allows the digital gain to increase image brightness first. This mode also allows the IR channels to reduce the effects of IR interference from sunlight. IR fill light with potentially high output power (e.g., structured light patterns) might still be exposed in the IR channels. Reducing the amount of sunlight might result in a darker RGB image. However, the AE adjustment system can adjust both the digital gain and tone curves to enhance brightness in the RGB image channels. Once the digital gain reaches its maximum, the shutter speed might need to be increased if adjusting the digital gain alone is insufficient.
[0027] In another example, in a low IR interference environment, the AE adjustment system can select a shutter-priority operation mode. Shutter-priority operation mode allows adjustment of the shutter speed to be longer than the shutter speed in digital gain-priority operation mode. For example, a longer shutter speed exposes the RGB-IR sensor to more lines of IR illumination (e.g., structured light patterns). Once the shutter speed reaches its maximum value, digital gain and tone mapping can be performed if adjusting the shutter speed alone is insufficient.
[0028] The AE system can be configured to determine when to switch between a digital gain-priority operating mode and a shutter-priority operating mode in response to the detected amount of IR interference. Embodiments of the invention can be configured to measure the amount of IR interference in the environment. In an example, an RGB image and an IR image (e.g., captured when a structured light pattern is turned off) can be used as input to determine the amount of IR interference. The average intensity of the two images can be calculated. For example, if the intensity of the IR image divided by the intensity of the RGB image exceeds a predetermined threshold, the IR interference can be determined to be high. In response to whether the IR interference is determined to be high, the emitting system selects an operating mode. By selecting an appropriate operating mode based on the IR interference conditions, the AE system enables the RGB-IR sensor to have consistent performance under various illumination conditions of visible and IR light. Providing consistent performance under various IR and visible light conditions overcomes the physical limitations of RGB-IR sensors.
[0029] Embodiments of the present invention can provide automatic exposure control for both RGB and IR channels. Automatic exposure control can be based on channel statistics. Different automatic exposure control strategies can be selected for different lighting scenarios (e.g., the contrast between visual light intensity and IR light intensity). Automatic exposure can provide intelligent automatic exposure control for RGB-IR sensors.
[0030] refer to Figure 1 , Figure 1 A diagram illustrates an example edge device according to an exemplary embodiment of the invention, which can utilize a processor configured to implement intelligent automatic exposure control for an RGB-IR sensor. A top view of region 50 is shown. In the illustrated example, region 50 may be an outdoor location. Streets, vehicles, and buildings are shown.
[0031] Devices 100a-100n are shown at different locations within region 50. Each of devices 100a-100n can independently implement an edge device. Edge devices 100a-100n may include smart IP cameras (e.g., camera systems). Edge devices 100a-100n may include low-power technologies (e.g., microprocessors running on sensors, cameras, or other battery-powered devices) in embedded platforms designed for deployment at the network edge, where power consumption is a critical issue. In the examples, edge devices 100a-100n may include various traffic cameras and Intelligent Transportation System (ITS) solutions.
[0032] Edge devices 100a-100n can be implemented for a variety of applications. In the illustrated example, edge devices 100a-100n may include an automatic license plate recognition (ANPR) camera 100a, a traffic camera 100b, a vehicle camera 100c, an access control camera 100d, an automatic teller machine (ATM) camera 100e, a bullet camera 100f, a dome camera 100n, etc. In the example, edge devices 100a-100n can be implemented as a traffic camera and Intelligent Transportation System (ITS) solution designed to enhance road safety by utilizing a combination of people and vehicle detection, vehicle brand / model recognition, and automatic license plate recognition (ANPR) functions.
[0033] In the illustrated example, region 50 can be an outdoor location. In some embodiments, edge devices 100a-100n can be implemented in various indoor locations. In the example, edge devices 100a-100n can incorporate convolutional neural networks for use in security (surveillance) applications and / or access control applications. In the example, edge devices 100a-100n implemented as security cameras and access control applications can include battery-powered cameras, doorbell cameras, outdoor cameras, indoor cameras, etc. According to embodiments of the invention, security camera and access control applications can achieve performance advantages from the application of convolutional neural networks. In the example, edge devices utilizing convolutional neural networks according to embodiments of the invention can acquire large amounts of image data and perform on-device inference to obtain useful information (e.g., multiple temporal instances of images performed by each network), thereby reducing bandwidth and / or power consumption. The design, type, and / or application performed by edge devices 100a-100n can vary depending on the design criteria of a particular implementation.
[0034] refer to Figure 2 The diagram illustrates an example edge device camera that implements an exemplary embodiment of the present invention. Camera systems 100a-100n are shown. Each camera device 100a-100n can have a different style and / or use case. For example, camera 100a can be an action camera, camera 100b can be a ceiling-mounted security camera, camera 100n can be a webcam, etc. Other types of cameras can be implemented (e.g., home security cameras, battery-powered cameras, doorbell cameras, stereo cameras, etc.). The design / style of cameras 100a-100n can vary depending on the design criteria of a particular implementation.
[0035] Each of the camera systems 100a-100n may include block (or circuit) 102, block (or circuit) 104, and / or block (or circuit) 106. Circuit 102 may implement a processor. Circuit 104 may implement a capture device. Circuit 106 may implement a structured light projector. Camera systems 100a-100n may include other components (not shown). They can be used with... Figure 3 The details of the components of the camera 100a-100n are described in relation to each other.
[0036] Processor 102 can be configured to implement an artificial neural network (ANN). In an example, the ANN may include a convolutional neural network (CNN). Processor 102 can be configured to implement a video encoder. Processor 102 can be configured to process pixel data arranged into video frames. Capture device 104 can be configured to capture pixel data that can be used by processor 102 to generate video frames. Structured light projector 106 can be configured to generate a structured light pattern (e.g., a speckle pattern). The structured light pattern can be projected onto a background (e.g., the environment). Capture device 104 can capture pixel data including a background image (e.g., the environment) with a speckle pattern.
[0037] Cameras 100a-100n can be edge devices. A processor 102 implemented by each of the cameras 100a-100n enables the cameras 100a-100n to perform various functions internally (e.g., at the local level). For example, processor 102 can be configured to perform on-device object / event detection (e.g., computer vision operations), 3D reconstruction, liveness detection, depth map generation, video encoding, and / or video transcoding. For example, processor 102 can even perform advanced processes such as computer vision and 3D reconstruction without uploading video data to a cloud service to offload computationally intensive functions (e.g., computer vision, video encoding, video transcoding, etc.).
[0038] In some embodiments, multiple camera systems may be implemented (e.g., camera systems 100a-100n may operate independently of each other). For example, each of cameras 100a-100n may individually analyze captured pixel data and perform event / object detection locally. In some embodiments, cameras 100a-100n may be configured as a camera network (e.g., security cameras that send video data to a central source such as network-attached storage and / or cloud services). The location and / or configuration of cameras 100a-100n may vary depending on the design criteria of the specific implementation.
[0039] The capture device 104 of each of the camera systems 100a-100n may include a single lens (e.g., a monocular camera). The processor 102 may be configured to accelerate the preprocessing of speckle structured light for monocular 3D reconstruction. Monocular 3D reconstruction can be performed to generate depth maps and / or parallax images without using a stereo camera.
[0040] refer to Figure 3 The diagram illustrates a block diagram of a camera system 100 implemented as an example. The camera system 100 can be a combination of... Figure 2 Representative examples of cameras 100a-100n are shown. Camera system 100 may include processor / SoC 102, capture device 104, and structured light projector 106.
[0041] The camera system 100 may further include blocks (or circuits) 150, 152, 154, 156, 158, 160, 162, 164, and / or 166. Circuit 150 may implement a memory. Circuit 152 may implement a battery. Circuit 154 may implement a communication device. Circuit 156 may implement a wireless interface. Circuit 158 may implement a general-purpose processor. Block 160 may implement an optical lens. Block 162 may implement a structured light patterning lens. Circuit 164 may implement one or more sensors. Circuit 166 may implement a human-machine interface (HID) device. In some embodiments, the camera system 100 may include a processor / SoC 102, a capture device 104, an IR structured light projector 106, a memory 150, a lens 160, an IR structured light projector 106, a structured light pattern lens 162, a sensor 164, a battery 152, a communication module 154, a wireless interface 156, and a processor 158. In another example, the camera system 100 may include the processor / SoC 102, the capture device 104, the structured light projector 106, the processor 158, the lens 160, the structured light pattern lens 162, and the sensor 164 as a single device, and the memory 150, the battery 152, the communication module 154, and the wireless interface 156 may be components of separate devices. The camera system 100 may include other components (not shown). The number, type, and / or arrangement of the components of the camera system 100 may vary depending on the design criteria of a particular implementation.
[0042] Processor 102 can be implemented as a video processor. In an example, processor 102 can be configured to receive three-sensor video input using a high-speed SLVS / MIPI-CSI / LVCMOS interface. In some embodiments, processor 102 can be configured to perform depth sensing in addition to generating video frames. In an example, depth sensing can be performed in response to depth information and / or vector light data captured in the video frames.
[0043] Memory 150 can store data. Memory 150 can be implemented in various types of memory, including but not limited to cache, flash memory, memory cards, random access memory (RAM), dynamic RAM (DRAM), etc. The type and / or size of memory 150 can vary depending on the design criteria of a particular implementation. The data stored in memory 150 may correspond to video files, motion information (e.g., readings from sensor 164), video fusion parameters, image stabilization parameters, user input, computer vision models, feature sets, and / or metadata information. In some embodiments, memory 150 can store reference images. Reference images can be used for computer vision operations, 3D reconstruction, etc. In some embodiments, reference images may include reference structured light images.
[0044] The processor / SoC 102 can be configured to execute computer-readable code and / or process information. In various embodiments, the computer-readable code may be stored within the processor / SoC 102 (e.g., microcode, etc.) and / or in memory 150. In an example, the processor / SoC 102 can be configured to execute one or more artificial neural network models (e.g., face recognition CNN, object detection CNN, object classification CNN, 3D reconstruction CNN, liveness detection CNN, etc.) stored in memory 150. In an example, memory 150 may store one or more directed acyclic graphs (DAGs) and one or more sets of weights and biases defining one or more artificial neural network models. The processor / SoC 102 can be configured to receive input from memory 150 and / or present output to memory 150. The processor / SoC 102 can be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of the processor / SoC 102 may vary depending on the design criteria of a particular implementation. The processor / SoC 102 can be configured for low-power (e.g., battery) operation.
[0045] Battery 152 can be configured to store and / or power components of camera system 100. The dynamic driver mechanism for the rolling shutter sensor can be configured to conserve power. Reduced power consumption allows camera system 100 to operate for extended periods using battery 152 without recharging. Battery 152 can be rechargeable. Battery 152 can be built-in (e.g., non-replaceable) or replaceable. Battery 152 can have an input for connecting to an external power source (e.g., for charging). In some embodiments, device 100 can be powered by an external power source (e.g., battery 152 can be omitted or implemented as a backup power source). Various battery technologies and / or chemistry can be used to implement battery 152. The type of battery 152 implemented can vary depending on the design criteria of a particular implementation.
[0046] The communication module 154 can be configured to implement one or more communication protocols. For example, the communication module 154 and the wireless interface 156 can be configured to implement one or more of the following: IEEE 102.11, IEEE 102.15, IEEE 102.15.1, IEEE 102.15.2, IEEE 102.15.3, IEEE 102.15.4, IEEE 102.15.5, IEEE 102.20, etc. and / or In some embodiments, the communication module 154 may be a hardwired data port (e.g., a USB port, a mini-USB port, a USB-C connector, an HDMI port, an Ethernet port, a DisplayPort interface, a Lightning port, etc.). In some embodiments, the wireless interface 156 may also implement one or more protocols associated with a cellular communication network (e.g., GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, SMS, etc.). In embodiments where the camera system 100 is implemented as a wireless camera, the protocol implemented by the communication module 154 and the wireless interface 156 may be a wireless communication protocol. The type of communication protocol implemented by the communication module 154 may vary depending on the design criteria of the specific implementation.
[0047] Communication module 154 and / or wireless interface 156 can be configured to generate broadcast signals as output from camera system 100. The broadcast signals can send video data, parallax data, and / or control signals to external devices. For example, the broadcast signals can be sent to cloud storage services (e.g., storage services capable of scaling on demand). In some embodiments, communication module 154 may not send data until the processor / SoC 102 has performed video analysis to determine that an object is in the field of view of camera system 100.
[0048] In some embodiments, the communication module 154 can be configured to generate a manual control signal. The manual control signal can be generated in response to a signal received from a user by the communication module 154. The manual control signal can be configured to activate the processor / SoC 102. Regardless of the power state of the camera system 100, the processor / SoC 102 can be activated in response to the manual control signal.
[0049] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive a feature set. The received feature set can be used to detect events and / or objects. For example, the feature set can be used to perform computer vision operations. The feature set information may include instructions for the processor 102 to determine which types of objects correspond to objects and / or events of interest.
[0050] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive user input. User input allows the user to adjust operating parameters of various features implemented by the processor 102. In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to interface with an application (e.g., an app) (e.g., using an application programming interface (API)). For example, the application may be implemented on a smartphone to enable the end user to adjust various settings and / or parameters for various features implemented by the processor 102 (e.g., setting video resolution, selecting frame rate, selecting output format, setting tolerance parameters for 3D reconstruction, etc.).
[0051] Processor 158 can be implemented using general-purpose processor circuitry. Processor 158 is operable to interact with video processing circuitry 102 and memory 150 to perform various processing tasks. Processor 158 can be configured to execute computer-readable instructions. In this example, the computer-readable instructions may be stored in memory 150. In some embodiments, the computer-readable instructions may include controller operations. Typically, input from sensor 164 and / or human-machine interface device 166 is shown to be received by processor 102. In some embodiments, general-purpose processor 158 can be configured to receive and / or analyze data from sensor 164 and / or HID 166 and make decisions in response to the input. In some embodiments, processor 158 can send data to and / or receive data from other components of camera system 100, such as battery 152, communication module 154, and / or wireless interface 156. Which functions of camera system 100 are performed by processor 102 and general-purpose processor 158 may vary depending on the design criteria of the specific implementation.
[0052] Lens 160 may be attached to capture device 104. Capture device 104 may be configured to receive an input signal (e.g., LIN) via lens 160. The signal LIN may be an optical input (e.g., an analog image). Lens 160 may be implemented as an optical lens. Lens 160 may provide zoom and / or focus features. In one example, capture device 104 and / or lens 160 may be implemented as a single lens assembly. In another example, lens 160 may be implemented separately from capture device 104.
[0053] Capture device 104 can be configured to convert input light LIN into computer-readable data. Capture device 104 can capture data received through lens 160 to generate raw pixel data. In some embodiments, capture device 104 can capture data received through lens 160 to generate a bitstream (e.g., generate video frames). For example, capture device 104 can receive focused light from lens 160. Lens 160 can be oriented, tilted, translated, scaled, and / or rotated to provide a target view from camera system 100 (e.g., a view of video frames, a view of panoramic video frames captured using multiple camera systems 100a-100n, a target image and reference image view for stereo vision, etc.). Capture device 104 can generate a signal (e.g., video). The signal VIDEO can be pixel data (e.g., a sequence of pixels that can be used to generate video frames). In some embodiments, the signal VIDEO can be video data (e.g., a sequence of video frames). The signal VIDEO can be presented to one of the inputs of processor 102. In some embodiments, the pixel data generated by the capture device 104 may be uncompressed and / or raw data generated in response to focused light from the lens 160. In some embodiments, the output of the capture device 104 may be a digital video signal.
[0054] In this example, capture device 104 may include block (or circuitry) 180, block (or circuitry) 182, and block (or circuitry) 184. Circuitry 180 may be an image sensor. Circuitry 182 may be a processor and / or logic unit. Circuitry 184 may be memory circuitry (e.g., a frame buffer). Lens 160 (e.g., a camera lens) may be directed to provide a view of the environment surrounding camera system 100. Lens 160 may be designed to capture ambient data (e.g., light input LIN). Lens 160 may be a wide-angle lens and / or a fisheye lens (e.g., a lens capable of capturing a wide field of view). Lens 160 may be configured to capture and / or focus light for capture device 104. Typically, image sensor 180 is located behind lens 160. Based on the light captured from lens 160, capture device 104 may generate bitstream and / or video data (e.g., a signal VIDEO).
[0055] Capture device 104 can be configured to capture video image data (e.g., light collected and focused by lens 160). Capture device 104 can capture data received through lens 160 to generate a video bitstream (e.g., pixel data of a video frame sequence). In various embodiments, lens 160 can be implemented as a fixed-focus lens. Fixed-focus lenses are generally advantageous for smaller size and lower power. In examples, fixed-focus lenses can be used in battery-powered, doorbell, and other low-power camera applications. In some embodiments, lens 160 can be oriented, tilted, panned, zoomed, and / or rotated to capture the environment around camera system 100 (e.g., capture data from the field of view). In examples, professional camera models can be implemented using an active lens system for enhanced functionality, remote control, etc.
[0056] The capture device 104 can convert received light into a digital data stream. In some embodiments, the capture device 104 can perform analog-to-digital conversion. For example, the image sensor 180 can perform photoelectric conversion on the light received by the lens 160. The processor / logic unit 182 can convert the digital data stream into a video data stream (or bitstream), a video file, and / or multiple video frames. In this example, the capture device 104 can present the video data as a digital video signal (e.g., VIDEO). The digital video signal can include video frames (e.g., continuous digital images and / or audio). In some embodiments, the capture device 104 can include a microphone for capturing audio. In some embodiments, the microphone can be implemented as a separate component (e.g., one of the sensors 164).
[0057] Video data captured by capture device 104 can be represented as a signal / bitstream / data VIDEO (e.g., a digital video signal). Capture device 104 can present the signal VIDEO to processor / SoC 102. The signal VIDEO can represent video frames / video data. The signal VIDEO can be a video stream captured by capture device 104. In some embodiments, the signal VIDEO may include pixel data operable by processor 102 (e.g., a video processing pipeline, image signal processor (ISP), etc.). Processor 102 can generate video frames in response to the pixel data in the signal VIDEO.
[0058] The signal VIDEO may include pixel data arranged as video frames. The signal VIDEO may be an image including a background (e.g., captured objects and / or environment) and a speckle pattern generated by the structured light projector 106. The signal VIDEO may include a single-channel source image. A single-channel source image may be generated in response to capturing pixel data using the monocular lens 160.
[0059] Image sensor 180 can receive input light LIN from lens 160 and convert the light LIN into digital data (e.g., a bitstream). For example, image sensor 180 can perform photoelectric conversion on light from lens 160. In some embodiments, image sensor 180 may have additional margins that are not used as part of the image output. In some embodiments, image sensor 180 may not have additional margins. In various embodiments, image sensor 180 can be configured to generate RGB-IR video signals. In a field of view illuminated only by infrared light, image sensor 180 can generate monochrome (B / W) video signals. In a field of view illuminated by both IR and visible light, image sensor 180 can be configured to generate color information in addition to generating monochrome video signals. In various embodiments, image sensor 180 can be configured to generate video signals in response to visible light and / or infrared (IR) light.
[0060] In some embodiments, the camera sensor 180 may include a rolling shutter sensor or a global shutter sensor. In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR sensor. In some embodiments, the capture device 104 may include a rolling shutter IR sensor and an RGB sensor (e.g., implemented as separate components). In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR rolling shutter complementary metal-oxide-semiconductor (CMOS) image sensor. In one example, the rolling shutter sensor 180 may be configured to assert a signal indicating the exposure time of the first row. In one example, the rolling shutter sensor 180 may apply a mask to a monochrome sensor. In an example, the mask may include multiple cells containing a red pixel, a green pixel, a blue pixel, and an IR pixel. The IR pixel may contain red, green, and blue filter materials that efficiently absorb all light in the visible spectrum while allowing longer infrared wavelengths to pass through with minimal loss. In the case of a rolling shutter, as each row (or line) of the sensor begins exposure, all pixels in that row (or line) may begin exposure simultaneously.
[0061] Processor / logic unit 182 can convert the bitstream into human-readable content (e.g., video data that an average person can understand regardless of image quality, such as video frames and / or pixel data that can be converted into video frames by processor 102). For example, processor / logic unit 182 can receive raw (e.g., raw) data from image sensor 180 and generate (e.g., encode) video data (e.g., bitstream) based on the raw data. Capture device 104 may have memory 184 to store raw data and / or processed bitstreams. For example, capture device 104 may implement frame memory and / or buffer 184 to store (e.g., provide temporary storage and / or cache) one or more video frames (e.g., digital video signals). In some embodiments, processor / logic unit 182 can perform analysis and / or correction on the video frames stored in memory / buffer 184 of capture device 104. Processor / logic unit 182 can provide status information about the captured video frames.
[0062] The structured light projector 106 may include a block (or circuit) 186. Circuit 186 may implement a structured light source. The structured light source 186 may be configured to generate a signal (e.g., a speckle pattern). The signal SLP may be a structured light pattern (e.g., a speckle pattern). The signal SLP may be projected onto the environment near the camera system 100. The structured light pattern SLP may be captured by a capture device 104 as part of a light input LIN.
[0063] The structured light patterning lens 162 can be a lens for the structured light projector 106. The structured light patterning lens 162 can be configured to allow the structured light SLP generated by the structured light source 186 of the structured light projector 106 to be emitted, while protecting the structured light source 186. The structured light patterning lens 162 can be configured to decompose the laser pattern generated by the structured light source 186 into a pattern array (e.g., a dense dot pattern array for speckle patterns).
[0064] In the example, the structured light source 186 can be implemented as an array of vertical cavity surface-emitting lasers (VCSELs) and lenses. However, other types of structured light sources can be implemented to meet the design criteria of specific applications. In the example, the VCSEL array is typically configured to generate laser patterns (e.g., signal SLP). The lenses are typically configured to decompose the laser pattern into an array of dense dot patterns. In the example, the structured light source 186 can be implemented as a near-infrared (NIR) light source. In various embodiments, the light source of the structured light source 186 can be configured to emit light with a wavelength of approximately 940 nanometers (nm), which is invisible to the human eye. However, other wavelengths can be utilized. In the example, wavelengths in the range of approximately 800 nm to 1000 nm can be utilized.
[0065] Sensor 164 can implement multiple sensors, including but not limited to motion sensors, ambient light sensors, proximity sensors (e.g., ultrasonic, radar, lidar, etc.), audio sensors (e.g., microphones), etc. In embodiments implementing a motion sensor, sensor 164 can be configured to detect motion anywhere (or some location outside the field of view) monitored by camera system 100. In various embodiments, motion detection can be used as a threshold for activating capture device 104. Sensor 164 can be implemented as an internal component of camera system 100 and / or an external component of camera system 100. In one example, sensor 164 can be implemented as a passive infrared (PIR) sensor. In another example, sensor 164 can be implemented as a smart motion sensor. In yet another example, sensor 164 can be implemented as a microphone. In embodiments implementing a smart motion sensor, sensor 164 may include a low-resolution image sensor configured to detect motion and / or people.
[0066] In various embodiments, sensor 164 may generate signals (e.g., SENS). The SENS may include various data (or information) collected by sensor 164. In an example, the SENS may include data collected in response to motion detected in the monitored field of view, ambient light levels in the monitored field of view, and / or sound picked up in the monitored field of view. However, other types of data may be collected and / or generated based on application-specific design criteria. The SENS may be presented to processor / SoC 102. In an example, sensor 164 may generate (assert) the SENS when motion is detected in the field of view monitored by camera system 100. In another example, sensor 164 may generate (assert) the SENS when audio is triggered in the field of view monitored by camera system 100. In yet another example, sensor 164 may be configured to provide directional information about motion and / or sound detected in the field of view. This directional information may also be transmitted to processor / SoC 102 via the SENS.
[0067] HID 166 can implement an input device. For example, HID 166 can be configured to receive human input. In one example, HID 166 can be configured to receive password input from a user. In another example, HID 166 can be configured to receive user input to provide various parameters and / or settings to processor 102 and / or memory 150. In some embodiments, camera system 100 may include a keyboard, touchpad (or screen), doorbell switch, and / or other human-machine interface device (HID) 166. In an example, sensor 164 can be configured to determine when an object approaches HID 166. In an example where camera system 100 is implemented as part of an access control application, capture device 104 can be activated to provide images for identifying a person attempting access, and a lock area and / or illumination for access touchpad 166 can be activated. For example, a combination of input from HID 166 (e.g., a password or PIN code) can be combined with activity determination and / or depth analysis performed by processor 102 to achieve two-factor authentication.
[0068] The processor / SoC 102 can receive a signal VIDEO and a signal SENS. The processor / SoC 102 can generate one or more video output signals (e.g., VIDOUT), one or more control signals (e.g., CTRL), and / or one or more depth data signals (e.g., DIMAGES) based on the signal VIDEO, signal SENS, and / or other inputs. In some embodiments, the signals VIDOUT, DIMAGES, and CTRL can be generated based on analysis of the signal VIDEO and / or objects detected in the signal VIDEO.
[0069] In various embodiments, the processor / SoC 102 may be configured to perform one or more of the following: feature extraction, object detection, object tracking, 3D reconstruction, liveness detection, and object recognition. For example, the processor / SoC 102 may determine motion information and / or depth information by analyzing frames from a signal VIDEO and comparing those frames with previous frames. The comparison may be used to perform digital motion estimation. In some embodiments, the processor / SoC 102 may be configured to generate a video output signal VIDOUT including video data and / or a depth data signal DIMAGES including a disparity map and a depth map from the signal VIDEO. The video output signal VIDOUT and / or the depth data signal DIMAGES may be presented to memory 150, communication module 154, and / or wireless interface 156. In some embodiments, the video signal VIDOUT and / or the depth data signal DIMAGES may be used internally by the processor 102 (e.g., not presented as output).
[0070] The signal VIDOUT can be presented to the communication device 156. In some embodiments, the signal VIDOUT may include encoded video frames generated by the processor 102. In some embodiments, the encoded video frames may include a complete video stream (e.g., encoded video frames representing all video captured by the capture device 104). The encoded video frames may be encoded, cropped, stitched, and / or enhanced versions of pixel data received from the signal VIDEO. In the example, the encoded video frames may be high-resolution, digital, encoded, de-distorted, stabilized, cropped, blended, stitched, and / or rolling shutter effect corrected versions of the signal VIDEO.
[0071] In some embodiments, the signal VIDOUT may be generated based on video analysis (e.g., computer vision operations) performed by processor 102 on generated video frames. Processor 102 may be configured to perform computer vision operations to detect objects and / or events in the video frames, and then convert the detected objects and / or events into statistical data and / or parameters. In one example, the data determined by the computer vision operations may be converted by processor 102 into a human-readable format. Data from the computer vision operations can be used to detect objects and / or events. The computer vision operations may be performed locally by processor 102 (e.g., without needing to communicate with external devices to offload computational operations). For example, locally executed computer vision operations allow the computer vision operations to be performed by processor 102 and avoid heavy video processing running on a backend server. Avoiding video processing on a backend (e.g., remote location) server can protect privacy.
[0072] In some embodiments, the signal VIDOUT can be data generated by processor 102 (e.g., video analysis results, audio / speech analysis results, etc.) that can be transmitted to a cloud computing service for information aggregation and / or to provide training data for machine learning (e.g., to improve object detection, improve audio detection, improve liveness detection, etc.). In some embodiments, the signal VIDOUT can be provided to a cloud service for mass storage (e.g., to enable users to retrieve encoded video using smartphones and / or desktop computers). In some embodiments, the signal VIDOUT can include data extracted from video frames (e.g., computer vision results) and can transmit the results to another device (e.g., a remote server, a cloud computing system, etc.) to offload the analysis of the results to another device (e.g., offloading the analysis of the results to a cloud computing service instead of performing all the analysis locally). The type of information transmitted by the signal VIDOUT can vary depending on the design criteria of a particular implementation.
[0073] The CTRL signal can be configured to provide a control signal. The CTRL signal can be generated in response to a decision made by processor 102. In an example, the CTRL signal can be generated in response to a detected object and / or features extracted from a video frame. The CTRL signal can be configured to enable, disable, or change the operating mode of another device. In one example, the CTRL signal can be used to lock / unlock a door controlled by an electronic lock. In another example, the device can be set to sleep mode (e.g., low power mode) and / or activated from sleep mode in response to the CTRL signal. In yet another example, the CTRL signal can be used to generate an alarm and / or notification. The type of device controlled by the CTRL signal and / or the response performed by the device in response to the CTRL signal can vary depending on the design criteria of the specific implementation.
[0074] The signal CTRL can be generated based on data received by sensor 164 (e.g., temperature readings, motion sensor readings, etc.). The signal CTRL can be generated based on input from HID 166. The signal CTRL can be generated based on human behavior detected by processor 102 in a video frame. The signal CTRL can be generated based on the type of detected object (e.g., human, animal, vehicle, etc.). The signal CTRL can be generated in response to the detection of a specific type of object at a specific location. The signal CTRL can be generated in response to the detection of a specific type of object at a specific location. The signal CTRL can be generated in response to user input to provide various parameters and / or settings to processor 102 and / or memory 150. Processor 102 can be configured to generate the signal CTRL in response to sensor fusion operations (e.g., aggregation of information received from different sources). Processor 102 can be configured to generate the signal CTRL in response to the result of a liveness detection performed by processor 102. The conditions used to generate the signal CTRL can vary depending on the design criteria of the specific implementation.
[0075] Signals DIMAGES may include one or more depth maps and / or disparity maps generated by processor 102. Signals DIMAGES may be generated in response to 3D reconstruction performed on a monocular single-channel image. Signals DIMAGES may be generated in response to analysis of captured video data and structured light pattern SLP.
[0076] A multi-step approach can be implemented to activate and / or disable the capture device 104 based on the output of the motion sensor 164 and / or any other power consumption characteristics of the camera system 100, thereby reducing the power consumption of the camera system 100 and extending the lifespan of the battery 152. The motion sensor in sensor 164 can have low power consumption on the battery 152 (e.g., less than 10W). In the example, the motion sensor of sensor 164 can be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 102. Video analysis performed by the processor / SoC 102 may have relatively high power consumption on the battery 152 (e.g., greater than that of the motion sensor 164). In the example, the processor / SoC 102 can be in a low-power state (or powered off) until some motion is detected by the motion sensor of sensor 164.
[0077] The camera system 100 can be configured to operate using various power states. For example, in a power-off state (e.g., sleep state, low power state), sensor 164 and the motion sensor of processor / SoC 102 can be turned on, and other components of the camera system 100 (e.g., image capture device 104, memory 150, communication module 154, etc.) can be turned off. In another example, the camera system 100 can operate in an intermediate state. In the intermediate state, image capture device 104 can be turned on, and memory 150 and / or communication module 154 can be turned off. In yet another example, the camera system 100 can operate in a powered-on (or high-power) state. In the powered-on state, sensor 164, processor / SoC 102, capture device 104, memory 150, and / or communication module 154 can be turned on. The camera system 100 can consume some power from battery 152 (e.g., a relatively small and / or minimal amount of power) in the power-off state. In the powered-on state, the camera system 100 can consume more power from battery 152. The number of power states and / or the number of components of the camera system 100 that are turned on when the camera system 100 operates in each power state can vary according to the design criteria of a particular implementation.
[0078] In some embodiments, the camera system 100 may be implemented as a system-on-a-chip (SoC). For example, the camera system 100 may be implemented as a printed circuit board including one or more components. The camera system 100 may be configured to perform intelligent video analysis on video frames of a video. The camera system 100 may be configured to crop and / or enhance the video.
[0079] In some embodiments, the video frame may be a view (or a derivative of a view) captured by the capture device 104. Pixel data signals may be enhanced by the processor 102 (e.g., color conversion, noise filtering, automatic exposure, automatic white balance, automatic focus, etc.). In some embodiments, the video frame may provide a series of cropped and / or enhanced video frames that improve the view from the perspective of the camera system 100 (e.g., providing night vision, providing high dynamic range (HDR) imaging, providing more viewing area, highlighting detected objects, providing additional data (e.g., digital distance to the detected object), etc.) to enable the processor 102 to see the location better than a human can see with human vision.
[0080] Encoded video frames can be processed locally. In one example, the encoded video can be stored locally by memory 150 so that processor 102 can facilitate computer vision analysis internally (e.g., without first uploading the video frames to a cloud service). Processor 102 can be configured to select video frames to be encapsulated into a video stream that can be transmitted over a network (e.g., a bandwidth-limited network).
[0081] In some embodiments, processor 102 may be configured to perform sensor fusion operations. The sensor fusion operations performed by processor 102 may be configured to analyze information from multiple sources (e.g., capture device 104, sensor 164, and HID 166). By analyzing various data from different sources, the sensor fusion operations may be able to make inferences about the data that might not be possible from just one data source. For example, the sensor fusion operations implemented by processor 102 may analyze video data (e.g., human mouth movements) and speech patterns from directional audio. Different sources can be used to develop scene models to support decision-making. For example, processor 102 may be configured to compare the synchronization of detected speech patterns with mouth movements in video frames to determine which person is speaking in the video frame. The sensor fusion operations may also provide temporal correlation, spatial correlation, and / or reliability of the received data.
[0082] In some embodiments, processor 102 may implement convolutional neural network (CNN) capabilities. CNN capabilities can be implemented using deep learning techniques for computer vision. CNN capabilities can be configured to perform pattern and / or image recognition using a training process involving multi-layer feature detection. Computer vision and / or CNN capabilities can be executed locally by processor 102. In some embodiments, processor 102 may receive training data and / or feature set information from external sources. For example, external devices (e.g., cloud services) may access various data sources to provide training data that the camera system 100 may not be able to obtain. However, computer vision operations performed using feature sets can be performed using the computational resources of processor 102 within the camera system 100.
[0083] The video pipeline of processor 102 can be configured to perform local dedistortion, cropping, enhancement, rolling shutter correction, stabilization, downsizing, packing, compression, conversion, mixing, synchronization, and / or other video operations. The video pipeline of processor 102 can enable multi-stream support (e.g., generating multiple bitstreams in parallel, each including a different bitrate). In the example, the video pipeline of processor 102 can implement an image signal processor (ISP) with an input pixel rate of 320 Mbps. The architecture of the video pipeline of processor 102 enables real-time and / or near real-time video operations on high-resolution video and / or high-bitrate video data. The video pipeline of processor 102 can implement computer vision processing, stereo vision processing, object detection, 3D noise reduction, fisheye lens correction (e.g., real-time 360-degree dedistortion and lens distortion correction), oversampling, and / or high dynamic range processing on 4K resolution video data. In one example, the video pipeline architecture can achieve 4K ultra-high resolution with H.264 encoding at dual real-time rates (e.g., 60fps) and 4K ultra-high resolution with H.265 / HEVC and / or 4KAVC encoding at 30fps (e.g., multi-stream 4KP30AVC and HEVC encoding). The type of video operation and / or the type of video data operated on by the processor 102 can vary depending on the design criteria of the specific implementation.
[0084] The camera sensor 180 can be a high-resolution sensor. Using the high-resolution sensor 180, the processor 102 can combine oversampling of the image sensor 180 with digital scaling within the cropped area. Each of oversampling and digital scaling can be one of the video operations performed by the processor 102. Oversampling and digital scaling can be implemented to provide a higher resolution image within the overall size constraints of the cropped area.
[0085] In some embodiments, lens 160 may be a fisheye lens. One of the video operations implemented by processor 102 may be a de-distortion operation. Processor 102 may be configured to de-distort generated video frames. De-distortion may be configured to reduce and / or remove severe distortion caused by fisheye lens and / or other lens characteristics. For example, de-distortion may reduce and / or eliminate bulging effects to provide a linear image.
[0086] Processor 102 can be configured to crop (e.g., trim) a region of interest from a full video frame (e.g., generate a region of interest video frame). Processor 102 can generate video frames and select regions. In the example, cropping the region of interest can generate a second image. The cropped image (e.g., the region of interest video frame) can be smaller than the original video frame (e.g., the cropped image can be a portion of the captured video).
[0087] The region of interest (ROI) can be dynamically adjusted based on the location of the audio source. For example, the detected audio source may be moving, and its location may shift as video frames are captured. Processor 102 can update the coordinates of the selected ROI and dynamically update the cropped portion (e.g., a directional microphone implemented as one or more sensors in sensor 164 can dynamically update its position based on captured directional audio). The cropped portion may correspond to the selected ROI. As the ROI changes, the cropped portion may change. For example, the selected coordinates of the ROI may change frame by frame, and processor 102 may be configured to crop the selected region in each frame.
[0088] Processor 102 can be configured to oversample image sensor 180. Oversampling of image sensor 180 can produce a higher resolution image. Processor 102 can also be configured to digitally magnify regions of video frames. For example, processor 102 can digitally magnify a cropped region of interest. For example, processor 102 can establish a region of interest based on directional audio, crop the region of interest, and then digitally magnify the cropped region of interest video frame.
[0089] The de-distortion operation performed by processor 102 can adjust the visual content of the video data. The adjustment performed by processor 102 can make the visual content look natural (e.g., look as if it were seen by a person viewing a position corresponding to the field of view of capture device 104). In the example, de-distortion can alter the video data to generate linear video frames (e.g., correcting artifacts caused by lens characteristics of lens 160). De-distortion operations can be implemented to correct distortions caused by lens 160. Adjusted visual content can be generated to achieve more accurate and / or reliable object detection.
[0090] Various features (e.g., de-distortion, digital scaling, cropping, etc.) can be implemented as hardware modules in processor 102. Implementing hardware modules can increase the video processing speed of processor 102 (e.g., faster than software implementation). Hardware implementation enables video to be processed while reducing latency. The hardware components used can vary depending on the design standards of the specific implementation.
[0091] Processor 102 is shown as including multiple blocks (or circuits) 190a-190n. Blocks 190a-190n can implement various hardware modules implemented by processor 102. Hardware modules 190a-190n can be configured to provide various hardware components to implement a video processing pipeline. Circuits 190a-190n can be configured to receive pixel data VIDEO, generate video frames from the pixel data, perform various operations on the video frames (e.g., de-distortion, rolling shutter correction, cropping, magnification, image stabilization, 3D reconstruction, liveness detection, etc.), prepare video frames for communication with external hardware (e.g., encoding, encapsulation, color correction, etc.), parse feature sets, and implement various operations for computer vision (e.g., object detection, segmentation, classification, etc.). Hardware modules 190a-190n can be configured to implement various security features (e.g., secure boot, I / O virtualization, etc.). Various implementations of processor 102 may not necessarily utilize all features of hardware modules 190a-190n. The features and / or functions of hardware modules 190a-190n may vary depending on the design criteria of a particular implementation. Details of hardware modules 190a-190n can be described in conjunction with U.S. Patent Application No. 16 / 831,549, filed April 16, 2020; U.S. Patent Application No. 16 / 288,922, filed February 28, 2019; U.S. Patent Application No. 15 / 593,493 (now U.S. Patent No. 10,437,600), filed May 12, 2017; U.S. Patent Application No. 15 / 931,942, filed May 14, 2020; U.S. Patent Application No. 16 / 991,344, filed August 12, 2020; and U.S. Patent Application No. 17 / 479,034, filed September 20, 2021, the appropriate portions of which are incorporated herein by reference in their entirety.
[0092] Hardware modules 190a-190n can be implemented as dedicated hardware modules. Compared to software implementation, using dedicated hardware modules 190a-190n to implement various functions of processor 102 allows processor 102 to be highly optimized and / or customized to limit power consumption, reduce heat generation, and / or increase processing speed. Hardware modules 190a-190n can be customizable and / or programmable to implement multiple types of operations. Implementing dedicated hardware modules 190a-190n allows the hardware used to perform each type of computation to be optimized for speed and / or efficiency. For example, hardware modules 190a-190n can implement several relatively simple operations frequently used in computer vision operations, which together enable computer vision operations to be performed in real time. The video pipeline can be configured to: identify objects. Objects can be identified by interpreting numerical and / or symbolic information to determine that visual data represents a specific type of object and / or feature. For example, the number of pixels and / or pixel color of video data can be used to identify portions of the video data as objects. Hardware modules 190a-190n enable computationally intensive operations (e.g., computer vision operations, video encoding, video transcoding, 3D reconstruction, depth map generation, liveness detection, etc.) to be performed locally by the camera system 100.
[0093] One of the hardware modules 190a-190n (e.g., 190a) can implement a scheduler circuit. Scheduler circuit 190a can be configured to store a directed acyclic graph (DAG). In the example, scheduler circuit 190a can be configured to generate and store a DAG in response to received (e.g., loaded) feature set information. The DAG can define video operations to be performed to extract data from video frames. For example, the DAG can define various mathematical weights (e.g., neural network weights and / or biases) to be applied when performing computer vision operations to classify various groups of pixels into specific objects.
[0094] Scheduler circuit 190a can be configured to parse acyclic graphs to generate various operators. Operators can be scheduled by scheduler circuit 190a in one or more hardware modules 190a-190n. For example, one or more hardware modules 190a-190n can implement a hardware engine configured to perform a specific task (e.g., a hardware engine designed to perform repetitive specific mathematical operations used for performing computer vision operations). Scheduler circuit 190a can schedule operators based on when they are ready to be processed by hardware engines 190a-190n.
[0095] Scheduler circuit 190a can multiplex task time across hardware modules 190a-190n based on their availability. Scheduler circuit 190a can parse a directed acyclic graph (DAG) into one or more data streams. Each data stream can include one or more operators. Once the DAG is parsed, scheduler circuit 190a can assign data streams / operators to hardware engines 190a-190n and send relevant operator configuration information to initiate the operators.
[0096] Each directed acyclic sphere Figure 2 The radix representation can be an ordered traversal of a directed acyclic graph, where descriptors and operators are intertwined based on data dependencies. Descriptors typically provide registers that link data buffers to specific operands in the associated operators. In various embodiments, operators may not appear in the directed acyclic graph representation until all associated descriptors have been declared for operands.
[0097] One of the hardware modules 190a-190n (e.g., 190b) can implement an Artificial Neural Network (ANN) module. The ANN module can be implemented as a fully connected neural network or a Convolutional Neural Network (CNN). In this example, the fully connected network is "structure-agnostic" because no special assumptions need to be made about the input. A fully connected neural network consists of a series of fully connected layers that connect each neuron in one layer to each neuron in another layer. In a fully connected layer, there are n*m weights for n inputs and m outputs. Each output node also has a bias value, resulting in a total of (n+1)*m parameters. In a trained neural network, (n+1)*m parameters have been determined during the training process. A trained neural network typically includes an architectural specification and a set of parameters (weights and biases) determined during the training process. In another example, a CNN architecture may explicitly assume that the input is an image in order to be able to encode specific attributes into the model architecture. A CNN architecture may include a sequence of layers, each layer transforming one activation into another through a differentiable function.
[0098] In the example shown, the artificial neural network 190b can implement a convolutional neural network (CNN) module. The CNN module 190b can be configured to perform computer vision operations on video frames. The CNN module 190b can be configured to recognize objects through multi-layer feature detection. The CNN module 190b can be configured to compute descriptors based on the performed feature detections. The descriptors enable the processor 102 to determine the probability that a pixel in a video frame corresponds to a specific object (e.g., a specific brand / model / year of a vehicle, identifying a person as a specific individual, detecting animal types, detecting facial features, etc.).
[0099] CNN module 190b can be configured to: implement convolutional neural network capabilities; implement computer vision using deep learning techniques; implement pattern and / or image recognition using a training process involving multi-layer feature detection; and perform inference on machine learning models.
[0100] CNN module 190b can be configured to perform feature extraction and / or matching solely in hardware. Feature points typically represent regions of interest (e.g., corners, edges, etc.) within a video frame. By tracking feature points over time, estimates of the self-motion of the capture platform or motion models of observed objects in the scene can be generated. To track feature points, the matching operation is typically incorporated into CNN module 190b by hardware to find the most probable correspondence between feature points in a reference and target video frame. During the matching of reference and target feature point pairs, each feature point can be represented by a descriptor (e.g., image patch, SIFT, BRIEF, ORB, FREAK, etc.). Implementing CNN module 190b using dedicated hardware circuitry enables real-time computation of descriptor matching distances.
[0101] The CNN module 190b can be configured to perform face detection, face recognition, and / or liveness detection. For example, face detection, face recognition, and / or liveness detection can be performed based on a trained neural network implemented by the CNN module 190b. In some embodiments, the CNN module 190b can be configured to generate a depth map from a structured light pattern. The CNN module 190b can be configured to perform various detection and / or recognition operations and / or perform 3D recognition operations.
[0102] CNN module 190b can be a dedicated hardware module configured to perform feature detection of video frames. Features detected by CNN module 190b can be used to compute descriptors. CNN module 190b can determine the probability that a pixel in a video frame belongs to a specific object and / or some objects in response to the descriptors. For example, using the descriptors, CNN module 190b can determine the probability that a pixel corresponds to a specific object (e.g., a person, a piece of furniture, a pet, a vehicle, etc.) and / or characteristics of the object (e.g., the shape of the eyes, the distance between facial features, a vehicle hood, body parts, a vehicle license plate, a face, clothing worn by a person, etc.). Implementing CNN module 190b as a dedicated hardware module of processor 102 enables device 100 to perform computer vision operations locally (e.g., on-chip) without relying on the processing power of a remote device (e.g., transmitting data to a cloud computing service).
[0103] The computer vision operations performed by the CNN module 190b can be configured to: perform feature detection on video frames to generate descriptors. The CNN module 190b can perform object detection to determine regions in the video frames with a high probability of matching a specific object. In one example, the type of object to be matched (e.g., a reference object) can be customized using an open operand stack (enabling the programmability of the processor 102 to implement various artificial neural networks defined by directed acyclic graphs, each providing instructions for performing various types of object detection). The CNN module 190b can be configured to: perform local masking on regions with a high probability of matching a specific object to detect the object.
[0104] In some embodiments, the CNN module 190b can determine the location (e.g., 3D coordinates and / or position coordinates) of various features (e.g., characteristics) of the detected object. In one example, 3D coordinates can be used to determine the position of a person's arms, legs, chest, and / or eyes. A position coordinate on a first axis representing the vertical position of a body part in 3D space and another coordinate on a second axis representing the horizontal position of a body part in 3D space can be stored. In some embodiments, the distance from the lens 160 can represent a coordinate of the depth position of the body part in 3D space (e.g., a position coordinate on a third axis). Using the various body parts' positions in 3D space, the processor 102 can determine the detected person's body position and / or body characteristics.
[0105] The CNN module 190b can be pre-trained (e.g., configured to perform computer vision to detect objects based on training data received to train the CNN module 190b). For example, the results of the training data (e.g., a machine learning model) can be pre-programmed and / or loaded into processor 102. The CNN module 190b can perform inference on the machine learning model (e.g., to perform object detection). Training can include determining weight values for each layer of the neural network model. For example, weight values can be determined for each layer for feature extraction (e.g., convolutional layers) and / or for classification (e.g., fully connected layers). The weight values learned by the CNN module 190b can vary depending on the design criteria of a particular implementation.
[0106] The CNN module 190b can perform feature extraction and / or object detection by performing convolution operations. These convolution operations can be hardware-accelerated for fast (e.g., real-time) computation that can be performed with low power consumption. In some embodiments, the convolution operations performed by the CNN module 190b can be used to perform computer vision operations. In some embodiments, the convolution operations performed by the CNN module 190b can be used for any function (e.g., 3D reconstruction) that may involve computing convolution operations and is performed by the processor 102.
[0107] Convolution operations can include sliding a feature detection window along a layer while performing computations (e.g., matrix operations). The feature detection window can apply filters to pixels and / or extract features associated with each layer. The feature detection window can be applied to a single pixel and multiple surrounding pixels. In the example, these layers can be represented as matrices representing the values of pixels and / or features of one of these layers, and the filters applied by the feature detection window can be represented as matrices. Convolution operations can apply matrix multiplication between regions of the current layer covered by the feature detection window. Convolution operations can slide the feature detection window along the region of the layer to generate a result representing each region. The size of the regions, the type of operation for applying filters, and / or the number of layers can vary depending on the design criteria of a particular implementation.
[0108] Using convolutional operations, the CNN module 190b can compute multiple features of pixels in the input image at each extraction step. For example, each layer can receive input from a set of features located in a small neighborhood (e.g., a region) of the previous layer (e.g., a local receptive field). Convolutional operations can extract basic visual features (e.g., oriented edges, endpoints, corners, etc.), which are then combined by higher layers. Because the feature extraction window operates on pixels and their nearby pixels (or subpixels), the results of the operations may be position-invariant. These layers can include convolutional layers, pooling layers, non-linear layers, and / or fully connected layers. In the example, convolutional operations can learn to detect edges from raw pixels (e.g., the first layer), then use features from the previous layer (e.g., detected edges) to detect shapes in the next layer, and then use these shapes to detect higher-level features in higher layers (e.g., facial features, pets, vehicles, vehicle parts, furniture, etc.), and the final layer may be a classifier using higher-level features.
[0109] The CNN module 190b can perform data stream operations for feature extraction and matching, including two-stage detection, transformation operators, component operators manipulating component lists (e.g., components can be regions of vectors sharing common attributes and can be combined with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, min-max-reduction operators, pooling operators, non-minimum and non-maximum suppression operators, non-maximum suppression operators based on scanning windows, aggregation operators, dispersion operators, statistical operators, classification operators, integral image operators, comparison operators, indexing operators, pattern matching operators, feature extraction operators, feature detection operators, two-stage object detection operators, score generation operators, block reduction operators, and upsampling operators. The types of operations performed by the CNN module 190b for extracting features from training data can vary depending on the design criteria of a particular implementation.
[0110] Each of the hardware modules 190a-190n can implement a processing resource (or hardware resource or hardware engine). Hardware engines 190a-190n can be used to perform specific processing tasks. In some configurations, hardware engines 190a-190n can operate in parallel and independently of each other. In other configurations, hardware engines 190a-190n can operate collaboratively with each other to perform assigned tasks. One or more hardware engines among the hardware engines 190a-190n can be homogeneous processing resources (all circuits 190a-190n can have the same capabilities) or heterogeneous processing resources (two or more circuits 190a-190n can have different capabilities).
[0111] refer to Figure 4 The diagram illustrates the processing circuitry of a camera system 100 configured to perform 3D reconstruction using a convolutional neural network. In this example, the processing circuitry of the camera system 100 can be configured for a variety of applications, including but not limited to: autonomous and semi-autonomous vehicles (e.g., cars, trucks, motorcycles, agricultural machinery, drones, aircraft, etc.), manufacturing, and / or security and surveillance systems. Compared to a general-purpose computer, the processing circuitry of the camera system 100 typically includes hardware circuitry optimized to provide high-performance image processing and computer vision pipelines with minimal area and power consumption. In this example, various operations for performing image processing, feature detection / extraction, 3D reconstruction, liveness detection, depth map generation, and / or object detection / classification for computer (or machine) vision can be implemented using hardware modules designed to reduce computational complexity and utilize resources efficiently.
[0112] In an example embodiment, processing circuitry 100 may include processor 102, memory 150, general-purpose processor 158, and / or memory bus 200. General-purpose processor 158 may implement a first processor. Processor 102 may implement a second processor. In this example, circuitry 102 may implement a computer vision processor. In this example, processor 102 may be an intelligent vision processor. Memory 150 may implement external memory (e.g., memory outside of circuitry 158 and 102). In this example, circuitry 150 may be implemented as dynamic random access memory (DRAM) circuitry. The processing circuitry of camera system 100 may include other components (not shown). The number, type, and / or arrangement of components in the processing circuitry of camera system 100 may vary depending on the design criteria of a particular implementation.
[0113] A general-purpose processor 158 can operate to interact with circuits 102 and 150 to perform various processing tasks. In an example, processor 158 can be configured as a controller for circuit 102. Processor 158 can be configured to execute computer-readable instructions. In one example, the computer-readable instructions may be stored by circuit 150. In some embodiments, the computer-readable instructions may include controller operations. Processor 158 can be configured to communicate with circuit 102 and / or access results produced by components of circuit 102. In an example, processor 158 can be configured to utilize circuit 102 to perform operations associated with one or more neural network models.
[0114] In the example, processor 102 typically includes scheduler circuitry 190a, block (or circuitry) 202, one or more blocks (or circuitries) 204a-204n, block (or circuitry) 206, and path 208. Block 202 may implement a directed acyclic graph (DAG) memory. DAG memory 202 may include CNN module 190b and / or weights / biases 210. Blocks 204a-204n may implement hardware resources (or engines). Block 206 may implement shared memory circuitry. In the example embodiment, one or more circuits in circuits 204a-204n may include blocks (or circuitries) 212a-212n. In the illustrated example, circuits 212a and 212b are implemented as representative examples in the corresponding hardware engines 204a-204b. One or more circuits in circuits 202, 204a-204n, and / or circuit 206 may be related to... Figure 3 Example implementations of hardware modules 190a-190n are shown in association.
[0115] In the example, processor 158 may be configured to program circuit 102 using one or more pre-trained artificial neural network models (ANNs), including a convolutional neural network (CNN) 190b having multiple output frames according to embodiments of the invention and weights / kernels (WGTS) 210 used by the CNN module 190b. In various embodiments, the CNN module 190b may be configured (trained) for operation in an edge device. In the example, the processing circuitry of camera system 100 may be coupled to a sensor (e.g., a video camera, etc.) configured to generate data input. The processing circuitry of camera system 100 may be configured to generate one or more outputs in response to data input from the sensor, based on one or more inferences made by the pre-trained CNN module 190b using weights / kernels (WGTS) 210. The operations performed by processor 158 may vary depending on the design criteria of a particular implementation.
[0116] In various embodiments, circuit 150 may implement dynamic random access memory (DRAM) circuitry. Circuit 150 is typically operable as a multidimensional array for storing input data elements and various forms of output data elements. Circuit 150 may exchange input data elements and output data elements with processor 158 and processor 102.
[0117] Processor 102 may implement computer vision processor circuitry. In examples, processor 102 may be configured to implement various functions for computer vision. Processor 102 is generally operable to perform specific processing tasks arranged by processor 158. In various embodiments, all or part of processor 102 may be implemented individually in hardware. Processor 102 may directly execute data streams involving the execution of CNN module 190b and generated by software (e.g., directed acyclic graphs, etc.) for specified processing tasks (e.g., computer vision, 3D reconstruction, liveness detection, etc.). In some embodiments, processor 102 may be a representative example of a multitude of computer vision processors implemented by the processing circuitry of camera system 100 and configured to operate together.
[0118] In one example, circuit 212a can implement a convolution operation. In another example, circuit 212b can be configured to provide a dot product operation. Convolution and dot product operations can be used to perform computer (or machine) vision tasks (e.g., as part of an object detection process). In yet another example, one or more of the circuits 204c-204n may include blocks (or circuits) 212c-212n (not shown) for providing multidimensional convolution computation. In yet another example, one or more of the circuits 204a-204n can be configured to perform a 3D reconstruction task.
[0119] In the example, circuit 102 can be configured to receive a directed acyclic graph (DAG) from processor 158. The DAG received from processor 158 can be stored in DAG memory 202. Circuit 102 can be configured to perform the DAG for CNN module 190b using circuits 190a, 204a-204n, and 206.
[0120] Multiple signals (e.g., OP_A-OP_N) can be exchanged between circuit 190a and corresponding circuits 204a-204n. Each signal in OP_A-OP_N can convey execution operation information and / or concession operation information. Multiple signals (e.g., MEM_A-MEM_N) can be exchanged between the respective circuits 204a-204n and circuit 206. Signals MEM_A-MEM_N can carry data. Signals (e.g., DRAM) can be exchanged between circuit 150 and circuit 206. Signal DRAM can transmit data between circuits 150 and 190a (e.g., on transmission path 208).
[0121] Scheduler circuit 190a is typically operable to schedule tasks within circuits 204a-204n to perform various computer vision-related tasks defined by processor 158. Individual tasks can be assigned by scheduler circuit 190a to circuits 204a-204n. Scheduler circuit 190a can assign individual tasks in response to parsing a directed acyclic graph (DAG) provided by processor 158. Scheduler circuit 190a can time-multiplex tasks across circuits 204a-204n based on their availability for performing work.
[0122] Each circuit 204a-204n can implement a processing resource (or hardware engine). Hardware engines 204a-204n are typically operable to perform a specific processing task. Hardware engines 204a-204n can be implemented as including dedicated hardware circuitry optimized for high performance and low power consumption while performing a specific processing task. In some configurations, hardware engines 204a-204n can operate in parallel and independently of each other. In other configurations, hardware engines 204a-204n can operate collaboratively with each other to perform assigned tasks.
[0123] Hardware engines 204a-204n can be homogeneous processing resources (e.g., all circuits 204a-204n can have the same capabilities) or heterogeneous processing resources (e.g., two or more circuits 204a-204n can have different capabilities). Hardware engines 204a-204n are typically configured to execute operators, which may include, but are not limited to, resampling operators, transformation operators, component operators manipulating a list of components (e.g., components can be vector regions sharing common attributes and can be grouped together with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, minimum-maximum-reduction operators, pooling operators, non-minimum, non-maximum suppression operators, aggregation operators, dispersion operators, statistical operators, classification operators, integral image operators, upsampling operators, and powers of two downsampling operators, etc.
[0124] In the example, hardware engines 204a-204n may include matrices stored in various memory buffers. The matrices stored in the memory buffers can be used to initialize convolution operators. The convolution operators can be configured to efficiently perform computations that are repeatedly executed on a convolution function. In the example, hardware engines 204a-204n implementing the convolution operators may include multiple mathematical circuits configured to process multi-bit input values and operate in parallel. Convolution operators can provide an efficient and general solution for computer vision and / or 3D reconstruction by using one-dimensional or higher-dimensional kernels to compute convolutions (also known as cross-correlation). Convolution can be used for computer vision operations such as object detection, object recognition, edge enhancement, image smoothing, etc. The techniques and / or architectures implemented by this invention are operable for computing convolutions of input arrays with kernels. Details of the convolution operators can be described in conjunction with U.S. Patent No. 10,310,768, filed January 11, 2017, the appropriate portions of which are incorporated herein by reference.
[0125] In various embodiments, hardware engines 204a-204n can be implemented as individual hardware circuits. In some embodiments, hardware engines 204a-204n can be implemented as general-purpose engines that can be configured to operate as dedicated machines (or engines) through circuit customization and / or software / firmware. In some embodiments, hardware engines 204a-204n can alternatively be implemented as one or more instances or threads of program code executing on processor 158 and / or one or more processors 102, including but not limited to vector processors, central processing units (CPUs), digital signal processors (DSPs), or graphics processing units (GPUs). In some embodiments, scheduler 190a can select one or more of hardware engines 204a-204n for a specific process and / or thread. Scheduler 190a can be configured to assign hardware engines 204a-204n to a specific task in response to parsing a directed acyclic graph stored in DAG memory 202.
[0126] Circuit 206 can implement shared memory circuitry. Shared memory 206 can be configured to store data in response to input requests and / or present data in response to output requests (e.g., requests from processor 158, DRAM 150, scheduler circuitry 190a, and / or hardware engines 204a-204n). In this example, shared memory circuitry 206 can implement on-chip memory for computer vision processor 102. Shared memory 206 is generally operable to store all or part of a multidimensional array (or vector) of input and output data elements generated and / or utilized by hardware engines 204a-204n. Input data elements can be transferred from DRAM circuitry 150 to shared memory 206 via memory bus 200. Output data elements can be sent from shared memory 206 to DRAM circuitry 150 via memory bus 200.
[0127] Path 208 can implement a transfer path within processor 102. Transfer path 208 is typically operable to move data from scheduler circuit 190a to shared memory 206. Transfer path 208 can also be operable to move data from shared memory 206 to scheduler circuit 190a.
[0128] Processor 158 is shown communicating with computer vision processor 102. Processor 158 may be configured as a controller of computer vision processor 102. In some embodiments, processor 158 may be configured to transmit instructions to scheduler 190a. For example, processor 158 may provide one or more directed acyclic graphs (DAGs) to scheduler 190a via DAG memory 202. Scheduler 190a may initialize and / or configure hardware engines 204a-204n in response to parsing the DAGs. In some embodiments, processor 158 may receive status information from scheduler 190a. For example, scheduler 190a may provide processor 158 with status information and / or readiness status from the outputs of hardware engines 204a-204n, enabling processor 158 to determine one or more next instructions to execute and / or decisions to make. In some embodiments, processor 158 may be configured to communicate with shared memory 206 (e.g., directly or via scheduler 190a, which receives data from shared memory 206 via path 208). Processor 158 can be configured to retrieve information from shared memory 206 to make a decision. The instructions executed by processor 158 in response to information from computer vision processor 102 can vary depending on the design criteria of a particular implementation.
[0129] refer to Figure 5A block diagram illustrating an intelligent automatic exposure control system is shown. An automatic exposure system 250 is shown. The automatic exposure system 250 may include a processor 102 and a capture device 104. A lens 160 is shown as receiving light input LIN. Typically, the intelligent automatic exposure control logic can be implemented by the processor 102.
[0130] Processor 102 may include blocks (or circuits) 260, 262, 264, and / or 266. Circuit 260 may implement an automatic exposure module (e.g., an AE module). Circuit 262 may implement sensor control logic. Circuit 264 may implement IR channel control logic. Circuit 266 may implement RGB channel control logic. Processor 102 may include other components (not shown). The number, type, and / or arrangement of components in processor 102 for performing intelligent automatic exposure control may vary depending on the design criteria of a particular implementation.
[0131] The AE module 260 can transmit signals (e.g., SENCTL) to sensor control logic 262. The AE module 260 can transmit signals (e.g., IRSLCTL) to IR channel control logic 264. The AE module 260 can transmit signals (e.g., RGBCTL) to RGB channel control logic 266. Other communication signals can be transmitted between the AE module 260, sensor control logic 262, IR channel control logic 264, and / or RGB channel control logic 266 (not shown). The data transmitted between the various components of processor 102 can vary depending on the design criteria of a particular implementation.
[0132] The AE module 260 can be configured to enable the RGB-IR sensor 180 to have consistent performance under various lighting conditions. The AE module 260 can provide adjustments to the sensor control logic 262, the IR channel control logic 264, and / or the RGB channel control logic 266 based on the visible and / or IR light in the environment. The adjustments provided by the AE module 260 can help overcome the physical limitations of the RGB-IR sensor 180.
[0133] The AE module 260 can be configured to provide and / or adjust individual control parameters for sensor settings of the RGB-IR sensor 180, the IR image channel, and / or the RGB image channel. The AE module 260 can be configured to determine and provide the parameters of the SENCTL signal to enable the sensor control logic 262 to adjust the control parameters for the RGB-IR sensor 180. The AE module 260 can be configured to determine and provide the parameters of the IRSLCTL signal to enable the IR channel control logic 264 to adjust the parameters of the IR image channel. The AE module 260 can be configured to determine and provide the parameters of the RGBCTL signal to enable the RGB channel control logic 266 to adjust the parameters of the RGB image channel. The individual adjustments determined by the AE module 260 enable the device 100 to dynamically respond to environmental conditions to ensure consistent image quality. For example, the AE module 260 can determine parameters and / or operating modes for selecting parameters for various control modules (e.g., sensor control logic 262, IR channel control logic 264, and RGB channel control logic 266) to enable tuning that can provide similar brightness for images generated under different visible and IR light conditions.
[0134] The AE module 260 can be configured to compensate for varying amounts of IR interference in the environment. The AE module 260 can be configured to select between two or more operating modes to overcome IR sensitivity to interference. The AE module 260 can select between a digital gain-priority operating mode and a shutter-priority operating mode. The AE module 260 can also receive user input (e.g., advance values) that can affect various operating modes.
[0135] Sensor control logic 262 can be configured to receive the signal SENCTL and generate a signal (e.g., SENPRM). The SENPRM signal can be transmitted to the capture device 104. Sensor control logic 262 can be configured to adjust the physical characteristics of the operation of the capture device 104. In this example, sensor control logic 262 can be configured to control the DC aperture of lens 160, the exposure time of the RGB-IR sensor 180, the shutter speed of lens 160, etc. Sensor control logic 262 can be configured to adjust at least three factors: aperture (or aperture diameter), exposure time, and / or analog gain control (AGC). AGC can be used to amplify the electrical signals of the RAW input data.
[0136] Since the response of the RGB-IR sensor 180 to the light input LIN can affect the physical characteristics of the input received by the device 100, adjustments made by the sensor control logic 262 can affect the input received by the RGB image channels and the IR image channels. In the example, a longer exposure time can allow for a larger exposure area of the structured light pattern SLP, which can provide a larger area of the depth map generated from the image in the IR image channel. However, a longer exposure time may also cause overexposure of the image in the RGB image channel. For example, the sensor control logic 262 can be shared by each image channel implemented by the processor 102. The sensor control logic 262 can be a source of adjustments that can be controlled by the AE module 260.
[0137] The IR channel control logic 264 can be configured to receive the signal IRSLCTL. The IR channel control logic 264 can also be configured to generate a signal (not shown) that can be configured to adjust characteristics of the image in the IR image channel with structured light. In the example, the characteristics adjusted by the IR channel control logic 264 could be post-processing adjustments, such as digital gain and tone curve adjustments. The signal IRSLCTL can be configured to dynamically and individually control pixel data in the IR image channel (with structured light).
[0138] The IR channel control logic 264 can be configured to provide adjustments to the generated IR image that has been captured by the structured light pattern SLP. For example, when the timing signal SL_TRIG toggles the structured light pattern SLP, the captured IR image may include a pattern of fill lights projected by the structured light pattern SLP. The IR channel control logic 264 can be used to adjust IR image channels with structured light. For example, since IR image channels may be sensitive to IR interference, the IR channel control logic 264 can provide individual adjustments to the IR image including the structured light pattern SLP in order to provide consistent brightness.
[0139] The RGB channel control logic 266 can be configured to receive the RGBCTL signal. The RGB channel control logic 266 can also be configured to generate a signal (not shown) that can be configured to adjust the characteristics of the image in the IR image channel (without structured light) and the RGB image channel. In the example, the characteristics adjusted by the RGB channel control logic 266 could be post-processing adjustments, such as digital gain and tone curve adjustments. The RGBCTL signal can be configured to dynamically and individually control the pixel data in the IR image channel (without structured light) and the RGB image channel.
[0140] The RGB channel control logic 266 can be configured to provide adjustments to the RGB image. The RGB channel control logic 266 can also be configured to provide adjustments to the generated IR image that does not have a structured light pattern SLP. For example, when the timing signal SL_TRIG turns off the structured light pattern SLP, the captured IR image may not include the dot pattern projected by the structured light pattern SLP. The RGB channel control logic 266 can be used to adjust IR image channels that do not have a structured light pattern. For example, since IR image channels without a structured light pattern are not affected by IR interference, the RGB channel control logic 266 can provide separate adjustments to the IR image and RGB image that do not include the structured light pattern SLP in order to provide consistent brightness.
[0141] Sensor control logic 262, IR channel control logic 264, and RGB channel control logic 266 enable adjustment of each image channel. For example, the IR channel (with structured light) may be highly sensitive to IR interference. A small amount of infrared light from an external light source (such as sunlight) can cause an overexposed IR image. In another example, if the shutter speed is reduced to prevent IR interference, the RGB image channel may be underexposed when the shutter speed is low. Even under different visible and IR light conditions, the individual channel control implemented by device 100 allows the RGB and IR images to be tuned to similar brightness.
[0142] AE module 260 may include block (or circuit) 270, block (or circuit) 272, and / or block (or circuit) 274. Circuit 270 may implement IR statistics. Circuit 272 may implement RGB statistics. Circuit 274 may implement an IR interference module. AE module 260 may include other components (not shown). The number, type, and / or arrangement of the components of AE module 260 may vary depending on the design criteria of a particular implementation.
[0143] IR statistics 270 can be configured to extract information about the IR image. In the example, AE module 260 can be configured to receive the IR image and analyze various parameters about it. IR statistics 270 can be used to determine various parameters for adjustment of IR channel control logic 264. AE module 260 can be configured to generate the signal IRSLCTL in response to the data extracted by IR statistics 270.
[0144] RGB statistics 272 can be configured to extract information about an RGB image. In the example, the AE module 260 can be configured to receive an RGB image (e.g., a visible light image) and analyze various parameters about the RGB image. RGB statistics 272 can be used to determine various parameters for adjustment of the RGB channel control logic 266. The AE module 260 can be configured to generate the signal RGBCTL in response to the data extracted by RGB statistics 272.
[0145] The IR interference module 274 can be configured to detect the amount of IR interference in the environment. The IR interference module 274 can also be configured to determine the operating mode in which the AE module 260 should operate based on the detected IR interference. In this example, the IR interference module 274 can be configured to switch between a shutter-priority operating mode and a digital gain-priority operating mode for the AE module 260. The operating mode selected by the IR interference module 274 can affect the parameters presented in the signals SENCTL, IRSLCTL, and / or RGBCTL. This can be associated with... Figure 7 To describe the details of the IR interference module 274.
[0146] The capture device 104 can be configured to present a signal VIDEO to the processor 102. The signal VIDEO may include pixel data captured by the RGB-IR sensor 180. In this example, the signal VIDEO may include an electrical signal of RAW input data from the RGB-IR sensor 180. The processor 102 can be configured to arrange the pixel data into video frames. The video frames may be analyzed by the AE module 260.
[0147] Processor 102 can be configured to present a signal SENPRM to capture device 104. The signal SENPRM can be generated by sensor control logic 262. Capture device 104 can be configured to adjust the physical operation of lens 160 and / or RGB-IR sensor 180 in response to the signal SENPRM.
[0148] Capture device 104 is shown as including block (or circuitry) 280. Circuitry 280 may implement actuators and / or control logic. Actuator and / or control logic 280 may be configured to adjust the physical operation of capture device 104. For example, actuator / control logic 280 may perform adjustments in response to the signal SENPRM. The operation and / or function of other components of capture device 104 (e.g., RGB-IR sensor 180, processor / logic 182, and / or frame buffer 184) may be adjusted in response to the signal SENPRM. In this example, actuator / control logic 280 may be configured to adjust the DC aperture, AGC, focus, zoom, tilt, and / or translation of capture device 104. For example, the amount of light input LIN reaching RGB-IR sensor 180 and / or the duration for which light input LIN is applied to RGB-IR sensor 180 may be adjusted in response to the signal SENPRM.
[0149] The physical constraint of the RGB-IR sensor 180 can be that the same RGB-IR sensor 180 simultaneously provides pixel data for both RGB and IR images. Since the pixel data for the RGB and IR images are captured simultaneously, the sensor control parameters generated by the sensor control logic 262 can affect both the pixel data from the RGB image and the pixel data from the IR image. Individually tuning the sensor control parameters may not overcome the physical limitations of the RGB-IR sensor 180 to balance the brightness of the RGB and IR image channels under various lighting conditions. The AE module 260 can be configured to individually tune the IR channel control logic 264 and / or the RGB channel control logic 266 to provide additional tuning that can overcome the physical constraints of the RGB-IR sensor 180. For example, the sensor control logic 262 can provide a layer of adjustment (e.g., to affect the light input LIN of the RGB-IR sensor 180), and the IR channel control logic 264 and the RGB channel control logic 266 can provide additional layer adjustments (e.g., post-processing, which can individually tune video frames in the RGB channel, IR channel (without structured light), and IR channel (with structured light)).
[0150] refer to Figure 6 , Figure 6A diagram illustrating a processor configured to extract IR video frames with structured light patterns, IR video frames without structured light patterns, and RGB video frames is shown. An example RGB / IR dispatcher 300 is shown. The example RGB / IR dispatcher 300 illustrates the extraction of RGB video frames and IR video frames from video frames generated by processor 102. The example RGB / IR dispatcher 300 may include processor 102, block (or circuit) 302, block (or circuit) 304, block (or circuit) 306, block (or circuit) 308, and / or block (or circuit) 310. Circuit 302 may implement an IR extraction module. Circuit 304 may implement an RGB extraction module. Block 306 may represent an IR structured light image channel (e.g., an IRSL image channel). Block 308 may represent an IR image channel without structured light (e.g., an IR image channel). Block 310 may represent an RGB image channel. The example RGB / IR dispatcher 300 may include other components (not shown). The number, type, and / or arrangement of the components in the RGB / IR distribution 300 can vary depending on the design criteria of the specific implementation.
[0151] Processor 102 may include video frames 320a-320n. Video frames 320a-320n may include pixel data received from a signal VIDEO presented by capture device 104. Processor 102 may be configured to generate video frames 320a-320n arranged from the pixel data. Although video frames 320a-320n are shown within processor 102, each of the IR extraction module 302, RGB extraction module 304, IRSL image channel 306, IR image channel 308, and / or RGB image channel 310 may be implemented within processor 102 and / or CNN module 190b. Circuits 302-310 may be implemented as a combination of discrete hardware modules and / or hardware engines 204a-204n, which are combined to perform a specific task. Circuits 302-310 may be conceptual blocks illustrating various techniques implemented by processor 102 and / or CNN module 190b.
[0152] Processor 102 can be configured to receive a signal VIDEO. The signal VIDEO may include RGB-IR pixel data generated by RGB-IR image sensor 180. The pixel data may include captured information about the environment and / or objects near device 104 and a structured light pattern SLP projected onto the environment and / or objects. Whether the pixel data in the signal VIDEO includes the structured light pattern SLP may depend on whether the structured light pattern SLP is currently active. The structured light pattern SLP may toggle between on and off in response to a timing signal SL_TRIG. Processor 102 may generate signals (e.g., FRAMES). The signal FRAMES may include video frames 320a-320n. Processor 102 can be configured to process pixel data of video frames 320a-320n arranged to include the structured light pattern SLP. In some embodiments, video frames 320a-320n may be presented to CNN module 190b (e.g., processed internally by processor 102 using CNN module 190b). Processor 102 can use video frames 320a-320n to perform other operations (e.g., generating encoded video frames for display, packaging video frames 320a-320n for transmission using communication module 154, etc.).
[0153] In the example shown, video frames 320a-320n can be generated by generating a structured light pattern SLP in one-third of the input frame using a structured light projector 106. For example, the timing signal SL_TRIG can be configured to: turn on the structured light pattern SLP for one of the video frames 320a-320n, then turn off the structured light pattern SLP for the next two video frames in 320a-320n, then turn on the structured light pattern SLP for the next video frame in 320a-320n, and so on. Details regarding the generation of the structured light pattern SLP using the timing signal SL_TRIG and / or the timing used for generating the structured light pattern SLP can be described in conjunction with U.S. Application No. 16 / 860,648, filed April 28, 2020, the appropriate portions of which are incorporated herein by reference. Point 322 is shown as an illustrative example of a structured light pattern SLP in video frames 320a and 320d (e.g., every three video frames). Although the dot pattern 322 is shown in the same pattern for illustrative purposes, the dot pattern can be different for each of the video frames 320a-320n to which the structured light pattern SLP is captured (e.g., depending on the object to which the structured light pattern SLP is projected when the image is captured). The dot pattern 322 can be captured in every three video frames 320a-320n by timing the projection of the structured light pattern SLP into one-third of the input frames. Although the timing of the one-third structured light pattern projection is shown, the apparatus 100 can be implemented with other rules to dispatch and / or extract IR and RGB images from the video frames 320a-320n.
[0154] Signal FRAMES, including video frames 320a-320n, can be presented to the IR extraction module 302 and the RGB extraction module 304. The IR extraction module 302 and the RGB extraction module 304 can be configured to extract appropriate content from the video frames 320a-320n for use by subsequent modules implemented by the processor 102. In the example shown, the IR extraction module 302 and the RGB extraction module 304 can extract appropriate content from the video frames 320a-320n for use in the IRSL image channel 306, the IR image channel 308, and the RGB image channel 310.
[0155] IR extraction module 302 can receive video frames 320a-320n. IR extraction module 302 can be configured to extract IR data (or IR video frames) from RGB-IR video frames 320a-320n. IR extraction module 302 can be configured to generate a signal (e.g., IRSL) and a signal (e.g., IR). Signal IRSL may include IR channel data extracted from video frames 320a-320n when the structured light pattern SLP is enabled. Signal IR may include IR channel data extracted from video frames 320a-320n when the structured light pattern SLP is disabled. Signal IRSL may include the structured light pattern SLP. Signal IR may not include the structured light pattern SLP. Both signal IRSL and signal IR may include a full-resolution IR image. Signal IRSL may be presented to IRSL image channel 306. Signal IRSL may transmit IRSL image channel 306 (e.g., a subset of the IR values of video frames 320a-320n with a valid structured light pattern SLP). The signal IR can be presented to the IR image channel 308. The signal IR can transmit the IR image channel 308 (e.g., an IR subset of video frames 320a-320n with an invalid structured light pattern SLP).
[0156] The RGB extraction module 304 can receive video frames 320a-320n. The RGB extraction module 304 can be configured to extract RGB data (or RGB video frames) from the RGB-IR video frames 320a-320n. The RGB extraction module 304 can be configured to generate a signal (e.g., RGB). The RGB signal may include RGB channel data extracted from the video frames 320a-320n. The RGB signal may include RGB information without a structured light pattern SLP (e.g., visible light data). The RGB signal can be presented to the RGB image channel 310. The RGB signal can transmit the RGB image channel 310 (e.g., a subset of the RGB values from the video frames 320a-320n).
[0157] IRSL image channel 306 may include IRSL images 330a-330k. IRSL images 330a-330k may each include a dot pattern 322. IRSL images 330a-330k may each include a full-resolution IR image with an effective structured light pattern SLP. IRSL image 330a may correspond to video frame 320a (e.g., having the same dot pattern 322), and IRSL image 330b may correspond to video frame 320d. Since the structured light pattern SLP is projected and captured for a subset of video frames 320a-320n, IRSL image channel 306 may include fewer IRSL images 330a-330k than the total number of generated video frames 320a-320n. In the example shown, IRSL image channel 306 may include one-third of video frames 320a-320n.
[0158] Structured light patterns (SLPs) can be displayed on IRSL channel data. IRSL data can be extracted from the output of RGB-IR sensor 180 by IR extraction module 302. IRSL images 330a-330k can be formatted as IR YUV images. IRSL images 330a-330k, including dot patterns 322 in IRYUV image format, can be presented to AE module 260 and / or other components of processor 102.
[0159] IR image channel 308 may include IR images 332a-332m. IR images 332a-332m may not include dot pattern 322. For example, structured light projector 106 may be timed to turn off during the capture of IR images 332a-332m. IR images 332a-332m may each include a full-resolution IR image with an invalid structured light pattern SLP. IR image 332a may correspond to video frame 320b (e.g., having the same IR pixel data), IR image 332b may correspond to video frame 320c, IR image 332c may correspond to video frame 320e, IR image 332d may correspond to video frame 320f, and so on. Since the structured light pattern SLP is projected and captured for a subset of video frames 320a-320n (e.g., IRSL images 330a-330k) and turned off for another subset of video frames 320a-320n (e.g., IR images 332a-332m), the IR image channel 308 can include fewer IR images 332a-332m than the total number of generated video frames 320a-320n but more than the total number of IRSL images 330a-330k.
[0160] The structured light pattern SLP can be disabled for IR channel data. IR data can be extracted from the output of the RGB-IR sensor 180 by the IR extraction module 302. IR images 332a-332m can be formatted as IR YUV images. IR images 332a-332m in IR YUV image format without the structured light pattern SLP can be presented to the AE module 260 and / or other components of the processor 102.
[0161] RGB image channel 310 may include RGB images 334a-334m. RGB images 334a-334m may not include dot pattern 322. For example, structured light projector 106 may be timed to turn off during the capture of RGB images 334a-334m. RGB images 334a-334m may each include a full-resolution RGB image with an invalid structured light pattern SLP. RGB image 334a may correspond to video frame 320b (e.g., having the same RGB pixel data), RGB image 334b may correspond to video frame 320c, RGB image 334c may correspond to video frame 320e, RGB image 334d may correspond to video frame 320f, and so on. Since the structured light pattern SLP is projected and captured for a subset of video frames 320a-320n (e.g., IRSL images 330a-330k) and turned off for another subset of video frames 320a-320n (e.g., RGB images 334a-334m), the RGB image channel 310 can include fewer RGB images 334a-334m than the total number of generated video frames 320a-320n but more than the total number of IRSL images 330a-330k.
[0162] RGB data can be extracted from the output of the RGB-IR sensor 180 by the RGB extraction module 304. RGB images 334a-334m can be formatted as RGB image formats. RGB images 334a-334m can be presented to the AE module 260 and / or other components of the processor 102.
[0163] IRSL image channel 306, IR image channel 308, and RGB image channel 310 can realize three different virtual channels from RAW images 320a-320n. Since the structured light pattern SLP is turned off more frequently than on, IRSL image channel 306 can include fewer IRSL images 330a-330k than IR images 332a-332m or RGB images 334a-334m. For example, IR image channel 308 and RGB image channel 310 can include the same frame number but different image content (e.g., IR light data and visible light data, respectively).
[0164] Data from IRSL image channel 306, IR image channel 308, and / or RGB image channel 310 can be used for various functions of device 100, processor 102, and / or CNN module 190b. In one example, various subsets of the data from image channels 306-310 can be used for 3D reconstruction and / or depth map generation. In another example, various subsets of the data from image channels 306-310 can be used to generate statistics for performing intelligent automatic exposure. In yet another example, various subsets of the data from image channels 306-310 can be used to generate output video (e.g., signal VIDOUT) for display. In yet another example, various subsets of the data from image channels 306-310 can be used to detect objects, classify objects, and / or analyze object characteristics. The use cases for the data provided by the various subsets of the data from image channels 306-310 can vary depending on the design criteria of a specific implementation.
[0165] refer to Figure 7 , Figure 7 A block diagram illustrating an IR interference module configured to determine the amount of IR interference in an environment is shown. An IR interference component 350 is shown. The IR interference component 350 may include various hardware, conceptual blocks, inputs, and / or outputs that can be used by the device 100 to perform IR interference measurements. The IR interference measurement can be used to select between a shutter-priority operating mode and a digital gain-priority operating mode. The IR interference component 350 can be represented as a block diagram illustrating the operation performed by the device 100 to select an operating mode in order to provide consistent brightness output for the individual channels 306-310.
[0166] The IR interference component 350 may include a processor 102, video frames 320a-320n, an IR extraction module 302, an RGB extraction module 304, and / or an IR interference module 274. For illustrative purposes, the processor 102, IR extraction module 302, RGB extraction module 304, and IR interference module 274 are shown as separate components. However, the IR extraction module 302, RGB extraction module 304, and / or IR interference module 274 may be components implemented by the processor 102. The IR interference module 274 may be a component of the AE module 260 (not shown).
[0167] Processor 102 can be configured to receive a signal VIDEO. The signal VIDEO may include RGB-IR pixel data generated by image sensor 180. Processor 102 can generate signal FRAMES that may include video frames 320a-320n. Processor 102 can be configured to process the pixel data, which is arranged into video frames 320a-320n that may include a structured light pattern SLP (depending on whether the timing signal SL_TRIG has activated the structured light projector 106). Video frames 320a-320n may be presented to (e.g., processed internally by processor 102 using IR extraction module 302 and RGB extraction module 304) IR extraction module 302 and RGB extraction module 304.
[0168] The IR interference module 274 can be configured to receive RGB signals (e.g., data from RGB image channel 310), IR signals (e.g., data from IR image channel 308), and / or signals (e.g., AGGR). The IR interference module 274 can also be configured to output signals (e.g., SF) and / or signals (e.g., DGF). While the IR extraction module 302 can generate data for the IRSL image channel 306, IR images with structured light patterns 330a-330k may not be used to measure the amount of IR interference in the environment. The IR interference module 274 can use visible light data (e.g., RGB images 334a-334m) and infrared light data without structured light patterns (e.g., IR images 332a-332m) to measure the amount of IR interference in the environment. The IR interference module 274 can transmit or receive other data (not shown). The number of inputs / outputs and / or the type of data transmitted by the IR interference module 274 can vary depending on the design criteria of the specific implementation.
[0169] The signal AGGR can be an incremental value. The signal AGGR can be a configurable value. In one example, the signal AGGR can be a user-input value. For example, the signal AGGR can be input to the IR interference module 274 using HID 166. In another example, the signal AGGR can be a pre-configured value (e.g., pre-stored in memory 150). The signal AGGR can provide a percentage value from 0% to 100%. The percentage value of the signal AGGR can provide a division or ratio setting for the shutter speed adjustment amount that can be applied to intelligent auto exposure using the capture device 104. For example, the signal AGGR can provide user input to provide a customized balance between the shutter speed control amount and the digital gain amount used for auto exposure adjustment. The signal AGGR can be applied using user input in digital gain-priority operation mode and / or shutter-priority operation mode.
[0170] Signals SF and DGF can be selected by the IR interference module 274 for the operating mode of the AE module 260. The IR interference module 274 can select signal SF for shutter priority operating mode or select signal DGF for digital gain priority operating mode. The selection of signal SF or signal DGF can depend on the amount of IR interference detected by the IR interference module 274.
[0171] In response to the signal DGF generated by the IR interference module 274, the AE module 260 can operate in a digital gain-priority operation mode. In this mode, the AE module 260 can generate signals SENCTL, IRSLCTL, and / or RGBCTL, which can provide parameters to keep the shutter speed of the capture device 104 at a low value (e.g., depending on the aggressive value in signal AGGR) and enable digital gain to prioritize increasing image brightness. In this mode, the IR image channel 308 may have reduced IR interference effects, while the IR fill light (e.g., structured light pattern SLP) may have high output power, but this high output power may still result in exposure in the IRSL image channel 306. In this mode, RGB images 334a-334m may appear dark, but the RGB channel control logic 266 can apply digital gain and / or tone curve adjustments to enhance the brightness of the RGB image channel 310. In the example, sunlight can provide strong infrared light, which may cause overexposure of IRSL images 330a-330k and / or IR images 332a-332m. A digital gain-priority operating mode can be selected in high IR interference environments.
[0172] Digital gain can have a maximum value. In the example, the maximum digital gain can be a user-defined limit (e.g., a limit based on the amount of digital gain that might cause video frames to look unnatural, video frames that do not provide sufficient depth data for liveness detection, and / or video frames that do not provide accurate color representation for computer vision operations). In digital gain-priority operating mode, digital gain can be increased until the maximum digital gain value is reached. After the maximum digital gain value is reached, the AE module 260 can adjust the SENCTL signal to allow the shutter speed to be increased. Digital gain-priority operating mode may preferably perform adjustments to increase digital gain and / or tone mapping first, and then increase the shutter speed (if this is beneficial based on IR statistics 270 and / or RGB statistics 272).
[0173] In response to the signal SF generated by the IR interference module 274, the AE module 260 can operate in shutter-priority mode. In shutter-priority mode, the AE module 260 can generate signals SENCTL, IRSLCTL, and / or RGBCTL, which can provide parameters to maintain the shutter time of the capture device 104 at a high value (e.g., depending on the aggressive value in signal AGGR) to increase the brightness of the generated image. Increasing the shutter time allows more rows of the RGB-IR sensor 190 to be exposed for IR illumination (e.g., structured light pattern SLP). Enabling a longer exposure time in shutter-priority mode can provide a larger exposure area for the structured light pattern SLP. With a larger exposure area in the IRSL image channel 306, a larger depth area can be generated (this can provide a larger area to cover the object to determine depth). However, increasing the exposure time may result in overexposure of the RGB images 334a-334m in the RGB image channel 310, which can be corrected by the RGB channel control logic 266. The shutter-priority mode can be selected in low IR interference environments.
[0174] Shutter speed can have a maximum value. In the example, the maximum shutter speed can be a physical limitation of the capture device 104 and / or the RGB-IR sensor 180. In shutter-priority operation mode, the shutter speed can be increased until the maximum shutter speed value is reached. After the maximum shutter speed value is reached, the AE module 260 can adjust the signals IRSLCTL and RGBCTL to enable IR channel control logic 264 and / or RGB channel control logic 266 to change the digital gain and tone mapping. The shutter-priority operation mode may preferably perform adjustments to increase the shutter speed first, and then perform digital adjustments (if this is beneficial based on IR statistics 270 and / or RGB statistics 272).
[0175] The IR interference module 274 may include blocks (or circuits) 352 and / or 354. Circuit 352 may implement an average intensity module. Circuit 354 may implement an automatic exposure decision module. The IR interference module 274 may include other components (not shown). Circuits 352-354 may be implemented as a combination of discrete hardware modules and / or hardware engines 204a-204n combined to perform a specific task. Circuits 352-354 may be conceptual blocks illustrating various techniques implemented by processor 102 and / or CNN module 190b. The number, type, and / or arrangement of the components of the IR interference module 274 may vary depending on the design criteria of a particular implementation.
[0176] The average intensity module 352 can be configured to receive signals RGB and IR. The average intensity module 352 can generate a signal (e.g., AVGINT) in response to the signals RGB and IR. The automatic exposure decision module 354 can receive signals AVGINT and AGGR. The automatic exposure decision module 354 can generate a signal SF or a signal DGF in response to signals AVGINT and / or AGGR.
[0177] The average intensity module 352 can be configured to calculate the average intensity (e.g., luminance) of two images. The average intensity can be determined based on the luminance components of the IR images 332a-332m in IR image channel 308 and the RGB images 334a-334m in RGB image channel 310. The signal AVGINT can output the average intensity measured by the average intensity module 352. The average intensity module 352 can be configured to compare one of the IR images (without structured light) 332a-332m with one of the RGB images 334a-334m. Since the IR images 332a-332m and the RGB images 334a-334m are captured simultaneously (e.g., when the structured light pattern SLP is turned off), the environmental conditions in the IR images 332a-332m and the RGB images 334a-334m can be the same. The average intensity module 352 can be configured to perform a division operation using the intensities of RGB images 334a-334m and the corresponding intensities of IR images 332a-332m to determine the average intensity. In the example, the average intensity can be determined by dividing the intensity of one of the IR images 332a-332m by the intensity of the corresponding one of the RGB images 334a-334m. The average intensity can be used to determine the amount of IR interference in the environment. The result of the division can be represented as the signal AVGINT.
[0178] The automatic exposure decision module 354 can be configured to analyze the average intensity calculated by the average intensity module 352. In an example, the automatic exposure decision module 354 can be configured to execute computer-readable instructions to select between a digital gain-priority operating mode and a signal-priority operating mode. In one example, the computer-readable instructions may be stored in memory 150. The automatic exposure decision module 354 can be configured to receive a signal AVGINT and compare the average intensity value with a predetermined threshold. In one example, if the average intensity value exceeds the predetermined threshold, the automatic exposure decision module 354 can determine that the IR interference is high. When the IR interference is determined to be high, the automatic exposure decision module 354 can select a digital gain-priority operating mode (e.g., a signal DGF may be generated). When the IR interference is determined to be low, the automatic exposure decision module 354 can select a shutter-priority operating mode (e.g., a signal SF may be generated).
[0179] A predetermined threshold can be determined based on the calibration of device 100. Various factors can influence the selection of the predetermined threshold. In one example, some factors that can be analyzed for calibrating the predetermined threshold of IR interference module 274 may be the energy intensity of structured light projector 106, the maximum duty cycle of structured light projector 106, and / or the maximum duration of SLP source 186. In another example, some factors that can be analyzed for calibrating the predetermined threshold may be the sensitivity of the RGB channel of RGB-IR sensor 180. In yet another example, some factors that can be analyzed for calibrating the predetermined threshold may be the sensitivity of the IR channel of RGB-IR sensor 180. When the average intensity of one of the IR images 332a-332m is higher than the average intensity of one of the RGB images 334a-334m (e.g., IR / RGB >= threshold), automatic exposure decision module 354 can switch AE module 260 to digital gain-priority operation mode. When the average intensity of one of the IR images 332a-332m is lower than the average intensity of one of the RGB images 334a-334m (e.g., IR / RGB < threshold), the automatic exposure decision module 354 can switch the AE module 260 to shutter priority operation mode. Various factors that can be used to calibrate the predetermined threshold may vary depending on the design criteria of the specific implementation.
[0180] As the IR interference measurement results change, the automatic exposure decision module 354 can switch between digital gain-priority operation mode and shutter-priority operation mode. By adjusting the operation mode in real time according to the changes in IR interference measurement results, the AE module 260 can be configured to adapt to changing lighting environments to ensure consistent brightness of the output images from each channel 306-310. The advance value provided by the AGGR signal can be used by the automatic exposure decision module 354 to determine the amount of parameter adjustment in each operation mode.
[0181] refer to Figure 8 , Figure 8A block diagram illustrating the adjustment of IR channel control and RGB channel control in response to statistics measured from video frames is shown. A processor 102, comprising components and / or a portion of a module, is shown. An AE module 260, IR channel control logic 264, RGB channel control logic 266, IRSL image channel 306, IR image channel 308, and RGB image channel 310 are shown. The AE module 260 can be configured to receive feedback data from the respective channels 306-310. The feedback provided by the various channels 306-310 can be used by the AE module 260 to provide adjustments to the sensor control logic 262, IR channel control logic 264, and / or RGB channel control logic 266. Adjustments to the sensor control logic 262, IR channel control logic 264, and / or RGB channel control logic 266 can affect video frames 320a-320n assigned in the respective channels 306-310.
[0182] Processor 102 can be configured to execute computer-readable instructions. In some embodiments, the computer-readable instructions executed by processor 102 can be stored in memory 150. In one example, the computer-readable instructions executed by processor 102 can process pixel data in a signal VIDEO arranged as video frames and extract an IR image with structured light patterns 330a-330k, an IR image without structured light patterns 332a-332m, and an RGB image 334a-334m. In another example, the computer-readable instructions executed by processor 102 can generate a sensor control signal SENCTL in response to the IR images 330a-330k and the RGB images 334a-334m. In yet another example, the computer-readable instructions executed by processor 102 can calculate an IR interference measurement in response to the IR images 332a-332m and the RGB images 334a-334m, and select the operating mode (as associated with) the IR channel control logic 264 and the RGB channel control logic 266 in response to the IR interference measurement. Figure 7 (As shown). The functionality of the processor 102, which responds to the execution of computer-readable instructions, can vary depending on the design criteria of a particular implementation.
[0183] Video frames 320a-320n generated by processor 102 can be extracted by IR extraction module 302 and / or RGB extraction module 304 to provide IRSL video frames 330a-330k with dot pattern 322 to IRSL image channel 306, IR video frames 332a-332m without dot pattern 322 to IR image channel 308, and RGB video frames 334a-334m to RGB image channel 310. IRSL image channel 306 can present a signal (e.g., IRSTAT) to AE module 260. The IRSTAT signal can include statistics about the IRSL images 330a-330k generated by IRSL image channel 306. RGB image channel 310 can present a signal (e.g., IRSTAT) to AE module 260. The RGBSTAT signal can include statistics about the RGB images 334a-334m generated by RGB image channel 310. IR image channel 308 can similarly generate statistics about IR images 332a-332m, but the statistics may not be used for the purposes of AE module 260.
[0184] IR statistics 270 can be configured to extract information about IRSL images 330a-330k in response to the signal IRSTAT. IR statistics 270 can be used by AE module 260 to determine various parameters for adjusting sensor control logic 262, IR channel control logic 264, and / or RGB channel control logic 266. AE module 260 can be configured to generate the signal IRSLCTL in response to data extracted by IR statistics 270 and / or the operating mode selected by IR interference module 274.
[0185] RGB statistics 272 can be configured to extract information about RGB images 334a-334m in response to the signal RGBSTAT. AE module 260 can use RGB statistics 272 to determine various parameters for adjusting sensor control logic 262, IR channel control logic 264, and / or RGB channel control logic 266. AE module 260 can be configured to generate the signal RGBCTL in response to data extracted by RGB statistics 272 and / or the operating mode selected by IR interference module 274.
[0186] IR statistics 270 and RGB statistics 272 may include analysis of video frames 320a-320n generated under various lighting conditions. In one example, the statistics used and / or extracted by IR statistics 270 and / or RGB statistics 272 may include the sum of the luminance values of all pixels in a slice of video frames 320a-320n. Each video frame 320a-320n (e.g., IRSL video frames 330a-330k and RGB video frames 334a-334m) may be divided into blocks with M columns and N rows, each block providing luminance values. In another example, the statistics used and / or extracted by IR statistics 270 and / or RGB statistics 272 may include 64-bit histogram data for each channel of the RGB-IR sensor 180 (e.g., R channel, G channel, B channel, and luminance channel). Each of the IRSL image channel 306, IR image channel 308, and RGB image channel 310 may have separate statistics. The statistics for IRSL image channel 306, IR image channel 308, and RGB image channel 310 may include multiple data entries, such as summation of luminance values, histograms, etc. Each of channels 306-310 may use some (or all) statistics generated based on a specific use case and / or function. For intelligent auto exposure implemented by AE module 260, IR statistics 270 and RGB statistics 272 may be used. The type of statistics used by AE module 260 for video frames 320a-320n may vary depending on the design criteria of the specific implementation.
[0187] In response to the operating mode selected by IR statistics 270, RGB statistics 272, and / or the IR interference module 274, the AE module 260 can determine various sensor parameters and present the SENCTL signal to the sensor control logic 262 (not shown). In response to the operating mode selected by IR statistics 270, RGB statistics 272, and / or the IR interference module 274, the AE module 260 can determine various IR parameters and present the IRSLCTL signal to the IR channel control logic 264. In response to the operating mode selected by IR statistics 270, RGB statistics 272, and / or the IR interference module 274, the AE module 260 can determine various RGB parameters and present the RGBCTL signal to the RGB channel control logic 266. The selected parameters enable the device 100 to respond in real time to changing lighting conditions, thereby providing intelligent automatic exposure control for the RGB-IR sensor 180.
[0188] IR channel control logic 264 can be configured to generate a signal (e.g., DGAIN-IR) in response to the signal IRSLCTL. The signal DGAIN-IR may include parameters that can be used by IR channel control logic 264 to provide post-processing adjustments. The signal DGAIN-IR can be provided to IRSL image channel 306. The signal DGAIN-IR can be configured to adjust the brightness and / or tone curves of IRSL image frames 330a-330k in IRSL image channel 306. IRSL image channel 306 can continuously provide statistics about IRSL image frames 330a-330k to AE module 260 to enable real-time adjustment 262 in response to feedback regarding adjustments made by IR channel control logic 264 (and sensor control logic).
[0189] The RGB channel control logic 266 can be configured to generate a signal (e.g., DGAIN-RGB) in response to the signal RGBCTL. The signal DGAIN-RGB may include parameters that can be used by the RGB channel control logic 266 to provide post-processing adjustments. The signal DGAIN-RGB can be provided to both the IR image channel 308 and the RGB image channel 310. Since both the IR image channel 308 and the RGB image channel 310 include pixel data from video frames 320a-320n generated when the structured light pattern SLP is turned off, similar adjustments can be applied to both the IR image channel 308 and the RGB image channel 310. The signal DGAIN-RGB can be configured to adjust the brightness and / or tone curves of the IR image frames 332a-332m in the IR image channel 308, and the brightness and / or tone curves of the RGB image frames 334a-334m in the RGB image channel 310. The RGB image channel 310 can continuously provide the AE module 260 with statistics about RGB image frames 334a-334m to enable real-time adjustment in response to feedback on adjustments made by the RGB channel control logic 266 (and sensor control logic 262).
[0190] The sensor control logic 262 can be adjusted in response to the sensor control signal SENCTL to automatically adjust the exposure of the RGB-IR sensor 180. This adjustment of the RGB-IR sensor 180's exposure can affect the performance (e.g., brightness) of each of the IRSL images 330a-330k, IR images 332a-332m, and RGB images 334a-334m. The IR channel control logic 264 can be adjusted in response to the IR control signal IRSLCTL to automatically adjust the exposure of the IRSL images 330a-330k. The RGB channel control logic 266 can be adjusted in response to the RGB control signal RGBCTL to automatically adjust the exposure of the IR images 332a-332m and the RGB images 334a-334m. This automatic adjustment performed by the AE module 260 enables the RGB-IR sensor 180 to provide consistent performance under various illumination conditions in both visible and infrared light (e.g., generating video frames that include both IR pixel data and RGB pixel data, which can be adjusted to have consistent brightness).
[0191] refer to Figure 9 , Figure 9 A diagram illustrating the exposure results of a video frame is shown. Exposure result 370 is shown. The exposure result may include an input video frame 320i, a smart auto exposure output 372, and a fixed auto exposure output 374. Input video frame 320i may be a representative example of input video frames 320a-320n. In response to implementing device 100 using auto exposure module 260, sensor control logic 262, IR channel control logic 264, and RGB channel control logic 266, smart auto exposure output 372 may provide an example output video frame (e.g., the signal VIDOUT). Fixed auto exposure output 374 may provide an example output video frame in response to performing auto exposure without implementing auto exposure module 260, sensor control logic 262, IR channel control logic 264, and RGB channel control logic 266.
[0192] Video frame 320i is shown as a representative example of pixel data generated by RGB-IR sensor 180 and arranged as a video frame. For illustrative purposes, video frame 320i is shown as a single video frame. However, video frames 320a-320n that include IR data with a structured light pattern SLP may have different frame numbers than video frames 320a-320n that include RGB data, as associated with... Figure 6 As described. Typically, the IR extraction module 302 and the RGB extraction module 304 can be used to assign multiple video frames 320a-320n to generate IRSL video frames 330a-330k and RGB video frames 334a-334m.
[0193] Video frame 320i may include an image of person 380. Video frame 320i may include data input from RGB-IR sensor 180. The face of person 380 is shown in video frame 320i. Dot pattern 322 is shown as captured in video frame 320i. When the pixel data of video frame 320i is captured by RGB-IR sensor 180, structured light pattern SLP can be projected onto the face of person 380. In one example, pixel data of the face of person 380 may be captured to provide input to device 100 for performing face detection-face recognition (FDFR). In another example, pixel data of the face of person 380 with dot pattern 322 may be captured to provide input to device 100 for generating depth information for liveness detection and / or 3D reconstruction. Various features and / or functions of device 100 that are performed in response to intelligent auto-exposure following video frames 320a-320n may vary depending on the design criteria of the specific implementation.
[0194] The intelligent auto-exposure output 372 may include an RGB image 334i and an IRSL image 330i. The intelligent auto-exposure of the RGB image 334i and the IRSL image 330i may be performed in part based on the operating mode selected for the sensor control logic 262 and the parameters of the SENCTL signal. The RGB image 334i may be generated in response to the RGB extraction module 304 extracting RGB pixel data from the video frame 320i (e.g., based on timing of a structured light pattern SLP), and the RGB image channel 310 performing post-processing adjustments in response to the DGAIN-RGB signal selected by the RGB channel control logic 266 in response to the operating mode selected for the AE module 260. The IRSL image 330i can be generated in response to the IR extraction module 302 extracting IRSL pixel data (e.g., timing based on the structured light pattern SLP, with a different frame number than the RGB image 334i) from the video frame 320i, and the IRSL image channel 306 performing post-processing adjustments in response to the signal DGAIN-IR selected by the IR channel control logic 264 in response to the operating mode selected for the AE module 260.
[0195] RGB image 334i may include a color image of person 380. RGB image 334i may include color 382. In the example shown, person 380's shirt is shown shaded to indicate color 382. Dot pattern 322 is not shown in RGB image 334i. IRSL image 330i may include an infrared image of person 380. Dot pattern 322 may be displayed in IRSL image 330i.
[0196] The fixed auto-exposure output 374 may include an RGB image 334i' and an IRSL image 330i'. Auto-exposure of the RGB image 334i' and IRSL image 330i' can be performed without the AE module 260, sensor control logic 262, IR channel control logic 264, or RGB channel control logic 266. The RGB image 334i' may include a color image of a person 380. The RGB image 334i' may include color 382'. In the example shown, the person 380's shirt is shown as shaded to indicate color 382'. The dot pattern 322 is not shown in the RGB image 334i'. The IRSL image 330i' may include an infrared image of the person 380'. The dot pattern 322' may be displayed in the IRSL image 330i'.
[0197] For the intelligent auto exposure output 372, the RGB images 334a-334m can be well exposed and the color 382 can be correct (e.g., accurately represented). Similarly, for the fixed auto exposure output 374, the RGB images 334a'-334m' can be well exposed and the color 382' can be correct. Both the RGB images 334a-334m of the intelligent auto exposure output 372 and the RGB images 334a'-334m' of the fixed auto exposure output 374 can be suitable inputs for face detection-face recognition and / or other object detection operations that can be performed by the processor 102 and / or the CNN module 190b. In the example shown, the color 382 of the RGB image 334i may be slightly darker than the color 382' of the RGB image 334i' (or have more noise due to analog gain control (AGC)). However, the slight difference between color 382 of RGB image 334i and color 382' of RGB image 334i' may not be sufficient to affect FDFR and / or other computer vision operations performed by processor 102.
[0198] IRSL image 330i can be a monochrome image with dot pattern 322. IRSL image 330i' can be a monochrome image with dot pattern 322'. IRSL image 330i can be well exposed and all structured light patterns (SLP) can be captured in dot pattern 322. Through intelligent AE control implemented by device 100, IRSL images 330a-330k can be well exposed to obtain depth information and can be used by processor 102 for liveness detection.
[0199] IRSL image 330i' is shown as having overexposed areas 390a-390n. Overexposed areas 390a-390n may lack detail (e.g., facial portions of person 380' may be missing) and portions of dot pattern 322' may be missing (e.g., structured light pattern SLP may not be fully captured). Due to the overexposed areas 390a-390n, IRSL images 330a'-330k' with fixed auto exposure output 374 may not provide processor 102 with sufficient information to generate depth information and / or perform liveness detection.
[0200] In some embodiments, the fixed auto exposure output 374 may be adjusted to prevent overexposed regions 390a-390n in the IRSL images 330a'-330k'. However, when using fixed auto exposure, correcting for overexposed regions 390a-390n may sacrifice the brightness of the RGB images 334a'-334m', which could affect FDFR and / or other computer vision operations. The fixed auto exposure output 374 may include a trade-off between an accurately generated RGB image and an accurately generated IRSL image.
[0201] The intelligent automatic exposure output 372 enables the simultaneous generation of well-exposed RGB images 334a-334m and well-exposed IRSL images 330a-330k (or nearly simultaneous generation based on specific frame numbers assigned to each image channel 306-310). The AE module 260, sensor control logic 262, IR channel control logic 264, and RGB channel control logic 266, operating based on the selected operating mode, allow both RGB images 334a-334m and IRSL images 330a-330k to be generated without compromising accuracy and / or brightness between them. Using the intelligent automatic exposure of the RGB-IR sensor 180, the device 100 can provide outputs suitable for both FDFR and / or other computer vision operations, as well as for generating depth information and / or performing liveness detection.
[0202] refer to Figure 10 The diagram illustrates method (or process) 400. Method 400 can perform intelligent automatic exposure control on an RGB-IR sensor. Method 400 generally includes step (or state) 402, decision step (or state) 404, step (or state) 406, step (or state) 408, step (or state) 410, step (or state) 412, step (or state) 414, step (or state) 416, step (or state) 418, step (or state) 420, and step (or state) 422.
[0203] Step 402 can initiate method 400. Next, method 400 can move to decision step 404. In decision step 404, processor 102 can determine whether the structured light projector 106 is turned on. In this example, processor 102 can turn the structured light pattern SLP on or off based on a timing signal SL_TRIG. If the structured light pattern SLP is turned on, method 400 can move to step 406. In step 406, RGB-IR sensor 180 can generate pixel data with the exposed structured light pattern SLP, and this pixel data with the structured light pattern SLP can be received by processor 102. Next, method 400 can move to step 410. If, in decision step 404, the structured light pattern SLP is turned off, method 400 can move to step 408. In step 408, RGB-IR sensor 180 can generate pixel data without the structured light pattern SLP, and this pixel data can be received by processor 102. Next, method 400 can move to step 410.
[0204] In step 410, processor 102 can process the pixel data arranged as video frames 320a-320n. Next, method 400 can perform extraction steps 412-416. Although extraction steps 412-416 are shown sequentially, they can be performed in parallel by processor 102 to extract the video frames into the respective channels 306-310. In step 412, IR extraction module 302 can extract IRSL video frames 330a-330k with dot patterns 322 into IRSL image channel 306. In step 414, IR extraction module 302 can extract IR video frames 332a-332m without dot patterns 322 into IR image channel 308. In step 416, RGB extraction module 304 can extract RGB video frames 334a-334m into RGB image channel 310. After performing various extractions, method 400 can move to step 418.
[0205] In step 418, the AE module 260 can generate a sensor control signal SENCTL in response to the RGB statistics 272 of the RGB images 334a-334m and the IR statistics 270 of the IRSL images 330a-330k. Next, in step 420, the IR interference module 274 can calculate IR interference measurements in response to the RGB images 334a-334m and the IR images 332a-332m. For example, the average intensity can be calculated by the average intensity module 352. In step 422, the IR interference module 274 can select an operating mode (e.g., shutter priority operating mode or digital gain priority operating mode) by generating a signal SF or a signal DGF. The operating mode can be selected for the sensor control logic 262, the IR channel control logic 264, and the RGB channel control logic 266 based on the performed IR measurements. Next, method 400 can return to decision step 404 (e.g., generate more video frames and perform intelligent auto exposure based on the incoming input pixel data).
[0206] refer to Figure 11 The diagram illustrates method (or process) 450. Method 450 can perform IR interference measurements. Method 450 typically includes steps (or states) 452, 454, 456, 458, a decision step (or state) 460, 410, 462, 464, and 466.
[0207] Step 452 can initiate method 450. In step 454, the IR interference module 274 can receive RGB images 334a-334m, which can be presented to the average intensity module 352. In step 456, the IR interference module 274 can receive IR images 332a-332m (without dot pattern 322), which can be presented to the average intensity module 352. Although steps 454 and 456 are shown sequentially, the IR interference module 274 can simultaneously receive pixel data from steps 454-456. Next, in step 458, the average intensity module 352 can divide the brightness of the IR images 332a-332m by the brightness of the RGB images 334a-334m to determine the average intensity. The average intensity can be presented to the automatic exposure decision module 354 via the signal AVGINT. Next, method 450 can move to decision step 460.
[0208] In decision step 460, the automatic exposure decision module 354 can determine whether the calculated average intensity is less than a predetermined threshold. In one example, the AE decision module 354 can compare the average intensity in the signal AVGINT with the predetermined threshold. If the average intensity is less than the predetermined threshold, method 450 can move to step 462. In step 462, the AE decision module 354 can select a shutter priority operation mode. For example, the AE decision module 354 can provide the signal SF. Next, method 450 can move to step 466. In decision step 460, if the average intensity is greater than or equal to the predetermined threshold, method 450 can move to step 464. In step 464, the AE decision module 354 can select a digital gain priority operation mode. For example, the AE decision module 354 can provide the signal DGF. Next, method 450 can move to step 466.
[0209] In step 466, the AE module 260 can provide automatic exposure adjustment for sensor control logic 262 (e.g., signal SENCTL), IR channel control logic 264 (e.g., signal IRSLCTL), and RGB channel control logic 266 (e.g., signal RGBCTL) based on the selected operating mode. Next, method 450 can return to step 454. Method 450 can continuously measure IR interference based on the average intensity of the RGB images 334a-334m and IR images 332a-332m to dynamically adjust the operating mode to suit the environment.
[0210] refer to Figure 12 The diagram illustrates method (or process) 500. Method 500 can implement automatic exposure control in shutter priority operation mode. Method 500 generally includes step (or state) 502, decision step (or state) 504, step (or state) 506, step (or state) 508, decision step (or state) 510, decision step (or state) 512, step (or state) 514, step (or state) 516, step (or state) 518, step (or state) 520, and step (or state) 522.
[0211] Step 502 can initiate method 500. Next, method 500 can move to decision step 504. In decision step 504, AE module 260 can determine whether AE decision module 354 has selected shutter priority operation mode. For example, shutter priority operation mode can be selected in response to signal SF. If AE module 260 is operating in shutter priority operation mode, method 500 can move to step 506. In step 506, AE module 260 can receive the advance value AGGR. Next, in step 508, AE module 260 can receive statistics. AE module 260 can receive statistics in parallel from IRSL image channel 306 and RGB image channel 310. AE module 260 can receive IR statistics 270 from IRSL image channel 306 via signal IRSTAT and RGB statistics 272 from RGB image channel 310 via signal RGBSTAT. Next, method 500 can move to decision step 510.
[0212] In decision step 510, AE module 260 can determine whether the output should be adjusted. In one example, the output can be adjusted to achieve consistent brightness. AE module 260 can analyze IR statistics 270, RGB statistics 272, and / or the advance value AGGR to determine whether the output should be adjusted. If adjusting the output does not benefit performance, method 500 can return to decision step 504. If adjusting the output would benefit performance, method 500 can move to decision step 512.
[0213] In decision step 512, AE module 260 can determine whether the shutter speed has reached its maximum value. In this example, memory 150 and / or feedback can track the current shutter speed and / or the maximum shutter speed value. If the shutter speed has not increased to its maximum value, method 500 can proceed to step 514. In step 514, AE module 260 can increase the shutter speed of the sensor control signal SENCTL to adjust the exposure of the video frames extracted in each of the IRSL image channel 306, IR image channel 308, and RGB image channel 310. Next, method 500 can return to decision step 504.
[0214] In decision step 512, if the shutter time has been increased to its maximum value, method 500 can proceed to step 516. In step 516, AE module 260 can generate the signal IRSLCTL for IR channel control logic 264 to perform post-processing on IRSL images 330a-330k in IRSL image channels 306. For example, IR channel control logic 264 can control the gain of IRSL images 330a-330k via the signal DGAIN-IR. Next, in step 518, AE module 260 can generate the signal RGBCTL for RGB channel control logic 266 to perform post-processing on IR images 332a-332m in IR image channels 308 and RGB images 334a-334m in RGB image channels 310. For example, RGB channel control logic 266 can control the gain of IR images 332a-332m and RGB images 334a-334m via the signal DGAIN-RGB. For example, once the shutter speed reaches its maximum value, digital gain and / or tone mapping can be performed for further adjustments. Although steps 516-518 are shown sequentially, the adjustments in steps 516-518 can be performed in parallel. Next, method 500 can return to decision step 504.
[0215] In decision step 504, if the AE module 260 is not operating in shutter priority mode, then method 500 can proceed to step 550. In step 550, the AE module 260 can operate in digital gain priority mode (which will be combined with...). Figure 13 (Description follows). Next, method 500 can proceed to step 522. Step 522 can end method 500.
[0216] refer to Figure 13 The diagram illustrates method (or process) 550. Method 550 can implement automatic exposure control in digital gain-priority operation mode. Method 550 generally includes step (or state) 552, decision step (or state) 554, step (or state) 556, step (or state) 558, decision step (or state) 560, decision step (or state) 562, step (or state) 564, step (or state) 566, step (or state) 568, step (or state) 570, step (or state) 572, step (or state) 574, and step (or state) 576.
[0217] Step 552 can initiate method 550. Next, method 550 can move to decision step 554. In decision step 554, AE module 260 can determine whether AE decision module 354 has selected the digital gain priority operating mode. For example, the digital gain priority operating mode can be selected in response to the signal DGF. If AE module 260 is operating in digital gain priority operating mode, method 550 can move to step 556. In step 556, AE module 260 can receive the gain value AGGR. Next, in step 558, AE module 260 can receive statistics. AE module 260 can receive statistics in parallel from IRSL image channel 306 and RGB image channel 310. AE module 260 can receive IR statistics 270 from IRSL image channel 306 via the signal IRSTAT, and RGB statistics 272 from RGB image channel 310 via the signal RGBSTAT. Next, method 550 can move to decision step 560.
[0218] In decision step 560, AE module 260 can determine whether the output should be adjusted. In one example, the output can be adjusted to achieve consistent brightness. AE module 260 can analyze IR statistics 270, RGB statistics 272, and / or the advance value AGGR to determine whether the output should be adjusted. If adjusting the output does not benefit performance, method 550 can return to decision step 554. If adjusting the output would benefit performance, method 550 can move to decision step 562.
[0219] In decision step 562, AE module 260 can determine whether the digital gain has reached its maximum value. In this example, memory 150 and / or feedback can track the current digital gain and / or the maximum digital gain value. If the digital gain has not increased to its maximum value, method 550 can proceed to step 564. In step 564, AE module 260 can select a low shutter speed for the sensor control signal SENCTL to adjust the exposure of the video frames extracted in each of the IRSL image channel 306, IR image channel 308, and RGB image channel 310. Next, in step 566, AE module 260 can increase the digital gain values of IR channel control logic 264 and RGB channel control logic 266. In step 568, AE module 260 can generate the signal IRSLCTL for IR channel control logic 264 to perform post-processing on the IRSL images 330a-330k in IRSL image channel 306. For example, IR channel control logic 264 can control the gain of IRSL images 330a-330k via the signal DGAIN-IR. In step 570, the AE module 260 can generate the signal RGBCTL for the RGB channel control logic 266 to perform post-processing on the IR images 332a-332m in the IR image channels 308 and the RGB images 334a-334m in the RGB image channels 310. For example, the RGB channel control logic 266 can control the gain of the IR images 332a-332m and the RGB images 334a-334m via the signal DGAIN-RGB. Next, method 550 can return to decision step 554.
[0220] In decision step 562, if the digital gain has been increased to its maximum value, method 550 can proceed to step 572. In step 572, AE module 260 can increase the shutter time of the sensor control signal SENCTL to adjust the exposure of the video frames extracted in each of the IRSL image channel 306, IR image channel 308, and RGB image channel 310. Next, method 550 can return to decision step 554.
[0221] In decision step 554, if the AE module 260 is not operating in digital gain priority mode, then method 550 can proceed to step 574. In step 574, the AE module 260 can operate in shutter priority mode (which will be combined with...). Figure 12 (Description follows). Next, method 550 can proceed to step 574. Step 574 can end method 550.
[0222] One or more of the following can be used to achieve the effect of Figures 1-13The functions performed by the diagram are as follows: general-purpose processors, digital computers, microprocessors, microcontrollers, RISC (Reduced Instruction Set Computer) processors, CISC (Complex Instruction Set Computer) processors, SIMD (Single Instruction Multiple Data) processors, signal processors, central processing units (CPUs), arithmetic logic units (ALUs), video digital signal processors (VDSPs), and / or similar computing machines are programmed according to the teachings of this manual, as will be apparent to those skilled in the art. Appropriate software, firmware, codes, routines, instructions, opcodes, microcode, and / or program modules can be readily prepared by a skilled programmer based on the teachings of this disclosure, as will also be apparent to those skilled in the art. The software is typically implemented by one or more processors and executed from one or more media.
[0223] The present invention can also be implemented by preparing an ASIC (Application-Specific Integrated Circuit), a platform ASIC, an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), a CPLD (Complex Programmable Logic Device), a gate sea, an RFIC (Radio Frequency Integrated Circuit), an ASSP (Application-Specific Standard Product), one or more monolithic integrated circuits, one or more chips or dies arranged as flip-chip modules and / or multi-chip modules, or by interconnecting appropriate networks of conventional component circuits as described herein, modifications of which will be obvious to those skilled in the art.
[0224] Therefore, the present invention may also include a computer product, which may be a storage medium or medium and / or a transmission medium or medium, including instructions that can be used to program the machine to execute one or more processes or methods according to the present invention. The machine executes the instructions contained in the computer product and the operation of surrounding circuitry, and can convert input data into one or more files on the storage medium and / or one or more output signals representing physical objects or substances, such as audio and / or visual descriptions. The storage medium may include, but is not limited to, any type of disk, including floppy disks, hard disks, magnetic disks, optical disks, CD-ROMs, DVDs, and magneto-optical disks, as well as circuitry and / or any type of medium suitable for storing electronic instructions, such as ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Electrically Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), UVPROM (Ultraviolet Erasable Programmable ROM), flash memory, magnetic cards, optical cards, etc.
[0225] The elements of this invention can form part or all of one or more devices, units, components, systems, machines, and / or apparatuses. These devices may include, but are not limited to, servers, workstations, storage array controllers, storage systems, personal computers, laptop computers, notebook computers, handheld computers, cloud servers, personal digital assistants, portable electronic devices, battery-powered devices, set-top boxes, encoders, decoders, transcoders, compressors, decompressors, preprocessors, post-processors, transmitters, receivers, transceivers, cryptographic circuits, cellular phones, digital cameras, positioning and / or navigation systems, medical devices, head-up displays, wireless devices, audio recording, audio storage and / or audio playback devices, video recording, video storage and / or video playback devices, gaming platforms, peripheral devices, and / or multi-chip modules. Those skilled in the art will understand that the elements of this invention can be implemented in other types of devices to meet specific application standards.
[0226] The terms “may” and “generally” are used herein in conjunction with “is” and verbs to convey the intention that the description is exemplary and considered broad enough to encompass both the specific examples presented in this disclosure and the alternative examples that may be derived based on this disclosure. The terms “may” and “generally” as used herein should not be construed as necessarily implying the desirability or possibility of omitting the corresponding element.
[0227] The designations “a” through “n” for various components, modules, and / or circuits, when used herein, disclose a single component, module, and / or circuit or multiple such components, modules, and / or circuits, wherein the designation “n” is used to indicate any particular integer. Each distinct component, module, and / or circuit having an instance (or event) designated as “a” through “n” may indicate that the distinct component, module, and / or circuit may have a matching number of instances or a different number of instances. An instance designated as “a” may represent the first of multiple instances, and an instance “n” may refer to the last of multiple instances without implying a specific number of instances.
[0228] Although the invention has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention.
Claims
1. An automatic exposure control device, comprising: A structured light projector configured to switch structured light patterns in response to a timing signal; An image sensor configured to generate pixel data; as well as The processor is configured to: (i) process pixel data arranged as video frames; (ii) extract (a) an IR image having the structured light pattern, (b) an IR image without the structured light pattern, and (c) an RGB image from the video frames; (iii) generate sensor control signals in response to the IR image having the structured light pattern and the RGB image; (iv) calculate IR interference measurements in response to the IR image without the structured light pattern and the RGB image; and (v) select an operating mode for IR channel control and RGB channel control in response to the IR interference measurements, wherein... (a) The sensor control signal is configured to adjust the exposure of the image sensor for both the IR image and the RGB image having the structured light pattern. (b) The IR channel control is configured to adjust the exposure of the IR image having the structured light pattern, and (c) The RGB channel control is configured to adjust the exposure of the RGB image and the IR image without the structured light pattern.
2. The automatic exposure control device according to claim 1, wherein, The operating modes include digital gain priority operating mode or shutter priority operating mode.
3. The automatic exposure control device according to claim 2, wherein, The digital gain priority operation mode selects a low shutter speed and increases the digital gain value based on the advance value.
4. The automatic exposure control device according to claim 3, wherein, In the digital gain priority operation mode, (i) the IR channel control is configured to: expose the structured light pattern of the IR image with the structured light pattern and adjust the digital gain of the structured light pattern, and the RGB channel control is configured to: adjust the digital gain and tone curve to increase brightness, and (ii) the low shutter speed is increased after the digital gain reaches its maximum value.
5. The automatic exposure control device according to claim 2, wherein, The shutter priority operation mode (i) selects a high shutter speed based on an aggressive value to increase brightness and expose the structured light pattern, and (ii) adjusts the digital gain and tone curve after the high shutter speed reaches its maximum value.
6. The automatic exposure control device according to claim 2, wherein, (i) The processor is configured to receive an advance value, and (ii) the advance value is a user input configured to select a ratio of shutter adjustment to digital gain value for the digital gain priority operation mode or the shutter priority operation mode.
7. The automatic exposure control device according to claim 1, wherein, The processor is configured to enable the image sensor to have consistent performance under various lighting conditions, including visible and infrared light.
8. The automatic exposure control device according to claim 1, wherein, The IR interference measurement includes (i) determining the intensity of one of the IR images that does not have the structured light pattern; and (ii) determining the intensity of one of the RGB images. (iii) Determine the average intensity by dividing the intensity of one of the IR images that does not have the structured light pattern by the intensity of one of the RGB images, and (iv) Compare the average intensity with a predetermined threshold.
9. The automatic exposure control device according to claim 8, wherein, The predetermined threshold is determined based on the following: (i) the energy intensity, maximum duty cycle, and maximum duration of the structured light projector; (ii) the sensitivity of the RGB channel of the image sensor; and (iii) the sensitivity of the IR channel of the image sensor.
10. The automatic exposure control device according to claim 8, wherein, The processor (i) selects a digital gain priority operation mode when the average intensity is greater than or equal to the predetermined threshold and (ii) selects a shutter priority operation mode when the average intensity is less than the predetermined threshold.
11. The automatic exposure control device according to claim 1, wherein, The device is configured to provide automatic exposure adjustment for the image sensor.
12. The automatic exposure control device according to claim 1, wherein, The image sensor is an RGB-IR sensor, which is configured to generate the pixel data simultaneously for an IR image having the structured light pattern, an IR image without the structured light pattern, and the RGB image.
13. The automatic exposure control device according to claim 1, wherein, The sensor control signal is configured to adjust the exposure of the image sensor by controlling the shutter value, thereby controlling the aperture used for the capture device and the exposure time and gain value used for the image sensor.
14. The automatic exposure control device according to claim 1, wherein, The processor is further configured to: (i) determine IR statistics in response to an IR image having the structured light pattern, and determine RGB statistics in response to the RGB image, and (ii) generate the sensor control signal in response to the IR statistics and the RGB statistics.
15. The automatic exposure control device according to claim 14, wherein, The IR statistics include the sum of the brightness values of all pixels in a slice of the IR image with the structured light pattern, and a histogram of each channel of the IR image with the structured light pattern.
16. The automatic exposure control device according to claim 14, wherein, The RGB statistics include the sum of the luminance values of all pixels in a slice of the RGB image and the histogram of each channel of the RGB image.
17. The automatic exposure control device according to claim 1, wherein, The sensor control signal adjusts the exposure of the image sensor by adjusting the amount of light arriving at the image sensor.
18. The automatic exposure control device according to claim 1, wherein, The exposure of the IR image with the structured light pattern, the exposure of the RGB image, and the exposure of the IR image without the structured light pattern include post-processing techniques performed by the processor.
19. The automatic exposure control device according to claim 1, wherein, The processor is configured to: provide automatic exposure adjustment to (i) adjust the exposure of the IR image having the structured light pattern to prevent overexposure that would lead to loss of detail and ensure sufficient depth information for liveness detection, and (ii) adjust the exposure of the RGB image to ensure the correct colors for face detection and face recognition FDFR.
Citation Information
Patent Citations
Convolution calculations in multiple dimensions
US10310768B1
Memory hierarchy to transfer vector data for operators of a directed acyclic graph
US10437600B1
Using camera data to manage a vehicle parked outside in cold climates
US11001231B1
940nm LED flash synchronization for DMS and OMS
US11140334B1
Generating training data for speed bump detection
US11586843B1