adding an adaptive offset term to the local adaptive binarization expression using convolution techniques
By using convolutional neural network technology on edge devices and adding adaptive offset terms to the local adaptive binarization expression, the problems of slow speed and low accuracy of traditional 3D reconstruction technology are solved, achieving faster and more accurate 3D reconstruction results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- AMBARELLA INT LP
- Filing Date
- 2021-08-16
- Publication Date
- 2026-04-28
AI Technical Summary
Traditional 3D reconstruction technology is slow and inaccurate in real-time applications. In particular, the back-end calculation method of monocular speckle structured light system is not fast and accurate enough, which affects the effect of 3D reconstruction.
By employing convolutional techniques to add adaptive offset terms to the local adaptive binarization expression, and using a convolutional neural network for processing, and accelerating the processing with dedicated hardware on edge devices, the generation of the local adaptive binarization expression and the removal of error points are realized, thereby improving the calculation speed and accuracy.
It improves the speed and accuracy of 3D reconstruction, reduces computation time, lowers the probability of incorrect matching, and enhances the real-time application capability of 3D reconstruction.
Smart Images

Figure CN115880349B_ABST
Abstract
Description
Technical Field
[0001] In general, the present invention relates to computer vision, and more specifically, to a method and / or apparatus for adding adaptive offset terms to a locally adaptive binarized expression using convolution techniques. Background Technology
[0002] Machine vision, optical technology, and artificial intelligence have developed rapidly. Three-dimensional (3D) reconstruction has become an important branch of machine vision. However, traditional 3D reconstruction techniques have problems in real-time applications. The speed of 3D reconstruction is not fast enough, and the accuracy of 3D reconstruction is not accurate enough.
[0003] A 3D reconstruction method is performed using a monocular speckle structured light system. The results of 3D reconstruction using monocular speckle structured light are affected by various factors, such as speckle projector power, temporal noise, spatial noise, and the reflectivity of the object being detected. Due to the insufficient speed and accuracy of 3D reconstruction, its application is generally limited to scenarios where high accuracy is not required, such as 3D face recognition and real-time face detection.
[0004] The performance of 3D reconstruction techniques using monocular speckle structured light systems is primarily limited by the matching speed and accuracy of the back-end computation methods. Research on the preprocessing of single-channel images obtained from the front-end speckle structured light for improving the accuracy and speed of back-end computation is incomplete. Traditional back-end computations for performing binarization operations are mainly based on simple global binarization, local binarization, and local adaptive binarization. Then, other methods are used to perform local or global binarization summation.
[0005] Convolutional techniques are needed to add adaptive offset terms to the local adaptive binarization expression. Summary of the Invention
[0006] This invention relates to an apparatus comprising an interface, a structured light projector, and a processor. The interface can receive pixel data. The structured light projector can generate a structured light pattern. The processor can process the pixel data arranged as video frames, perform operations using a convolutional neural network to determine a binarization result and an offset value, and generate a disparity map and a depth map in response to the video frame, the structured light pattern, the binarization result, the offset value, and the removal of error points. The convolutional neural network can perform partial block summation and averaging on the video frame to generate a convolution result, compare the convolution result with an ideal speckle value to determine the offset value, generate an adaptive result in response to performing a convolution operation to add the offset value to the video frame, compare the video frame with the adaptive result to generate the binarization result for the video frame, and remove the error points from the binarization result. Attached Figure Description
[0007] Embodiments of the invention will become apparent from the following detailed description, as well as the appended claims and drawings.
[0008] Figure 1 This is a diagram illustrating an example of an edge device that can utilize a processor configured to implement a convolutional neural network, according to an exemplary embodiment of the present invention.
[0009] Figure 2 This is a diagram illustrating an example camera that implements an example embodiment of the present invention.
[0010] Figure 3 This is a block diagram showing a camera system.
[0011] Figure 4 This is a diagram showing the processing circuitry of a camera system configured to perform 3D reconstruction using a convolutional neural network.
[0012] Figure 5 This is a diagram illustrating the preprocessing of video frames using partial block summation performed by a neural network implemented by a processor.
[0013] Figure 6 This is a graph illustrating the determination of offset values performed by a neural network implemented by a processor.
[0014] Figure 7 This diagram illustrates how a neural network implemented by a processor combines video frames with offset values to determine the adaptive offset result.
[0015] Figure 8 This is a graph showing the error points removed using a quadruple-domain method executed by a processor-implemented neural network.
[0016] Figure 9 This is a diagram showing an example speckle image.
[0017] Figure 10 This is a diagram showing the disparity map generated from a speckle image after binarization without adding adaptive offset values.
[0018] Figure 11 This is a diagram showing the depth map generated from a speckle image after binarization without adding adaptive offset values.
[0019] Figure 12 This is a graph showing the binarized result generated from the speckle image in response to adding adaptive offset values and removing error points.
[0020] Figure 13This is a diagram showing the disparity map generated from the speckle image after binarization in response to the addition of adaptive offset values and the removal of error points.
[0021] Figure 14 This is a diagram showing the depth map generated from the speckle image after binarization in response to the addition of adaptive offset values and the removal of error points.
[0022] Figure 15 This is a flowchart illustrating a method for preprocessing video frames by adding adaptive offset terms to a local adaptive binarization expression using convolution techniques.
[0023] Figure 16 This is a flowchart illustrating a method for performing partial block summation and averaging using a convolutional neural network.
[0024] Figure 17 This is a flowchart illustrating a method for determining offset values.
[0025] Figure 18 This is a flowchart illustrating a method for generating adaptive results by summing offset values and producing a binarized result.
[0026] Figure 19 This is a flowchart illustrating a method for removing error points to generate binary data. Detailed Implementation
[0027] Embodiments of the present invention include providing an adaptive offset term added to a locally adaptive binarized expression using convolution techniques, which can (i) be executed by a processor on an edge device, (ii) utilize a convolutional neural network implemented on the processor, (iii) reduce the generation of error points, (iv) separate speckle patterns from a background image, (v) perform area summation using convolution operations and add an offset operation to accelerate execution, (vi) eliminate error points after binarization, (vii) reduce the probability of false matches, (viii) reduce the proportional error parallax for pixels by adding an adaptive bias, (ix) improve Z accuracy, (x) replace the averaging operation with convolution, and / or (xi) be implemented as one or more integrated circuits.
[0028] Embodiments of the present invention can be implemented using a video processor. The video processor may include hardware dedicated to implementing convolutional neural networks. This dedicated hardware can be configured to provide accelerated processing for convolution operations. The hardware acceleration provided by the dedicated hardware enables edge devices to compute locally binarized expressions using convolution techniques to add adaptive offset terms. Without hardware acceleration, convolution techniques may not be feasible (e.g., performance may be too slow for real-time applications).
[0029] Dedicated hardware for implementing convolutional neural networks can be configured to accelerate the preprocessing of speckle structured light for monocular 3D reconstruction. Embodiments of the invention can be configured to generate locally adaptive binary expressions with adaptive offset terms. Compared to conventional methods for monocular structured light matching preprocessing, generating locally adaptive binary expressions with adaptive offset terms can reduce the number of computations performed and / or improve the accuracy of the performed backend matching.
[0030] Summation can be performed using preprocessing methods implemented for monocular structured light matching techniques. Adaptive offset terms can be added to provide accuracy and / or richness to the feature representation of the binary image after binarization. Local summation can be performed using ordinary convolution methods implemented in dedicated hardware for convolutional neural networks. Offset terms can be added to improve the speed of the summation and subsequent offset addition process. Preprocessing using convolution can be performed to provide beneficial (e.g., optimal) conditions for generating locally adaptive binarized expressions based on the addition of adaptive offset terms.
[0031] Determining an adaptive offset term enables the separation of the speckle pattern from the background image. The adaptive offset term limits the generation of error points in the four-field domain. The remaining error points after binarization can be eliminated by implementing a four-field domain method. The four-field domain method can be implemented to reduce the probability of mismatches in subsequent stages.
[0032] Using convolution operations to perform binarization operations can speed up computation (e.g., reduce processing time). Convolution operations can be fully utilized to perform area summation and / or offset addition operations.
[0033] refer to Figure 1 The diagram illustrates an example of an edge device that can utilize a convolutional neural network according to an exemplary embodiment of the present invention. A top view of region 50 is shown. In the example shown, region 50 may be an outdoor location. Streets, vehicles, and buildings are shown.
[0034] Devices 100a-100n are shown at different locations within region 50. Each of devices 100a-100n can independently implement an edge device. Edge devices 100a-100n may include smart IP cameras (e.g., camera systems). Edge devices 100a-100n may include low-power technologies (e.g., microprocessors running on sensors, cameras, or other battery-powered devices) in embedded platforms designed for deployment at the network edge, where power consumption is a critical issue. In the example, edge devices 100a-100n may include various traffic cameras and Intelligent Transportation System (ITS) solutions.
[0035] Edge devices 100a-100n can be implemented for a variety of applications. In the illustrated example, edge devices 100a-100n may include an Automatic License Plate Recognition (ANPR) camera 100a, a traffic camera 100b, a vehicle camera 100c, an access control camera 100d, an ATM camera 100e, a bullet camera 100f, a dome camera 100n, etc. In the example, edge devices 100a-100n can be implemented as a traffic camera and Intelligent Transportation System (ITS) solution designed to enhance road safety by leveraging a combination of people and vehicle detection, vehicle brand / model recognition, and Automatic License Plate Recognition (ANPR) capabilities.
[0036] In the illustrated example, region 50 can be an outdoor location. In some embodiments, edge devices 100a-100n can be implemented in various indoor locations. In the example, edge devices 100a-100n can incorporate convolutional neural networks for use in security (surveillance) applications and / or access control applications. In the example, edge devices 100a-100n implemented as security cameras and access control applications can include battery-powered cameras, doorbell cameras, outdoor cameras, indoor cameras, etc. According to embodiments of the invention, security camera and access control applications can achieve performance advantages from the application of convolutional neural networks. In the example, edge devices utilizing convolutional neural networks according to embodiments of the invention can acquire large amounts of image data and perform on-device inference to obtain useful information (e.g., multiple temporal instances of images performed by each network), thereby reducing bandwidth and / or power consumption. The design, type, and / or application performed by edge devices 100a-100n can vary depending on the design criteria of a particular implementation.
[0037] refer to Figure 2The diagram illustrates an example edge device camera that implements an exemplary embodiment of the present invention. Camera systems 100a-100n are shown. Each camera device 100a-100n can have a different style and / or use case. For example, camera 100a can be an action camera, camera 100b can be a ceiling-mounted security camera, camera 100n can be a webcam, etc. Other types of cameras can be implemented (e.g., home security cameras, battery-powered cameras, doorbell cameras, stereo cameras, etc.). The design / style of cameras 100a-100n can vary depending on the design criteria of a particular implementation.
[0038] Each of the camera systems 100a-100n may include block (or circuit) 102, block (or circuit) 104, and / or block (or circuit) 106. Circuit 102 may implement a processor. Circuit 104 may implement a capture device. Circuit 106 may implement a structured light projector. Camera systems 100a-100n may include other components (not shown). They can be used with... Figure 3 The details of the components of the cameras 100a-100n are described in relation to each other.
[0039] Processor 102 can be configured to implement an artificial neural network (ANN). In an example, the ANN may include a convolutional neural network (CNN). Processor 102 can be configured to implement a video encoder. Processor 102 can be configured to process pixel data arranged as video frames. Capture device 104 can be configured to capture pixel data that can be used by processor 102 to generate video frames. Structured light projector 106 can be configured to generate a structured light pattern (e.g., a speckle pattern). The structured light pattern can be projected onto a background (e.g., the environment). Capture device 104 can capture pixel data including a background image (e.g., the environment) with a speckle pattern.
[0040] Cameras 100a-100n can be edge devices. A processor 102 implemented by each of the cameras 100a-100n enables the cameras 100a-100n to perform various functions internally (e.g., at the local level). For example, processor 102 can be configured to perform object / event detection (e.g., computer vision operations), 3D reconstruction, video encoding, and / or video transcoding on the device. For example, processor 102 can even perform advanced processes such as computer vision and 3D reconstruction without uploading video data to a cloud service to offload computationally intensive functions (e.g., computer vision, video encoding, video transcoding, etc.).
[0041] In some embodiments, multiple camera systems may be implemented (e.g., camera systems 100a-100n may operate independently of each other). For example, each of cameras 100a-100n may individually analyze captured pixel data and perform event / object detection locally. In some embodiments, cameras 100a-100n may be configured as a camera network (e.g., security cameras that send video data to a central source such as network-attached storage and / or cloud services). The location and / or configuration of cameras 100a-100n may vary depending on the design criteria of the specific implementation.
[0042] The capture device 104 of each of the camera systems 100a-100n may include a single lens (e.g., a monocular camera). The processor 102 may be configured to accelerate the preprocessing of speckle structured light for monocular 3D reconstruction. Monocular 3D reconstruction can be performed to generate depth and / or parallax images without using a stereo camera.
[0043] refer to Figure 3 The diagram illustrates a block diagram of an example implementation of a camera system 100. In this example, the camera system 100 may include a processor / SoC 102, a capture device 104, and a structured light projector 106, as shown below. Figure 2As shown in the associated diagram, the camera system 100 may further include blocks (or circuits) 150, 152, 154, 156, 158, 160, 162, 164, and / or 166. Circuit 150 may implement a memory. Circuit 152 may implement a battery. Circuit 154 may implement a communication device. Circuit 156 may implement a wireless interface. Circuit 158 may implement a general-purpose processor. Block 160 may implement an optical lens. Block 162 may implement a structured light patterning lens. Circuit 164 may implement one or more sensors. Circuit 166 may implement a human-machine interface (HID) device. In some embodiments, the camera system 100 may include a processor / SoC 102, a capture device 104, an IR structured light projector 106, a memory 150, a lens 160, an IR structured light projector 106, a structured light pattern lens 162, a sensor 164, a battery 152, a communication module 154, a wireless interface 156, and a processor 158. In another example, the camera system 100 may include the processor / SoC 102, the capture device 104, the structured light projector 106, the processor 158, the lens 160, the structured light pattern lens 162, and the sensor 164 as a single device, and the memory 150, the battery 152, the communication module 154, and the wireless interface 156 may be components of separate devices. The camera system 100 may include other components (not shown). The number, type, and / or arrangement of the components of the camera system 100 may vary depending on the design criteria of a particular implementation.
[0044] Processor 102 can be implemented as a video processor. In an example, processor 102 can be configured to receive three-sensor video input using a high-speed SLVS / MIPI-CSI / LVCMOS interface. In some embodiments, processor 102 can be configured to perform depth sensing in addition to generating video frames. In an example, depth sensing can be performed in response to depth information and / or vector light data captured in the video frames.
[0045] Memory 150 can store data. Memory 150 can be implemented in various types of memory, including but not limited to cache, flash memory, memory cards, random access memory (RAM), dynamic RAM (DRAM), etc. The type and / or size of memory 150 can vary depending on the design criteria of a particular implementation. The data stored in memory 150 may correspond to video files, motion information (e.g., readings from sensor 164), video fusion parameters, image stabilization parameters, user input, computer vision models, feature sets, and / or metadata information. In some embodiments, memory 150 can store reference images. Reference images can be used for computer vision operations, 3D reconstruction, etc.
[0046] Processor / SoC 102 can be configured to execute computer-readable code and / or process information. In various embodiments, the computer-readable code may be stored within processor / SoC 102 (e.g., microcode, etc.) and / or in memory 150. In an example, processor / SoC 102 can be configured to execute one or more artificial neural network models (e.g., face recognition CNN, object detection CNN, object classification CNN, 3D reconstruction CNN, etc.) stored in memory 150. In an example, memory 150 may store one or more directed acyclic graphs (DAGs) and a set of or more sets of weights and biases defining one or more artificial neural network models. Processor / SoC 102 can be configured to receive input from memory 150 and / or present output to memory 150. Processor / SoC 102 can be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of processor / SoC 102 may vary depending on the design criteria of a particular implementation. Processor / SoC 102 can be configured for low-power (e.g., battery) operation.
[0047] Battery 152 can be configured to store and / or power components of camera system 100. The dynamic driver mechanism for the rolling shutter sensor can be configured to conserve power. Reduced power consumption allows camera system 100 to operate for extended periods using battery 152 without recharging. Battery 152 can be rechargeable. Battery 152 can be built-in (e.g., non-replaceable) or replaceable. Battery 152 can have an input for connecting to an external power source (e.g., for charging). In some embodiments, device 100 can be powered by an external power source (e.g., battery 152 can be omitted or implemented as a backup power source). Various battery technologies and / or chemistry can be used to implement battery 152. The type of battery 152 implemented can vary depending on the design criteria of a particular implementation.
[0048] The communication module 154 can be configured to implement one or more communication protocols. For example, the communication module 154 and the wireless interface 156 can be configured to implement one or more of the following: IEEE 102.11, IEEE 102.15, IEEE 102.15.1, IEEE 102.15.2, IEEE 102.15.3, IEEE 102.15.4, IEEE 102.15.5, IEEE 102.20, etc. and / or In some embodiments, the communication module 154 may be a hardwired data port (e.g., a USB port, a mini-USB port, a USB-C connector, an HDMI port, an Ethernet port, a DisplayPort interface, a Lightning port, etc.). In some embodiments, the wireless interface 156 may also implement one or more protocols associated with a cellular communication network (e.g., GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, SMS, etc.). In embodiments where the camera system 100 is implemented as a wireless camera, the protocol implemented by the communication module 154 and the wireless interface 156 may be a wireless communication protocol. The type of communication protocol implemented by the communication module 154 may vary depending on the design criteria of the specific implementation.
[0049] Communication module 154 and / or wireless interface 156 can be configured to generate broadcast signals as output from camera system 100. The broadcast signals can send video data, parallax data, and / or control signals to external devices. For example, the broadcast signals can be sent to cloud storage services (e.g., storage services capable of scaling on demand). In some embodiments, communication module 154 may not send data until the processor / SoC 102 has performed video analysis to determine that an object is in the field of view of camera system 100.
[0050] In some embodiments, the communication module 154 can be configured to generate a manual control signal. The manual control signal can be generated in response to a signal received from a user by the communication module 154. The manual control signal can be configured to activate the processor / SoC 102. Regardless of the power state of the camera system 100, the processor / SoC 102 can be activated in response to the manual control signal.
[0051] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive a feature set. The received feature set can be used to detect events and / or objects. For example, the feature set can be used to perform computer vision operations. The feature set information may include instructions for the processor 102 to determine which types of objects correspond to objects and / or events of interest.
[0052] Processor 158 can be implemented using general-purpose processor circuitry. Processor 158 is operable to interact with video processing circuitry 102 and memory 150 to perform various processing tasks. Processor 158 can be configured to execute computer-readable instructions. In this example, the computer-readable instructions may be stored in memory 150. In some embodiments, the computer-readable instructions may include controller operations. Typically, input from sensor 164 and / or human-machine interface device 166 is shown to be received by processor 102. In some embodiments, general-purpose processor 158 can be configured to receive and / or analyze data from sensor 164 and / or HID 166 and make decisions in response to the input. In some embodiments, processor 158 can send data to and / or receive data from other components of camera system 100, such as battery 152, communication module 154, and / or wireless interface 156. Which functions of camera system 100 are performed by processor 102 and general-purpose processor 158 may vary depending on the design criteria of the specific implementation.
[0053] Lens 160 may be attached to capture device 104. Capture device 104 may be configured to receive an input signal (e.g., LIN) via lens 160. The signal LIN may be an optical input (e.g., an analog image). Lens 160 may be implemented as an optical lens. Lens 160 may provide zoom and / or focus features. In one example, capture device 104 and / or lens 160 may be implemented as a single lens assembly. In another example, lens 160 may be implemented separately from capture device 104.
[0054] Capture device 104 can be configured to convert input light LIN into computer-readable data. Capture device 104 can capture data received through lens 160 to generate raw pixel data. In some embodiments, capture device 104 can capture data received through lens 160 to generate a bitstream (e.g., generate video frames). For example, capture device 104 can receive focused light from lens 160. Lens 160 can be oriented, tilted, translated, scaled, and / or rotated to provide a target view from camera system 100 (e.g., a view of video frames, a view of panoramic video frames captured using multiple camera systems 100a-100n, a target image and reference image view for stereo vision, etc.). Capture device 104 can generate a signal (e.g., video). The signal VIDEO can be pixel data (e.g., a sequence of pixels that can be used to generate video frames). In some embodiments, the signal VIDEO can be video data (e.g., a sequence of video frames). The signal VIDEO can be presented to one of the inputs of processor 102. In some embodiments, the pixel data generated by the capture device 104 may be uncompressed and / or raw data generated in response to focused light from the lens 160. In some embodiments, the output of the capture device 104 may be a digital video signal.
[0055] In this example, capture device 104 may include block (or circuitry) 180, block (or circuitry) 182, and block (or circuitry) 184. Circuitry 180 may be an image sensor. Circuitry 182 may be a processor and / or logic unit. Circuitry 184 may be memory circuitry (e.g., a frame buffer). Lens 160 (e.g., a camera lens) may be directed to provide a view of the environment surrounding camera system 100. Lens 160 may be designed to capture ambient data (e.g., light input LIN). Lens 160 may be a wide-angle lens and / or a fisheye lens (e.g., a lens capable of capturing a wide field of view). Lens 160 may be configured to capture and / or focus light for capture device 104. Typically, image sensor 180 is located behind lens 160. Based on the light captured from lens 160, capture device 104 may generate bitstream and / or video data (e.g., a signal VIDEO).
[0056] Capture device 104 can be configured to capture video image data (e.g., light collected and focused by lens 160). Capture device 104 can capture data received through lens 160 to generate a video bitstream (e.g., pixel data of a video frame sequence). In various embodiments, lens 160 can be implemented as a fixed-focus lens. Fixed-focus lenses are generally advantageous for smaller size and lower power. In examples, fixed-focus lenses can be used in battery-powered, doorbell, and other low-power camera applications. In some embodiments, lens 160 can be oriented, tilted, panned, zoomed, and / or rotated to capture the environment around camera system 100 (e.g., capture data from the field of view). In examples, professional camera models can be implemented using an active lens system for enhanced functionality, remote control, etc.
[0057] The capture device 104 can convert received light into a digital data stream. In some embodiments, the capture device 104 can perform analog-to-digital conversion. For example, the image sensor 180 can perform photoelectric conversion on the light received by the lens 160. The processor / logic unit 182 can convert the digital data stream into a video data stream (or bitstream), a video file, and / or multiple video frames. In the example, the capture device 104 can present the video data as a digital video signal (e.g., VIDEO). The digital video signal may include video frames (e.g., continuous digital images and / or audio). In some embodiments, the capture device 104 may include a microphone for capturing audio. In some embodiments, the microphone may be implemented as a separate component (e.g., one of the sensors 164).
[0058] Video data captured by capture device 104 can be represented as a signal / bitstream / data VIDEO (e.g., a digital video signal). Capture device 104 can present the signal VIDEO to processor / SoC 102. The signal VIDEO can represent video frames / video data. The signal VIDEO can be a video stream captured by capture device 104. In some embodiments, the signal VIDEO may include pixel data operable by processor 102 (e.g., a video processing pipeline, image signal processor (ISP), etc.). Processor 102 can generate video frames in response to the pixel data in the signal VIDEO.
[0059] The signal VIDEO may include pixel data arranged as video frames. The signal VIDEO may be an image including a background (e.g., captured objects and / or environment) and a speckle pattern generated by the structured light projector 106. The signal VIDEO may include a single-channel source image. A single-channel source image may be generated in response to capturing pixel data using the monocular lens 160.
[0060] Image sensor 180 can receive input light LIN from lens 160 and convert the light LIN into digital data (e.g., a bitstream). For example, image sensor 180 can perform photoelectric conversion on light from lens 160. In some embodiments, image sensor 180 may have additional margins that are not used as part of the image output. In some embodiments, image sensor 180 may not have additional margins. In various embodiments, image sensor 180 can be configured to generate RGB-IR video signals. In a field of view illuminated only by infrared light, image sensor 180 can generate monochrome (B / W) video signals. In a field of view illuminated by both IR and visible light, image sensor 180 can be configured to generate color information in addition to generating monochrome video signals. In various embodiments, image sensor 180 can be configured to generate video signals in response to visible light and / or infrared (IR) light.
[0061] In some embodiments, the camera sensor 180 may include a rolling shutter sensor or a global shutter sensor. In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR sensor. In some embodiments, the capture device 104 may include a rolling shutter IR sensor and an RGB sensor (e.g., implemented as separate components). In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR rolling shutter complementary metal-oxide-semiconductor (CMOS) image sensor. In one example, the rolling shutter sensor 180 may be configured to assert a signal indicating the exposure time of the first row. In one example, the rolling shutter sensor 180 may apply a mask to a monochrome sensor. In an example, the mask may include multiple cells containing a red pixel, a green pixel, a blue pixel, and an IR pixel. The IR pixel may contain red, green, and blue filter materials that efficiently absorb all light in the visible spectrum while allowing longer infrared wavelengths to pass through with minimal loss. In the case of a rolling shutter, as each row (or line) of the sensor begins exposure, all pixels in that row (or line) may begin exposure simultaneously.
[0062] Processor / logic unit 182 can convert the bitstream into human-readable content (e.g., video data that an average person can understand regardless of image quality, such as video frames and / or pixel data that can be converted into video frames by processor 102). For example, processor / logic unit 182 can receive raw (e.g., raw) data from image sensor 180 and generate (e.g., encode) video data (e.g., bitstream) based on the raw data. Capture device 104 may have memory 184 to store raw data and / or processed bitstreams. For example, capture device 104 may implement frame memory and / or buffer 184 to store (e.g., provide temporary storage and / or cache) one or more video frames (e.g., digital video signals). In some embodiments, processor / logic unit 182 can perform analysis and / or correction on the video frames stored in memory / buffer 184 of capture device 104. Processor / logic unit 182 can provide status information about the captured video frames.
[0063] The structured light projector 106 may include a block (or circuit) 186. Circuit 186 may implement a structured light source. The structured light source 186 may be configured to generate a signal (e.g., a speckle pattern). The signal SLP may be a structured light pattern (e.g., a speckle pattern). The signal SLP may be projected onto the environment near the camera system 100. The structured light pattern SLP may be captured by a capture device 104 as part of the light input LIN.
[0064] The structured light patterning lens 162 can be a lens for the structured light projector 106. The structured light patterning lens 162 can be configured to allow the structured light SLP generated by the structured light source 186 of the structured light projector 106 to be emitted, while protecting the structured light source 186. The structured light patterning lens 162 can be configured to decompose the laser pattern generated by the structured light source 186 into a pattern array (e.g., a dense dot pattern array for speckle patterns).
[0065] In the example, the structured light source 186 can be implemented as an array of vertical cavity surface-emitting lasers (VCSELs) and lenses. However, other types of structured light sources can be implemented to meet the design criteria of specific applications. In the example, the VCSEL array is typically configured to generate laser patterns (e.g., signal SLP). The lenses are typically configured to decompose the laser pattern into an array of dense dot patterns. In the example, the structured light source 186 can be implemented as a near-infrared (NIR) light source. In various embodiments, the light source of the structured light source 186 can be configured to emit light with a wavelength of approximately 940 nanometers (nm), which is invisible to the human eye. However, other wavelengths can be utilized. In the example, wavelengths in the range of approximately 800 nm to 1000 nm can be utilized.
[0066] Sensor 164 can implement multiple sensors, including but not limited to motion sensors, ambient light sensors, proximity sensors (e.g., ultrasonic, radar, lidar, etc.), audio sensors (e.g., microphones), etc. In embodiments implementing a motion sensor, sensor 164 can be configured to detect motion anywhere (or some location outside the field of view) monitored by camera system 100. In various embodiments, motion detection can be used as a threshold for activating capture device 104. Sensor 164 can be implemented as an internal component of camera system 100 and / or an external component of camera system 100. In one example, sensor 164 can be implemented as a passive infrared (PIR) sensor. In another example, sensor 164 can be implemented as a smart motion sensor. In yet another example, sensor 164 can be implemented as a microphone. In embodiments implementing a smart motion sensor, sensor 164 may include a low-resolution image sensor configured to detect motion and / or people.
[0067] In various embodiments, sensor 164 may generate signals (e.g., SENS). The SENS may include various data (or information) collected by sensor 164. In an example, the SENS may include data collected in response to motion detected in the monitored field of view, ambient light levels in the monitored field of view, and / or sound picked up in the monitored field of view. However, other types of data may be collected and / or generated based on application-specific design criteria. The SENS may be presented to processor / SoC 102. In an example, sensor 164 may generate (assert) the SENS when motion is detected in the field of view monitored by camera system 100. In another example, sensor 164 may generate (assert) the SENS when audio is triggered in the field of view monitored by camera system 100. In yet another example, sensor 164 may be configured to provide directional information about motion and / or sound detected in the field of view. This directional information may also be transmitted to processor / SoC 102 via the SENS.
[0068] HID 166 can implement an input device. For example, HID 166 can be configured to receive human input. In one example, HID 166 can be configured to receive password input from a user. In some embodiments, camera system 100 may include a keyboard, touchpad (or screen), doorbell switch, and / or other human-machine interface device (HID) 166. In an example, sensor 164 may be configured to determine when an object approaches HID 166. In an example where camera system 100 is implemented as part of an access control application, capture device 104 may be activated to provide an image for identifying a person attempting access, and a lock area and / or illumination for the access touchpad 166 may be activated. For example, a combination of input from HID 166 (e.g., a password or PIN code) may be combined with activity determination and / or depth analysis performed by processor 102 to achieve two-factor authentication.
[0069] The processor / SoC 102 can receive a signal VIDEO and a signal SENS. The processor / SoC 102 can generate one or more video output signals (e.g., VIDOUT), one or more control signals (e.g., CTRL), and / or one or more depth data signals (e.g., DIMAGES) based on the signal VIDEO, the signal SENS, and / or other inputs. In some embodiments, the signals VIDOUT, DIMAGES, and CTRL can be generated based on analysis of the signal VIDEO and / or objects detected in the signal VIDEO.
[0070] In various embodiments, the processor / SoC 102 may be configured to perform one or more of the following: feature extraction, object detection, object tracking, 3D reconstruction, and object recognition. For example, the processor / SoC 102 may determine motion information and / or depth information by analyzing frames from a signal VIDEO and comparing those frames with previous frames. The comparison may be used to perform digital motion estimation. In some embodiments, the processor / SoC 102 may be configured to generate a video output signal VIDOUT including video data and / or a depth data signal DIMAGES including a disparity map and a depth map from the signal VIDEO. The video output signal VIDOUT and / or the depth data signal DIMAGES may be presented to memory 150, communication module 154, and / or wireless interface 156. In some embodiments, the video signal VIDOUT and / or the depth data signal DIMAGES may be used internally by the processor 102 (e.g., not presented as output).
[0071] The signal VIDOUT can be presented to the communication device 156. In some embodiments, the signal VIDOUT may include encoded video frames generated by the processor 102. In some embodiments, the encoded video frames may include a complete video stream (e.g., encoded video frames representing all video captured by the capture device 104). The encoded video frames may be encoded, cropped, stitched, and / or enhanced versions of pixel data received from the signal VIDEO. In the example, the encoded video frames may be high-resolution, digital, encoded, de-distorted, stabilized, cropped, blended, stitched, and / or rolling shutter effect corrected versions of the signal VIDEO.
[0072] In some embodiments, the signal VIDOUT may be generated based on video analysis (e.g., computer vision operations) performed by processor 102 on generated video frames. Processor 102 may be configured to perform computer vision operations to detect objects and / or events in the video frames, and then convert the detected objects and / or events into statistical data and / or parameters. In one example, the data determined by the computer vision operations may be converted by processor 102 into a human-readable format. Data from the computer vision operations can be used to detect objects and / or events. The computer vision operations may be performed locally by processor 102 (e.g., without needing to communicate with external devices to offload computational operations). For example, locally executed computer vision operations allow the computer vision operations to be performed by processor 102 and avoid heavy video processing running on a backend server. Avoiding video processing on a backend (e.g., remote location) server can protect privacy.
[0073] In some embodiments, the signal VIDOUT can be data generated by processor 102 (e.g., video analysis results, audio / speech analysis results, etc.) that can be transmitted to a cloud computing service for information aggregation and / or to provide training data for machine learning (e.g., to improve object detection, improve audio detection, etc.). In some embodiments, the signal VIDOUT can be provided to a cloud service for mass storage (e.g., to enable users to retrieve encoded video using smartphones and / or desktop computers). In some embodiments, the signal VIDOUT can include data extracted from video frames (e.g., computer vision results) and can transmit the results to another device (e.g., a remote server, a cloud computing system, etc.) to offload the analysis of the results to another device (e.g., offloading the analysis of the results to a cloud computing service instead of performing all the analysis locally). The type of information transmitted by the signal VIDOUT can vary depending on the design criteria of a particular implementation.
[0074] The CTRL signal can be configured to provide a control signal. The CTRL signal can be generated in response to a decision made by processor 102. In an example, the CTRL signal can be generated in response to a detected object and / or features extracted from a video frame. The CTRL signal can be configured to enable, disable, or change the operating mode of another device. In one example, the CTRL signal can be used to lock / unlock a door controlled by an electronic lock. In another example, the device can be set to sleep mode (e.g., low power mode) and / or activated from sleep mode in response to the CTRL signal. In yet another example, the CTRL signal can be used to generate an alarm and / or notification. The type of device controlled by the CTRL signal and / or the response performed by the device in response to the CTRL signal can vary depending on the design criteria of the specific implementation.
[0075] The signal CTRL can be generated based on data received by sensor 164 (e.g., temperature readings, motion sensor readings, etc.). The signal CTRL can be generated based on input from HID 166. The signal CTRL can be generated based on human behavior detected by processor 102 in a video frame. The signal CTRL can be generated based on the type of detected object (e.g., person, animal, vehicle, etc.). The signal CTRL can be generated in response to the detection of a specific type of object at a specific location. Processor 102 can be configured to generate the signal CTRL in response to sensor fusion operations (e.g., aggregating information received from different sources). The conditions used to generate the signal CTRL can vary depending on the design criteria of the specific implementation.
[0076] Signals DIMAGES may include one or more depth maps and / or disparity maps generated by processor 102. Signals DIMAGES may be generated in response to 3D reconstruction performed on a monocular single-channel image. Signals DIMAGES may be generated in response to analysis of captured video data and structured light pattern SLP.
[0077] A multi-step approach can be implemented to activate and / or disable the capture device 104 based on the output of the motion sensor 164 and / or any other power consumption characteristics of the camera system 100, thereby reducing the power consumption of the camera system 100 and extending the lifespan of the battery 152. The motion sensor in sensor 164 can have low power consumption on the battery 152 (e.g., less than 10W). In the example, the motion sensor of sensor 164 can be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 102. Video analysis performed by the processor / SoC 102 may have relatively high power consumption on the battery 152 (e.g., greater than that of the motion sensor 164). In the example, the processor / SoC 102 can be in a low-power state (or powered off) until some motion is detected by the motion sensor of sensor 164.
[0078] The camera system 100 can be configured to operate using various power states. For example, in a power-off state (e.g., sleep state, low power state), sensor 164 and the motion sensor of processor / SoC 102 can be turned on, and other components of the camera system 100 (e.g., image capture device 104, memory 150, communication module 154, etc.) can be turned off. In another example, the camera system 100 can operate in an intermediate state. In the intermediate state, image capture device 104 can be turned on, and memory 150 and / or communication module 154 can be turned off. In yet another example, the camera system 100 can operate in a powered-on (or high-power) state. In the powered-on state, sensor 164, processor / SoC 102, capture device 104, memory 150, and / or communication module 154 can be turned on. The camera system 100 can consume some power from battery 152 (e.g., a relatively small and / or minimal amount of power) in the power-off state. In the powered-on state, the camera system 100 can consume more power from battery 152. The number of power states and / or the number of components of the camera system 100 that are turned on when the camera system 100 operates in each power state can vary according to the design criteria of a particular implementation.
[0079] In some embodiments, the camera system 100 may be implemented as a system-on-a-chip (SoC). For example, the camera system 100 may be implemented as a printed circuit board including one or more components. The camera system 100 may be configured to perform intelligent video analysis on video frames of a video. The camera system 100 may be configured to crop and / or enhance the video.
[0080] In some embodiments, the video frame may be a view (or a derivative of a view) captured by the capture device 104. Pixel data signals may be enhanced by the processor 102 (e.g., color conversion, noise filtering, automatic exposure, automatic white balance, automatic focus, etc.). In some embodiments, the video frame may provide a series of cropped and / or enhanced video frames that improve the view from the perspective of the camera system 100 (e.g., providing night vision, providing high dynamic range (HDR) imaging, providing more viewing area, highlighting detected objects, providing additional data (e.g., digital distance to the detected object), etc.) to enable the processor 102 to see the location better than a human can see with human vision.
[0081] Encoded video frames can be processed locally. In one example, the encoded video can be stored locally by memory 150 so that processor 102 can facilitate computer vision analysis internally (e.g., without first uploading the video frames to a cloud service). Processor 102 can be configured to select video frames to be encapsulated into a video stream that can be transmitted over a network (e.g., a bandwidth-limited network).
[0082] In some embodiments, processor 102 may be configured to perform sensor fusion operations. The sensor fusion operations performed by processor 102 may be configured to analyze information from multiple sources (e.g., capture device 104, sensor 164, and HID 166). By analyzing various data from different sources, the sensor fusion operations may be able to make inferences about the data that might not be possible from just one data source. For example, the sensor fusion operations implemented by processor 102 may analyze video data (e.g., human mouth movements) and speech patterns from directional audio. Different sources can be used to develop scene models to support decision-making. For example, processor 102 may be configured to compare the synchronization of detected speech patterns with mouth movements in video frames to determine which person is speaking in the video frame. The sensor fusion operations may also provide temporal correlation, spatial correlation, and / or reliability of the received data.
[0083] In some embodiments, processor 102 may implement convolutional neural network (CNN) capabilities. CNN capabilities can be implemented using deep learning techniques for computer vision. CNN capabilities can be configured to perform pattern and / or image recognition using a training process involving multi-layer feature detection. Computer vision and / or CNN capabilities can be executed locally by processor 102. In some embodiments, processor 102 may receive training data and / or feature set information from external sources. For example, external devices (e.g., cloud services) may access various data sources to provide training data that the camera system 100 may not be able to obtain. However, computer vision operations performed using feature sets can be performed using the computational resources of processor 102 within the camera system 100.
[0084] The video pipeline of processor 102 can be configured to perform local dedistortion, cropping, enhancement, rolling shutter correction, stabilization, downsizing, packing, compression, conversion, mixing, synchronization, and / or other video operations. The video pipeline of processor 102 can enable multi-stream support (e.g., generating multiple bitstreams in parallel, each including a different bitrate). In the example, the video pipeline of processor 102 can implement an image signal processor (ISP) with an input pixel rate of 320 Mbps. The architecture of the video pipeline of processor 102 enables real-time and / or near real-time video operations on high-resolution video and / or high-bitrate video data. The video pipeline of processor 102 can implement computer vision processing, stereo vision processing, object detection, 3D noise reduction, fisheye lens correction (e.g., real-time 360-degree dedistortion and lens distortion correction), oversampling, and / or high dynamic range processing on 4K resolution video data. In one example, the video pipeline architecture can achieve 4K ultra-high resolution with H.264 encoding at dual real-time rates (e.g., 60fps) and 4K ultra-high resolution with H.265 / HEVC and / or 4K AVC encoding at 30fps (e.g., multi-stream 4KP30 AVC and HEVC encoding). The type of video operation and / or the type of video data operated on by the processor 102 can vary depending on the design criteria of the specific implementation.
[0085] The camera sensor 180 can be a high-resolution sensor. Using the high-resolution sensor 180, the processor 102 can combine oversampling of the image sensor 180 with digital scaling within the cropped area. Each of oversampling and digital scaling can be one of the video operations performed by the processor 102. Oversampling and digital scaling can be implemented to provide a higher resolution image within the overall size constraints of the cropped area.
[0086] In some embodiments, lens 160 may be a fisheye lens. One of the video operations performed by processor 102 may be a de-distortion operation. Processor 102 may be configured to de-distort the generated video frames. De-distortion may be configured to reduce and / or remove severe distortion caused by fisheye lens and / or other lens characteristics. For example, de-distortion may reduce and / or eliminate bulging effects to provide a linear image.
[0087] Processor 102 can be configured to crop (e.g., trim) a region of interest from a full video frame (e.g., generate a region of interest video frame). Processor 102 can generate video frames and select regions. In the example, cropping the region of interest can generate a second image. The cropped image (e.g., the region of interest video frame) can be smaller than the original video frame (e.g., the cropped image can be a portion of the captured video).
[0088] The region of interest (ROI) can be dynamically adjusted based on the location of the audio source. For example, the detected audio source may be moving, and its location may shift as video frames are captured. Processor 102 can update the coordinates of the selected ROI and dynamically update the cropped portion (e.g., a directional microphone implemented as one or more sensors in sensor 164 can dynamically update its position based on captured directional audio). The cropped portion may correspond to the selected ROI. As the ROI changes, the cropped portion may change. For example, the selected coordinates of the ROI may change frame by frame, and processor 102 may be configured to crop the selected region in each frame.
[0089] Processor 102 can be configured to oversample image sensor 180. Oversampling of image sensor 180 can produce a higher resolution image. Processor 102 can also be configured to digitally magnify regions of video frames. For example, processor 102 can digitally magnify a cropped region of interest. For example, processor 102 can establish a region of interest based on directional audio, crop the region of interest, and then digitally magnify the cropped region of interest video frame.
[0090] The de-distortion operation performed by processor 102 can adjust the visual content of the video data. The adjustment performed by processor 102 can make the visual content look natural (e.g., look as if it were seen by a person viewing a position corresponding to the field of view of capture device 104). In the example, de-distortion can alter the video data to generate linear video frames (e.g., correcting artifacts caused by lens characteristics of lens 160). De-distortion operations can be implemented to correct distortions caused by lens 160. Adjusted visual content can be generated to achieve more accurate and / or reliable object detection.
[0091] Various features (e.g., de-distortion, digital scaling, cropping, etc.) can be implemented as hardware modules in processor 102. Implementing hardware modules can increase the video processing speed of processor 102 (e.g., faster than software implementation). Hardware implementation enables video to be processed while reducing latency. The hardware components used can vary depending on the design standards of the specific implementation.
[0092] Processor 102 is shown as comprising multiple blocks (or circuits) 190a-190n. Blocks 190a-190n can implement various hardware modules implemented by processor 102. Hardware modules 190a-190n can be configured to provide various hardware components to implement a video processing pipeline. Circuits 190a-190n can be configured to receive pixel data VIDEO, generate video frames from the pixel data, perform various operations on the video frames (e.g., de-distortion, rolling shutter correction, cropping, magnification, image stabilization, 3D reconstruction, etc.), prepare video frames for communication with external hardware (e.g., encoding, encapsulation, color correction, etc.), parse feature sets, and implement various operations for computer vision (e.g., object detection, segmentation, classification, etc.). Hardware modules 190a-190n can be configured to implement various security features (e.g., secure boot, I / O virtualization, etc.). Various implementations of processor 102 may not necessarily utilize all features of hardware modules 190a-190n. The features and / or functions of hardware modules 190a-190n may vary depending on the design criteria of a particular implementation. Details of hardware modules 190a-190n can be described in conjunction with U.S. Patent Application No. 16 / 831,549, filed April 16, 2020; U.S. Patent Application No. 16 / 288,922, filed February 28, 2019; U.S. Patent Application No. 15 / 593,493 (now U.S. Patent No. 10,437,600), filed May 12, 2017; U.S. Patent Application No. 15 / 931,942, filed May 14, 2020; and U.S. Patent Application No. 16 / 991,344, filed August 12, 2020, the appropriate portions of which are incorporated herein by reference in their entirety.
[0093] Hardware modules 190a-190n can be implemented as dedicated hardware modules. Compared to software implementation, using dedicated hardware modules 190a-190n to implement various functions of processor 102 allows processor 102 to be highly optimized and / or customized to limit power consumption, reduce heat generation, and / or increase processing speed. Hardware modules 190a-190n can be customizable and / or programmable to implement multiple types of operations. Implementing dedicated hardware modules 190a-190n allows the hardware used to perform each type of computation to be optimized for speed and / or efficiency. For example, hardware modules 190a-190n can implement several relatively simple operations frequently used in computer vision operations, which together enable computer vision operations to be performed in real time. The video pipeline can be configured to: identify objects. Objects can be identified by interpreting numerical and / or symbolic information to determine that visual data represents a specific type of object and / or feature. For example, the number of pixels and / or pixel color of video data can be used to identify portions of the video data as objects. Hardware modules 190a-190n enable computationally intensive operations (e.g., computer vision operations, video encoding, video transcoding, 3D reconstruction, etc.) to be performed locally by the camera system 100.
[0094] One of the hardware modules 190a-190n (e.g., 190a) can implement a scheduler circuit. Scheduler circuit 190a can be configured to store a directed acyclic graph (DAG). In the example, scheduler circuit 190a can be configured to generate and store a DAG in response to received (e.g., loaded) feature set information. The DAG can define video operations to be performed to extract data from video frames. For example, the DAG can define various mathematical weights (e.g., neural network weights and / or biases) to be applied when performing computer vision operations to classify various groups of pixels into specific objects.
[0095] Scheduler circuit 190a can be configured to parse acyclic graphs to generate various operators. Operators can be scheduled by scheduler circuit 190a in one or more hardware modules 190a-190n. For example, one or more hardware modules 190a-190n can implement a hardware engine configured to perform a specific task (e.g., a hardware engine designed to perform repetitive specific mathematical operations used for performing computer vision operations). Scheduler circuit 190a can schedule operators based on when they are ready to be processed by hardware engines 190a-190n.
[0096] Scheduler circuit 190a can multiplex task time across hardware modules 190a-190n based on their availability. Scheduler circuit 190a can parse a directed acyclic graph (DAG) into one or more data streams. Each data stream can include one or more operators. Once the DAG is parsed, scheduler circuit 190a can assign data streams / operators to hardware engines 190a-190n and send relevant operator configuration information to initiate the operators.
[0097] Each directed acyclic sphere Figure 2 The radix representation can be an ordered traversal of a directed acyclic graph, where descriptors and operators are intertwined based on data dependencies. Descriptors typically provide registers that link data buffers to specific operands in the associated operators. In various embodiments, operators may not appear in the directed acyclic graph representation until all associated descriptors have been declared for operands.
[0098] One of the hardware modules 190a-190n (e.g., 190b) can implement a convolutional neural network (CNN) module. CNN module 190b can be configured to perform computer vision operations on video frames. CNN module 190b can be configured to recognize objects through multi-layer feature detection. CNN module 190b can be configured to compute descriptors based on the performed feature detections. The descriptors enable processor 102 to determine the probability that a pixel in a video frame corresponds to a specific object (e.g., a specific brand / model / year of a vehicle, identifying a person as a specific individual, detecting an animal type, etc.).
[0099] CNN module 190b can be configured to: implement convolutional neural network capabilities; implement computer vision using deep learning techniques; implement pattern and / or image recognition using a training process involving multi-layer feature detection; and perform inference against machine learning models.
[0100] CNN module 190b can be configured to perform feature extraction and / or matching solely in hardware. Feature points typically represent regions of interest (e.g., corners, edges, etc.) within a video frame. By tracking feature points over time, estimates of the self-motion of the capture platform or motion models of observed objects in the scene can be generated. To track feature points, the matching operation is typically incorporated into CNN module 190b by hardware to find the most probable correspondence between feature points in a reference and target video frame. During the matching of reference and target feature point pairs, each feature point can be represented by a descriptor (e.g., image patch, SIFT, BRIEF, ORB, FREAK, etc.). Implementing CNN module 190b using dedicated hardware circuitry enables real-time computation of descriptor matching distances.
[0101] The CNN module 190b can be configured to perform face detection, face recognition, and / or liveness detection. For example, face detection, face recognition, and / or liveness detection can be performed based on a trained neural network implemented by the CNN module 190b. In some embodiments, the CNN module 190b can be configured to generate a depth image from a structured light pattern. The CNN module 190b can be configured to perform various detection and / or recognition operations and / or perform 3D recognition operations.
[0102] CNN module 190b can be a dedicated hardware module configured to perform feature detection of video frames. Features detected by CNN module 190b can be used to compute descriptors. CNN module 190b can determine the probability that a pixel in a video frame belongs to a specific object and / or some objects in response to the descriptors. For example, using the descriptors, CNN module 190b can determine the probability that a pixel corresponds to a specific object (e.g., a person, a piece of furniture, a pet, a vehicle, etc.) and / or characteristics of the object (e.g., the shape of the eyes, the distance between facial features, a vehicle hood, body parts, a vehicle license plate, a face, clothing worn by a person, etc.). Implementing CNN module 190b as a dedicated hardware module of processor 102 enables device 100 to perform computer vision operations locally (e.g., on-chip) without relying on the processing power of a remote device (e.g., transmitting data to a cloud computing service).
[0103] The computer vision operations performed by the CNN module 190b can be configured to: perform feature detection on video frames to generate descriptors. The CNN module 190b can perform object detection to determine regions in the video frames with a high probability of matching a specific object. In one example, the type of object to be matched (e.g., a reference object) can be customized using an open operand stack (enabling the programmability of the processor 102 to implement various artificial neural networks defined by directed acyclic graphs, each providing instructions for performing various types of object detection). The CNN module 190b can be configured to: perform local masking on regions with a high probability of matching a specific object to detect the object.
[0104] In some embodiments, the CNN module 190b can determine the location (e.g., 3D coordinates and / or position coordinates) of various features (e.g., characteristics) of the detected object. In one example, 3D coordinates can be used to determine the position of a person's arms, legs, chest, and / or eyes. A position coordinate on a first axis representing the vertical position of a body part in 3D space and another coordinate on a second axis representing the horizontal position of a body part in 3D space can be stored. In some embodiments, the distance from the lens 160 can represent a coordinate of the depth position of the body part in 3D space (e.g., a position coordinate on a third axis). Using the various body parts' positions in 3D space, the processor 102 can determine the detected person's body position and / or body characteristics.
[0105] The CNN module 190b can be pre-trained (e.g., configured to perform computer vision to detect objects based on training data received to train the CNN module 190b). For example, the results of the training data (e.g., a machine learning model) can be pre-programmed and / or loaded into processor 102. The CNN module 190b can perform inference on the machine learning model (e.g., to perform object detection). Training can include determining weight values for each layer of the neural network model. For example, weight values can be determined for each layer for feature extraction (e.g., convolutional layers) and / or for classification (e.g., fully connected layers). The weight values learned by the CNN module 190b can vary depending on the design criteria of a particular implementation.
[0106] The CNN module 190b can perform feature extraction and / or object detection by performing convolution operations. These convolution operations can be hardware-accelerated for fast (e.g., real-time) computation that can be performed with low power consumption. In some embodiments, the convolution operations performed by the CNN module 190b can be used to perform computer vision operations. In some embodiments, the convolution operations performed by the CNN module 190b can be used for any function (e.g., 3D reconstruction) that may involve computing convolution operations and is performed by the processor 102.
[0107] Convolution operations can include sliding a feature detection window along a layer while performing computations (e.g., matrix operations). The feature detection window can apply filters to pixels and / or extract features associated with each layer. The feature detection window can be applied to a single pixel and multiple surrounding pixels. In the example, these layers can be represented as matrices representing the values of pixels and / or features of one of these layers, and the filters applied by the feature detection window can be represented as matrices. Convolution operations can apply matrix multiplication between regions of the current layer covered by the feature detection window. Convolution operations can slide the feature detection window along the region of the layer to generate a result representing each region. The size of the regions, the type of operation for applying filters, and / or the number of layers can vary depending on the design criteria of a particular implementation.
[0108] Using convolutional operations, the CNN module 190b can compute multiple features of pixels in the input image at each extraction step. For example, each layer can receive input from a set of features located in a small neighborhood (e.g., a region) of the previous layer (e.g., a local receptive field). Convolutional operations can extract basic visual features (e.g., oriented edges, endpoints, corners, etc.), which are then combined by higher layers. Because the feature extraction window operates on pixels and their nearby pixels (or subpixels), the results of the operations may be position-invariant. These layers can include convolutional layers, pooling layers, non-linear layers, and / or fully connected layers. In the example, convolutional operations can learn to detect edges from raw pixels (e.g., the first layer), then use features from the previous layer (e.g., detected edges) to detect shapes in the next layer, and then use these shapes to detect higher-level features in higher layers (e.g., facial features, pets, vehicles, vehicle parts, furniture, etc.), and the final layer may be a classifier using higher-level features.
[0109] The CNN module 190b can perform data stream operations for feature extraction and matching, including two-stage detection, transformation operators, component operators manipulating component lists (e.g., components can be regions of vectors sharing common attributes and can be combined with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, min-max-reduction operators, pooling operators, non-minimum and non-maximum suppression operators, non-maximum suppression operators based on scanning windows, aggregation operators, dispersion operators, statistical operators, classification operators, integral image operators, comparison operators, indexing operators, pattern matching operators, feature extraction operators, feature detection operators, two-stage object detection operators, score generation operators, block reduction operators, and upsampling operators. The types of operations performed by the CNN module 190b for extracting features from training data can vary depending on the design criteria of a particular implementation.
[0110] Each of the hardware modules 190a-190n can implement a processing resource (or hardware resource or hardware engine). Hardware engines 190a-190n can be used to perform specific processing tasks. In some configurations, hardware engines 190a-190n can operate in parallel and independently of each other. In other configurations, hardware engines 190a-190n can operate collaboratively with each other to perform assigned tasks. One or more hardware engines among the hardware engines 190a-190n can be homogeneous processing resources (all circuits 190a-190n can have the same capabilities) or heterogeneous processing resources (two or more circuits 190a-190n can have different capabilities).
[0111] refer to Figure 4 The diagram illustrates the processing circuitry of a camera system 100 configured to perform 3D reconstruction using a convolutional neural network. In this example, the processing circuitry of the camera system 100 can be configured for a variety of applications, including but not limited to: autonomous and semi-autonomous vehicles (e.g., cars, trucks, motorcycles, agricultural machinery, drones, aircraft, etc.), manufacturing, and / or security and surveillance systems. Compared to a general-purpose computer, the processing circuitry of the camera system 100 typically includes hardware circuitry optimized to provide high-performance image processing and computer vision pipelines with minimal area and power consumption. In this example, various operations for performing image processing, feature detection / extraction, 3D reconstruction, and / or object detection / classification for computer (or machine) vision can be implemented using hardware modules designed to reduce computational complexity and utilize resources efficiently.
[0112] In an example embodiment, processing circuitry 100 may include processor 102, memory 150, general-purpose processor 158, and / or memory bus 200. General-purpose processor 158 may implement a first processor. Processor 102 may implement a second processor. In this example, circuitry 102 may implement a computer vision processor. In this example, processor 102 may be an intelligent vision processor. Memory 150 may implement external memory (e.g., memory outside of circuitry 158 and 102). In this example, circuitry 150 may be implemented as dynamic random access memory (DRAM) circuitry. The processing circuitry of camera system 100 may include other components (not shown). The number, type, and / or arrangement of components in the processing circuitry of camera system 100 may vary depending on the design criteria of a particular implementation.
[0113] A general-purpose processor 158 can operate to interact with circuits 102 and 150 to perform various processing tasks. In an example, processor 158 can be configured as a controller for circuit 102. Processor 158 can be configured to execute computer-readable instructions. In one example, the computer-readable instructions may be stored by circuit 150. In some embodiments, the computer-readable instructions may include controller operations. Processor 158 can be configured to communicate with circuit 102 and / or access results produced by components of circuit 102. In an example, processor 158 can be configured to utilize circuit 102 to perform operations associated with one or more neural network models.
[0114] In the example, processor 102 typically includes scheduler circuitry 190a, block (or circuitry) 202, one or more blocks (or circuitries) 204a-204n, block (or circuitry) 206, and path 208. Block 202 may implement a directed acyclic graph (DAG) memory. DAG memory 202 may include CNN module 190b and / or weights / biases 210. Blocks 204a-204n may implement hardware resources (or engines). Block 206 may implement shared memory circuitry. In the example embodiment, one or more circuits in circuits 204a-204n may include blocks (or circuitries) 212a-212n. In the illustrated example, circuits 212a and 212b are implemented as representative examples in the corresponding hardware engines 204a-204b. One or more circuits in circuits 202, 204a-204n, and / or circuit 206 may be related to... Figure 3 Example implementations of hardware modules 190a-190n are shown in association.
[0115] In the example, processor 158 may be configured to program circuit 102 using one or more pre-trained artificial neural network models (ANNs), including a convolutional neural network (CNN) 190b having multiple output frames according to embodiments of the invention and weights / kernels (WGTS) 210 used by the CNN module 190b. In various embodiments, the CNN module 190b may be configured (trained) for operation in an edge device. In the example, the processing circuitry of camera system 100 may be coupled to a sensor (e.g., a video camera, etc.) configured to generate data input. The processing circuitry of camera system 100 may be configured to generate one or more outputs in response to data input from the sensor, based on one or more inferences made by the pre-trained CNN module 190b using weights / kernels (WGTS) 210. The operations performed by processor 158 may vary depending on the design criteria of a particular implementation.
[0116] In various embodiments, circuit 150 may implement dynamic random access memory (DRAM) circuitry. Circuit 150 is typically operable as a multidimensional array for storing input data elements and various forms of output data elements. Circuit 150 may exchange input data elements and output data elements with processor 158 and processor 102.
[0117] Processor 102 may implement computer vision processor circuitry. In examples, processor 102 may be configured to implement various functions for computer vision. Processor 102 is generally operable to perform specific processing tasks arranged by processor 158. In various embodiments, all or part of processor 102 may be implemented individually in hardware. Processor 102 may directly execute data streams involving the execution of CNN module 190b and generated by software (e.g., directed acyclic graph, etc.) for a specified processing task (e.g., computer vision, 3D reconstruction, etc.). In some embodiments, processor 102 may be a representative example of a multitude of computer vision processors implemented by the processing circuitry of camera system 100 and configured to operate together.
[0118] In one example, circuit 212a can implement a convolution operation. In another example, circuit 212b can be configured to provide a dot product operation. Convolution and dot product operations can be used to perform computer (or machine) vision tasks (e.g., as part of an object detection process). In yet another example, one or more of the circuits 204c-204n may include blocks (or circuits) 212c-212n (not shown) for providing multidimensional convolution computation. In yet another example, one or more of the circuits 204a-204n can be configured to perform a 3D reconstruction task.
[0119] In the example, circuit 102 can be configured to receive a directed acyclic graph (DAG) from processor 158. The DAG received from processor 158 can be stored in DAG memory 202. Circuit 102 can be configured to perform the DAG for CNN module 190b using circuits 190a, 204a-204n, and 206.
[0120] Multiple signals (e.g., OP_A-OP_N) can be exchanged between circuit 190a and corresponding circuits 204a-204n. Each signal in OP_A-OP_N can convey execution operation information and / or concession operation information. Multiple signals (e.g., MEM_A-MEM_N) can be exchanged between the respective circuits 204a-204n and circuit 206. Signals MEM_A-MEM_N can carry data. Signals (e.g., DRAM) can be exchanged between circuit 150 and circuit 206. Signal DRAM can transmit data between circuits 150 and 190a (e.g., on transmission path 208).
[0121] Circuit 190a can implement a scheduler circuit. Scheduler circuit 190a is typically operable to schedule tasks within circuits 204a-204n to perform various computer vision-related tasks defined by processor 158. Individual tasks can be assigned to circuits 204a-204n by scheduler circuit 190a. Scheduler circuit 190a can assign individual tasks in response to parsing a directed acyclic graph (DAG) provided by processor 158. Scheduler circuit 190a can time-multiplex tasks across circuits 204a-204n based on their availability for performing tasks.
[0122] Each circuit 204a-204n can implement a processing resource (or hardware engine). Hardware engines 204a-204n are typically operable to perform a specific processing task. Hardware engines 204a-204n can be implemented as including dedicated hardware circuitry optimized for high performance and low power consumption while performing a specific processing task. In some configurations, hardware engines 204a-204n can operate in parallel and independently of each other. In other configurations, hardware engines 204a-204n can operate collaboratively with each other to perform assigned tasks.
[0123] Hardware engines 204a-204n can be homogeneous processing resources (e.g., all circuits 204a-204n can have the same capabilities) or heterogeneous processing resources (e.g., two or more circuits 204a-204n can have different capabilities). Hardware engines 204a-204n are typically configured to execute operators, which may include, but are not limited to, resampling operators, transformation operators, component operators manipulating a list of components (e.g., components can be vector regions sharing common attributes and can be grouped together with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, minimum-maximum-reduction operators, pooling operators, non-minimum, non-maximum suppression operators, aggregation operators, dispersion operators, statistical operators, classification operators, integral image operators, upsampling operators, and powers of two downsampling operators, etc.
[0124] In various embodiments, hardware engines 204a-204n can be implemented as individual hardware circuits. In some embodiments, hardware engines 204a-204n can be implemented as general-purpose engines that can be configured to operate as dedicated machines (or engines) through circuit customization and / or software / firmware. In some embodiments, hardware engines 204a-204n can alternatively be implemented as one or more instances or threads of program code executing on processor 158 and / or one or more processors 102, including but not limited to vector processors, central processing units (CPUs), digital signal processors (DSPs), or graphics processing units (GPUs). In some embodiments, scheduler 190a can select one or more of hardware engines 204a-204n for a specific process and / or thread. Scheduler 190a can be configured to assign hardware engines 204a-204n to a specific task in response to parsing a directed acyclic graph stored in DAG memory 202.
[0125] Circuit 206 can implement shared memory circuitry. Shared memory 206 can be configured to store data in response to input requests and / or present data in response to output requests (e.g., requests from processor 158, DRAM 150, scheduler circuitry 190a, and / or hardware engines 204a-204n). In this example, shared memory circuitry 206 can implement on-chip memory for computer vision processor 102. Shared memory 206 is generally operable to store all or part of a multidimensional array (or vector) of input and output data elements generated and / or utilized by hardware engines 204a-204n. Input data elements can be transferred from DRAM circuitry 150 to shared memory 206 via memory bus 200. Output data elements can be sent from shared memory 206 to DRAM circuitry 150 via memory bus 200.
[0126] Path 208 can implement a transfer path within processor 102. Transfer path 208 is typically operable to move data from scheduler circuit 190a to shared memory 206. Transfer path 208 can also be operable to move data from shared memory 206 to scheduler circuit 190a.
[0127] Processor 158 is shown communicating with computer vision processor 102. Processor 158 may be configured as a controller of computer vision processor 102. In some embodiments, processor 158 may be configured to transmit instructions to scheduler 190a. For example, processor 158 may provide one or more directed acyclic graphs (DAGs) to scheduler 190a via DAG memory 202. Scheduler 190a may initialize and / or configure hardware engines 204a-204n in response to parsing the DAGs. In some embodiments, processor 158 may receive status information from scheduler 190a. For example, scheduler 190a may provide processor 158 with status information and / or readiness status from the outputs of hardware engines 204a-204n, enabling processor 158 to determine one or more next instructions to execute and / or decisions to make. In some embodiments, processor 158 may be configured to communicate with shared memory 206 (e.g., directly or via scheduler 190a, which receives data from shared memory 206 via path 208). Processor 158 can be configured to retrieve information from shared memory 206 to make a decision. The instructions executed by processor 158 in response to information from computer vision processor 102 can vary depending on the design criteria of a particular implementation.
[0128] refer to Figure 5 The diagram illustrates the preprocessing of video frames using partial block summation performed by a neural network implemented by a processor. Visualization 300 is shown. Visualization 300 can represent an example of operations performed using a neural network implemented by processor 102. Visualization 300 can illustrate partial block summation and averaging operations performed using convolution techniques natively implemented by edge device 100 to generate convolution results.
[0129] The operations shown in the visualization 300 for partial block summation (e.g., area summation) and averaging operations can be performed using the CNN module 190b. Visualization 300 can represent multiple operations for performing preprocessing of video frames and / or reference frames. In one example, the partial block summation and averaging operations can be implemented by one or more of hardware engines 204a-204n. The partial block summation and averaging operations can be implemented using additional components (not shown).
[0130] CNN module 190b can receive a signal VIDEO. The signal VIDEO may include a single-channel image with a structured light pattern (e.g., a newly incoming video frame captured by camera system 100). In one example, the signal VIDEO may include a reference image (e.g., preprocessed offline). The reference image can be used as a baseline for depth data, which can be computed using camera system 100 under known conditions. For example, operations performed in visualization 300 can be implemented online (e.g., preprocessing performed on a source image captured in real time) or offline (e.g., preprocessing performed on a reference image). The preprocessing operations shown in visualization 300 can provide some of the operations for generating data for other (e.g., upcoming) operations performed by processor 102 (e.g., generating depth maps and / or disparity maps).
[0131] CNN module 190b is illustrated as blocks 304, 306, and / or 308. Blocks 304-308 can illustrate various inputs, operations, and / or outputs performed by CNN module 190b to perform partial block summation and / or averaging operations. In one example, blocks 304-308 can represent computations performed by hardware engines 204a-204n. For example, hardware engines 204a-204n can be specifically tailored to perform various computations described by and / or associated with blocks 304-308. Block 304 can represent partial block summation of an input video image (e.g., an input image VIDEO, which may include source video frames or reference images). Block 306 can represent an averaging block. Block 308 can represent a convolution result.
[0132] Processor 102 can be configured to leverage the processing power of CNN module 190b to accelerate computations performed on partial block summations and / or generate convolution results. In the example, the partial block summation of input 304 can include a block size of 9x9. CNN module 190b can perform a partial block summation of 9x9 on the collected raw single-channel image (e.g., a source image or a reference image). CNN module 190b can then divide the partial block summation 304 by the average block 306. The average block 306 can have a block size of 9x9 and all values can be 81. The result of dividing the partial block summation 304 by the average block 306 can provide the average value of the block image. Instead of performing the averaging operation after the summation, CNN module 190b can accelerate computation using a 9x9 ordinary convolution (which uses hardware resources 204a-204n). The weights of the convolution can be 9x9, and each value in the label can be 81 for the convolution operation. The stride of a convolution can be 1 (for example, the kernel data is 1 / 81 to get a convolution result of 308).
[0133] The result of partial block summation 304 divided by average block 306 can be a convolution result 308. The convolution result 308 can be a 9x9 value. CNN module 190b can generate an output signal (e.g., CONVRES). The signal CONVRES can include the convolution result 308. Although visualization 300 is shown as having the signal CONVRES as output, the convolution result 308 can be used internally by CNN module 190b. The convolution result 308 can be generated in response to the hardware of CNN module 190b, which can perform partial block summation and averaging on video frames.
[0134] refer to Figure 6 A diagram illustrating the determination of offset values performed by a neural network implemented by a processor is shown. Visualization 350 is shown. Visualization 350 can represent an example of an operation performed using a neural network implemented by processor 102. Visualization 350 can illustrate the generation of offset values.
[0135] The CNN module 190b can be used to perform the operations shown in visualization 350 for generating offset values. Visualization 350 can represent multiple operations for performing preprocessing of video frames and / or reference frames. The preprocessing operations shown in visualization 350 can provide some of the operations for generating data for other (e.g., upcoming) operations performed by processor 102 (e.g., generating depth maps and / or disparity maps). Additional components (not shown) can be used to determine the offset values. CNN module 190b can receive a signal CONVRES and a signal (e.g., IDEAL). The signal IDEAL can include an ideal speckle image. Visualization 350 can represent neural network operations performed by processor 102 after the generation of convolution result 308, such as... Figure 5 As shown in the related diagram.
[0136] CNN module 190b is shown as illustrating convolution results 308, block 352, and / or block 354. Blocks 308 and / or blocks 352-354 can illustrate various inputs, operations, and / or outputs performed by CNN module 190b to generate offset values. In one example, blocks 308 and / or blocks 352-354 can represent computations performed by hardware engines 204a-204n. For example, hardware engines 204a-204n can be specifically tailored to perform various computations described and / or associated with blocks 308 and / or blocks 352-354. Block 352 can represent an ideal (e.g., reference) speckle image (e.g., the signal IDEAL). Block 354 can represent adaptive offset values.
[0137] After obtaining the convolution result 308, the CNN module 190b can be configured to compare each result (after convolution) with an ideal speckle value 352. The ideal speckle value can be determined based on the maximum ideal projection distance of the structured light projector 106. In the example, the magnitude of the ideal speckle value can be obtained by projecting the structured light pattern SLP onto a white wall at the maximum ideal distance (e.g., the furthest distance at which the structured light SLP is visible when generated by the structured light projector 106 and / or the ideal parameters of the structured light source 186). The signal IDEAL can include a video frame generated using the ideal speckle value. The signal IDEAL can be used as a reference image including the ideal speckle value, which can be compared with a single-channel source video frame or a reference video frame (e.g., the signal VIDEO).
[0138] The CNN module 190b can be configured to calculate the difference between the entire speckle pattern (e.g., the speckle pattern captured in the convolution result 308 generated from the signal VIDEO) and each corresponding value of the ideal speckle pattern image 352. The CNN module 190b can construct a histogram to calculate the difference results. The difference at the point with the most counts in the histogram can be used as the offset value 354.
[0139] The offset value 354 can be the output of the CNN module 190b. The CNN module 190b can be configured to generate an output signal (e.g., OFFSET). The signal OFFSET can include the offset value 354. Although visualization 350 is shown as having the signal OFFSET as output, the offset value 354 can be used internally by the CNN module 190b. The offset value 354 can be generated by the CNN module 190b comparing the convolution result 308 with the ideal speckle value 352.
[0140] by and Figure 5 The visualization of 300 and related features is shown in association. Figure 6 The operations illustrated in the visualization 350 may include calculations performed using hardware acceleration engines 204a-204n to determine offset values 354 based on the source image and the reference image. Both the source image and the ground reference image can be used in the visualization 350 to determine the offset values. Both the source image and the ground reference image can be used to determine the convolution result 308 for comparison with a reference image 352 having ideal speckle values (e.g., signal IDEAL). The source image can be compared to the ideal speckle values at a different time than the comparison between the ground reference image and the reference image with ideal speckle values.
[0141] In one example, the operations shown by visualizations 300 and 350 can be performed outside the system (e.g., offline). In another example, the operations shown by visualizations 300 and 350 can be performed when the system is online and / or in real time within the system. For example, the source image can be preprocessed to determine the offset value 354 in real time when the system is online (e.g., real-time operation of camera system 100). In another example, a reference image can be preprocessed to determine the offset value 354 when the system is offline (e.g., during offline training). Whether an operation is performed online or offline may depend on whether the scene used is determined (e.g., features known in advance and / or distances to various objects to be used as a real reference). Once the system and scene are determined, the adaptive offset phase 354 can also be determined. The CNN module 190b can then directly use a 9x9 convolutional kernel for additional operations (which will be combined with...). Figure 7 (Describe in relation to each other).
[0142] Reference Figure 7 The diagram illustrates how a combination of video frames and offset values, performed by a neural network implemented by a processor, determines an adaptive offset result. A visualization 400 is shown. Visualization 400 can represent an example of operations performed using a neural network implemented by processor 102. Visualization 400 can illustrate the generation of adaptive results. Visualization 400 can represent multiple operations for performing preprocessing on video frames and / or reference frames. The preprocessing operations shown in visualization 400 can provide some of the operations to generate data for other (e.g., upcoming) operations performed by processor 102 (e.g., generating depth maps and / or disparity maps).
[0143] The operations shown in visualization 400 for generating adaptive results can be performed using CNN module 190b, block (or circuit) 402, and / or block (or circuit) 404. Block 402 can implement a binarization module. Block 404 can implement a quadruple domain module. In one example, binarization module 402 and / or quadruple domain module 404 can be implemented by one or more of hardware engines 204a-204n. In another example, binarization module 402 and / or quadruple domain module 404 can be implemented as part of the video processing pipeline of processor 102. Additional components (not shown) can be used to determine the adaptive results. CNN module 190b can receive signal VIDEO and signal OFFSET. For example, visualization 400 can represent the neural network operations performed by processor 102 after offset value 354 has been generated, such as with... Figure 6 As shown in connection.
[0144] In some embodiments, the operations performed in visualization 400 (e.g., adding an offset value of 354) can be performed on the source image. In some embodiments, the operations performed in visualization 400 can be performed on a real reference image. In some embodiments, the adaptive offset value of 354 may not be added to the real reference image. For example, the real reference image may have been prepared offline to determine the binarization result for the real reference image.
[0145] CNN module 190b is illustrated as averaging 306, offset (or bias) 354, block 410, and / or block 412. Blocks 306, 354, and / or blocks 410-412 can illustrate various inputs, operations, and / or outputs performed by CNN module 190b to generate adaptive results. In one example, blocks 306, 354, and / or blocks 410-412 can represent computations performed by hardware engines 204a-204n. For example, hardware engines 204a-204n can be specifically tailored to perform various computations described and / or associated with blocks 306, 354, and / or blocks 410-412. Block 410 can represent a source image (e.g., a signal video). In an example, block 410 can include the sum of blocks (e.g., the sum of areas) of source video data or reference video data (e.g., a 9x9 block size). Block 412 can represent the adaptive result.
[0146] After obtaining the offset value of 354, the CNN module 190b can be configured to determine the adaptive result 412. The CNN module 190b can implement 9x9 convolutional kernels to determine the adaptive result 412. The weights of the convolution can be 9x9, and the values inside can all be 81. The convolution operation can be performed directly to obtain the adaptive result 412.
[0147] The convolution operation performed by hardware engines 204a-204n can be configured to: determine an adaptive offset of 354 (e.g., as in combination with...) Figure 6 Following the visualization shown (350), an adaptive offset value 354 is added to the source image 410 (or the real reference image). The average of the source image 410 can be determined by dividing the source image 410 by the average block 306 (e.g., a block size of 9x9, and all values can be 81). The offset value 354 can be added to the average. Hardware-accelerated convolution operations can be used to add the offset value 354. The adaptive offset phase can be included in the offset value 354. In the example, the adaptive result 412 can include the average of a 9x9 block of a single-channel image plus the offset value 354.
[0148] The adaptive result 412 can be the output of CNN module 190b. CNN module 190b can be configured to generate an output signal (e.g., ADVRES). The signal ADVRES can include the adaptive result 412. Although visualization 400 is shown as having the signal ADVRES as output, the adaptive result 412 can be used internally by CNN module 190b. CNN module 190b can be configured to generate the adaptive result 412 in response to performing a convolution operation to add the offset value 354 to video frames (e.g., the sum of video frames 410).
[0149] The signal ADVRES can be presented to the binarization module 402. The binarization module 402 can be configured to receive the signal ADVRES and a single-channel input image from the signal VIDEO. The binarization module 402 can be configured to perform a comparison between the adaptive result 412 from the signal ADVRES and the single-channel input image from the signal VIDEO. In some embodiments, the binarization module 402 can be configured to perform a comparison between the adaptive result 412 generated from the real reference image and the real reference image.
[0150] A comparison can be performed between the adaptive result 412 performed by binarization module 402 and the source image in the signal VIDEO to determine the binarization expression for the source image. Similarly, a comparison can be performed between the adaptive result 412 of the real reference image performed by binarization module 402 and the real reference image in the signal VIDEO to determine the binarization expression for the real reference image. The comparison can be performed by analyzing corresponding points from the adaptive result 412 and the video data. For example, if the comparison result is greater than or equal to 1, the output for the binarization expression can be 0; and if the comparison result is less than 1, the output for the binarization expression can be 1. The analysis of the comparison can form the binarized result.
[0151] In the example shown, binarization module 402 can generate a binarized result for the input image. A similar operation can be performed by binarization module 402 to binarize a real reference image (e.g., to form a binarized result of the reference image). Typically, the real reference image may have been prepared offline to enable the generation of a binarized expression for the real reference image. An adaptive offset value 354 can be determined online or offline. Using the adaptive offset value 354, an adaptive result 412 can be determined. Then, using the adaptive result 412, a binarized expression can be determined for the source image in the signal VIDEO.
[0152] Binarization module 402 can generate a signal (e.g., BINVID). The signal BINVID can include a binarization expression for the source image or a binarization result for the real reference image. The signal BINVID can be presented to quadruple domain module 404.
[0153] The quadruple domain module 404 can be configured to remove error points from the binarized result of the source image BINVID (or the binarized result of the real reference image). The quadruple domain module 404 can also be configured to implement the quadruple domain method (which will be combined with...). Figure 8 (Described in association). Although a quadruple domain method is shown, in some embodiments, the quadruple domain module 404 may be configured to perform point separation using a four-connected component labeling operation. The quadruple domain module 404 may generate a signal (e.g., ERR). The signal ERR may include a binarized result of the source image with error points removed. Similarly, the signal ERR may include a binarized result of the true reference image with error points removed.
[0154] Error point removal may include fitting error points to values of 0 instead of 1. Preprocessing performed by the CNN module 190b after the signal ERR is generated can be completed (e.g., binary data for the source image and binary data for the ground reference image may have been determined). The processor 102 can be configured to use the binary data for the source image and the binary data for the ground reference image for subsequent operations. In one example, the binary data for the source image and the binary data for the ground reference image (e.g., the signal ERR) can be used as input to a matching operation. The matching operation can be used to generate a depth map and a disparity map. This can be combined with... Figures 9-14 The details of generating depth maps and disparity maps in response to binary data determined by preprocessing performed by CNN module 190b are described in connection with this description.
[0155] refer to Figure 8 The diagram illustrates the removal of error points using a quadruple-domain method performed by a neural network. A visualization of error point 450 is shown. The visualization of error point 450 can be represented using a quadruple-domain method with... Figure 7 The type of error points removed by the quadruple-domain method performed by the quadruple-domain module 404 is shown in association.
[0156] After obtaining the binarized result (e.g., for a source image or a ground reference image), the quaddomain module 404 can receive a signal BINVID. The signal BINVID may include a binarized expression of the source image or the ground reference image, which may include a result of the same size as the original image (e.g., the signal VIDEO). The quaddomain module 404 can be configured to remove various error points. In one example, error points may include isolated points. In another example, error points may include connection points. In yet another example, error points may include glitch points. The type of error points removed can vary depending on the design criteria of a particular implementation.
[0157] Visualization 450 can include example error points 452-456. Error points 452-456 can affect the final accuracy of the depth image. Error points 452-456 can be eliminated using the quadruple domain method. Error point 452 can represent an isolated point. Error point 454 can represent a connected point. Error point 456 can represent a spurious point.
[0158] Isolated point 452 may include point 460. Point 460 may not be near any other points. The amount of space between point 460 and other points that is considered an isolated point can vary depending on the design criteria of a particular implementation. Isolated point 460 may be an error that can be removed. In the example, isolated point 460 may be assigned a value of 0 in response to detection by the quadruple domain method.
[0159] Connection point 454 may include a set of points 462a-462n, connection point 464, and a set of points 466a-466n. Connection point 464 may connect to both the set of points 462a-462n and the set of points 466a-466n. However, connection point 464 may not be part of a set of points. Connection point 464 may be an error that can be removed. In the example, in response to being detected by the quadruple field method, connection point 464 may be assigned the value 2.
[0160] Spurt point 456 may include a set of points 468a-468n and spur point 470. Spurt point 470 may be adjacent to one of the points in the set 468a-468n, but is not part of the set. Spurt point 470 may be an error that can be removed. In the example, in response to being detected by the quadruple-domain method, spur point 470 may be assigned the value 1.
[0161] The signal ERR may include binary data (e.g., a preprocessed result) of a source image or a true reference image with error points removed. CNN module 190b can generate the preprocessed result. The preprocessed result can be used for upcoming operations performed by processor 102 and / or CNN module 190. In one example, processor 102 can perform logical operations on the preprocessed result. For example, processor 102 can perform an XOR operation between binary data for the source image and binary data for the reference image. The preprocessed result can be used as input to generate a disparity map and / or a depth map for 3D reconstruction.
[0162] refer to Figure 9 A diagram illustrating an example speckle image is shown. A dashed box 500 is shown. The dashed box 500 can represent a video frame. The video frame 500 can be an example input video frame captured by the capture device 104. In the example, the video frame 500 can represent a signal VIDEO generated by the capture device 104 and presented to the processor 102. The video frame 500 can represent an example of a speckle image. The video frame 500 can capture a structured light pattern SLP generated by the structured light projector 106.
[0163] Video frame 500 may include a wall 504 and a box 506. Box 506 may have a front 508, sides 510, and a top 512. The front 508 of box 506 may generally face the direction of the capture device 104 that captures video frame 500. For example, the front 508 may be the side of box 506 closest to the capture device 104. In the example shown, box 506 may not directly face the capture device 104. For example, the sides 510 and top 512 may be farther from the capture device 104 than the distance from the front 508 to the capture device 104. The white wall 504 may be located further away from the capture device 104 than the front 508, sides 510, and / or top 512 of box 506.
[0164] Typically, a white wall is used as the evaluation scene to assess the accuracy of a depth sensing system. Wall 504 can be a white wall used for the evaluation scene. Box 506 can be located in front of white wall 504. The accuracy of depth and / or disparity of white wall 504 and box 506 can be determined. In this example, an accurate depth map and / or disparity map may be more accurate in distinguishing the edges of white wall 504 and / or box 506 than a less accurate one.
[0165] The structured light projector 106 can be configured to project a structured light pattern SLP onto a white wall 504 and a box 506. In one example, the structured light pattern SLP can be implemented as a speckle pattern comprising dots of a predetermined size. Generally, when the structured light pattern SLP is projected onto an object closer to the lens 160 of the capture device 104, the dots of the structured light pattern SLP can have a larger size than the dots of the structured light pattern SLP projected onto an object farther from the lens 160 of the capture device 104. For clarity and illustration purposes, to show the difference in speckle patterns on the white wall 504 and the box 506, the speckle pattern of the dots of the structured light pattern SLP is shown only as projected onto the white wall 504 and the box 506. Typically, the speckle pattern of the dots of the structured light pattern SLP can be projected onto the entire video frame 500 (e.g., on the floor / ground, on the ceiling, on any surface next to the white wall 504, etc.).
[0166] The speckle pattern of the structured light pattern SLP is shown as pattern 514 on white wall 504, pattern 516 on the front 508 of box 506, pattern 518 on the top 512 of box 506, and pattern 520 on the side 510 of box 506. The dots in pattern 514 may include small dots. The dots in pattern 516 may include large dots. The dots in pattern 518 and pattern 520 may include medium-sized dots.
[0167] Since the front 508 of box 506 can be the surface closest to lens 160, pattern 516 can include the largest dot in video frame 500. The sides 510 and top 512 of box 506 can be further away from lens 160 than the front 508, but closer to lens 160 than the white wall 504. The dots of pattern 518 on top 512 and pattern 520 on side 510 can be smaller than the dots of pattern 516 on front 508. Which dot is larger, pattern 518 or pattern 520, can depend on which surface (e.g., side 510 or top 512) is closer to lens 160. Since white wall 504 can be the surface furthest from lens 160, pattern 514 can include the smallest dot in video frame 500.
[0168] The dimensions of the points in patterns 514-520 of the structured light pattern SLP in video frame 500 can be used by processor 102 to determine the distance and / or depth of various objects captured in the video frame. The dimensions of the points in patterns 514-520 enable processor 102 to generate a disparity map. Depth maps and / or disparity maps can be generated in response to video frames captured using monocular lens 160 and analysis performed on speckle patterns 514-520.
[0169] In one example, processor 102 can be configured to perform adaptive binarization on a 480x272 IR channel map. By implementing the video pipeline of processor 102 using neural network module 190b (e.g., using adaptive offset terms and a four-connection domain approach), adaptive binarization can be performed in approximately 74µs (e.g., where Net_id:0, Dags:1 / 1, vp_ticks:911). On a general-purpose processor (e.g., an ARM processor), the generation time for adaptive binarization (e.g., without adaptive offset terms) can be approximately 2ms. Compared to using a general-purpose processor, the use of convolutions performed by hardware modules 204a-204n for binarization and / or the use of four-connection domains for isolated points, connected points, and glitch points can provide a significant speed advantage.
[0170] refer to Figure 10 A diagram illustrating the disparity map generated from a speckle image after binarization without adding an adaptive offset value is shown. Dashed box 550 is shown. Dashed box 550 can represent a disparity map. Disparity map 550 can be an example output (e.g., the signal DIMAGES) generated by processor 102. In the example shown, disparity map 550 can represent a disparity map generated without using adaptive offset value 354. For example, processor 102 can be configured to: use with... Figures 5-7 The associated neural network operations described are 300-400 (e.g., high accuracy, efficient operation, low power consumption) or do not use [the following]. Figures 5-7 The neural network operations described in association are 300-400 (e.g., less accurate, less efficient, more power-consuming operations) to generate signals DIMAGES.
[0171] White wall 504 and box 506 are shown in disparity map 550. White wall 504 and box 506 are shown without speckle patterns 514-520. For example, binarization can extract speckle patterns 514-520 to enable processor 102 to perform disparity calculations.
[0172] Box 506 is shown with edge 552. Edge 552 may be inaccurate. In the example shown, edge 552 is shown as largely blurred to illustrate the inaccuracy. This inaccuracy may occur due to the lack of use of adaptive offset terms and / or other neural network operations (e.g., other methods may produce inaccuracies when using...). Figures 5-7 (These inaccuracies may be corrected when describing neural network operations 300-400 in association), and inaccuracies at edge 552 may exist in disparity map 550.
[0173] Disparity error can be calculated for each pixel in the disparity map. In the example, a pixel with disparity error can be considered to have disparity error if the disparity for that pixel is greater than 1 compared to the true disparity. The true disparity can be determined based on a true reference image. In the example, processor 102 can be configured to perform disparity error calculation.
[0174] The processor 102 can also be configured to calculate the proportion of disparity error for each pixel. The proportion of disparity error for each pixel can be determined by summing the total disparity error pixels and dividing the sum by the image size. The proportion of disparity error for each pixel can provide a measure of disparity quality for the disparity map 550. In the example shown, the proportion of disparity error for each pixel can be approximately 5.7%. Disparity error can cause inaccuracies in the edges 552 of the disparity map 550.
[0175] refer to Figure 11 A diagram illustrating the depth map generated from a speckle image after binarization without adding an adaptive offset value is shown. Dashed box 560 is shown. Dashed box 560 can represent a depth map. Depth map 560 can be an example output (e.g., the signal DIMAGES) generated by processor 102. In the example shown, depth map 560 can represent a depth map generated without using adaptive offset value 354. For example, processor 102 can be configured to: use with... Figures 5-7 The associated neural network operations described are 300-400 (e.g., high accuracy, efficient operation, low power consumption) or do not use [the following]. Figures 5-7 The neural network operations described in association are 300-400 (e.g., less accurate, less efficient, more power-consuming operations) to generate signals DIMAGES.
[0176] White wall 504 and box 506 are shown in depth map 560. White wall 504 and box 506 are shown without speckle patterns 514-520. For example, binarization can extract speckle patterns 514-520 to enable processor 102 to perform depth calculations.
[0177] Box 506 is shown as having edge 562. Edge 562 may be inaccurate. In the example shown, edge 562 is shown as largely blurred to illustrate the inaccuracy. This inaccuracy may occur due to the lack of use of adaptive offset terms and / or other neural network operations (e.g., other methods may produce inaccuracies when using...). Figures 5-7 (These inaccuracies may be corrected when describing neural network operations 300-400 in association), and inaccuracies at edge 562 may exist in depth map 560.
[0178] Z-accuracy can be calculated to evaluate the accuracy of depth data in a depth image. Z-accuracy measures how close a reported depth value in a depth image is to the true value. The true value can be determined based on a real reference image. In this example, processor 102 can be configured to perform Z-accuracy calculations.
[0179] A fill rate can be calculated to measure the proportion of a depth image containing valid pixels. Valid pixels can be pixels with non-zero depth values. The fill rate measure can be independent of the accuracy of the depth data. Processor 102 can be configured to perform fill rate calculations.
[0180] In the example shown, the Z-accuracy of depth image 560 can be approximately 94.3%. In the example shown, the fill rate of depth image 560 can be approximately 97.3%. Low measured Z-accuracy values and / or low fill rates may result in inaccurate edges 562 in depth map 560.
[0181] refer to Figure 12 The figure illustrates the binarization result generated from the speckle image in response to adding adaptive offset values and removing error points. Binarization result 570 is shown. It can be responded to with... Figures 5-7 The neural network operations described in association 300-400 generate a binarized result 570.
[0182] The speckle pattern 514-520 is shown in the binarized result 570. The processor 102 can be configured to implement neural network operations 300-400, the binarization module 402, and / or the quadruple domain module 404 to generate the binarized result 570. The binarized result 570 can be the signal ERR generated in response to the removal of error points from the signal BINVID after an adaptive offset value 354 has been added to the signal VIDEO (e.g., the source image or a true reference image) to generate the adaptive result ADVRES.
[0183] The binarization result 570 can extract the speckle patterns 514-520 from the captured image. In the example shown, the speckle patterns 514-520 are shown as having the same... Figure 9 The dots shown in the associated captured video frame 500 have the same dot size. However, objects (e.g., the front 508, side 510, and top 512 of white wall 504 and box 506) may not be directly visible. Processor 102 can be configured to determine the position, size, and / or depth of the front 508, side 510, and top 512 of white wall 504 and box 506 by inference based on the size of the dots in speckle patterns 514-520.
[0184] Binarization result 570 can represent the binarization result generated in response to performing neural network operations 300-400. Binarization results can also be generated without using neural network operations 300-400. For example, disparity image 550 (and...) Figure 10 (shown in association) and depth image 560 (with) Figure 11 (As shown in association) can be generated based on the binarization result produced without neural network operations 300-400. The binarization result 570 generated using neural network operations 300-400 can provide higher accuracy and / or quality than when performing binarization without neural network operations 300-400.
[0185] A single-point ratio can be calculated to evaluate the binary quality of the binarization result. The single-point ratio can include a calculation of dividing the number of single points by the total number of points. The total number of points can include the sum of the number of single points, isolated points 452, connected points 454, and spur points 456 in an image after preprocessing the binary. The processor 102 can be configured to perform the single-point ratio calculation. In the example shown, the binarization result 570 can have a single-point ratio of 95.4%. Without performing neural network operations 300-400, the binarization result used to generate the disparity image 550 and the depth image 560 can have a single-point ratio of 90.2%. By implementing neural network operations 300-400 to add an adaptive bias 354, the binarization result 570 can have a 5.2% improvement in the single-point ratio metric.
[0186] refer to Figure 13 A diagram illustrating the disparity map generated from a speckle image after binarization in response to the addition of an adaptive offset value and the removal of error points is shown. Dashed box 580 is shown. Dashed box 580 can represent a disparity map. Disparity map 580 can be an example output (e.g., the signal DIMAGES) generated by processor 102. In the example shown, disparity map 580 can represent a disparity map generated using an adaptive offset value 354. For example, processor 102 can be configured to: use with… Figures 5-7 The neural network operations 300-400 described in connection are used to generate signals DIMAGES.
[0187] The disparity map 580 can represent the result generated by the processor 102 after the adaptive offset 354 has been added to the source or reference image, after the binarization result 570 has been generated, and after error points 452-456 have been removed. In the example, binary data (e.g., binary results in the signal ERR) can be generated for the source and reference images during preprocessing. The preprocessing result can be used to generate the disparity map 580. In the example, one or more of the hardware modules 190a-190n implemented by the processor 102 can be configured to perform a matching operation using the preprocessing result as input to generate the disparity map 580.
[0188] White wall 504 and box 506 are shown in disparity map 580. White wall 504 and box 506 are shown without speckle patterns 514-520. For example, the binarization result 570 can extract speckle patterns 514-520 so that processor 102 can perform disparity calculation.
[0189] Box 506 is shown as having edge 582. Edge 582 can be clearly represented. In the example shown, edge 582 is shown differently to illustrate the accuracy of the resulting disparity image. This is due to the use of adaptive offset terms and / or other neural network operations (e.g., with...). Figures 5-7 The associated neural network operations (300-400) are described, so the accuracy of edge 582 can exist in disparity map 580.
[0190] In the example shown, for disparity map 580, the disparity error percentage per pixel could be approximately 1.8%. In contrast to... Figure 10 In the associated disparity image 550, the proportion of disparity error for each pixel can be approximately 5.7%. By implementing neural network operations 300-400, processor 102 can generate a disparity map 580 with a 3.9% reduction in error disparity after adding an adaptive bias 354. This reduction in error disparity makes it possible to define distinct edges 582 for boxes 506 in the disparity image 580.
[0191] refer to Figure 14 A diagram illustrating the depth map generated from a speckle image after binarization in response to the addition of an adaptive offset value and the removal of error points is shown. Dashed box 580 is shown. Dashed box 580 can represent a depth map. Depth map 600 can be an example output (e.g., the signal DIMAGES) generated by processor 102. In the example shown, depth map 600 can represent a depth map generated using an adaptive offset value 354. For example, processor 102 can be configured to: use with... Figures 5-7 The neural network operations 300-400 described in connection are used to generate signals DIMAGES.
[0192] Depth map 600 can represent the result generated by processor 102 after adaptive offset 354 has been added to the source or reference image, after binarization result 570 has been generated, and after error points 452-456 have been removed. In the example, binary data (e.g., binary results in the signal ERR) can be generated for the source and reference images during preprocessing. The preprocessing result can be used to generate depth map 600. In the example, one or more of the hardware modules 190a-190n implemented by processor 102 can be configured to perform a matching operation using the preprocessing result as input to generate depth map 600.
[0193] White wall 504 and box 506 are shown in depth map 600. White wall 504 and box 506 are shown without speckle patterns 514-520. For example, the binarization result 570 can extract speckle patterns 514-520 to enable processor 102 to perform depth calculations.
[0194] Box 506 is shown with edge 602. Edge 602 can be clearly represented. In the example shown, edge 602 is shown differently to illustrate the accuracy of the resulting depth image. This is due to the use of adaptive offset terms and / or other neural network operations (e.g., with...). Figures 5-7 The neural network operations described in association (300-400) are such that the accuracy of edge 602 can be found in depth map 600.
[0195] In the example shown, the Z-accuracy of the depth map 600 can be approximately 96.4%, and the fill rate can be approximately 98.4%. (In contrast to...) Figure 11 In the depth image 560 shown in association, the Z-accuracy can be approximately 94.3% and the fill rate can be approximately 97.3%. By implementing neural network operations 300-400, the processor 102 can generate a depth map 600 with a 2.1% improvement in Z-accuracy and a 1.1% improvement in fill rate after adding an adaptive bias 354. The improvement in Z-accuracy and fill rate makes it possible for boxes 506 in the depth map 600 to define different edges 602.
[0196] refer to Figure 15The diagram illustrates method (or process) 620. Method 620 preprocesses video frames by adding adaptive offset terms to a locally adaptive binarized expression using convolution techniques. Method 620 generally includes steps (or states) 622, 624, 626, 628, decision steps (or states) 630, 632, 634, 636, 638, 640, 642, 644, 646, and 648.
[0197] Step 622 can initiate method 620. In step 624, the structured light projector 106 can generate a structured light pattern SLP. In an example, the SLP source 186 can generate a signal SLP including a speckle pattern that can be projected onto the environment near device 100. Next, in step 626, the processor 102 can receive pixel data from a monocular camera. In an example, the capture device 104 can implement a monocular camera. The monocular camera 104 can receive a signal LIN including light via lens 160. The RGB-IR sensor 180 can convert the input light into pixel data and / or video frames. The pixel data can include information about the environment near device 100 and capture the structured light pattern generated by the structured light projector 106. The monocular camera 104 can present a signal VIDEO to the processor 102. In step 628, the processor 102 can be configured to process the pixel data arranged as video frames. In one example, the processor 102 can convert the pixel data into video frames. In another example, capture device 104 can convert pixel data into video frames and present the video frames to processor 102. The video frames may include a single-channel source image or a single-channel reference image. Next, method 620 can move to decision step 630.
[0198] In decision step 630, processor 102 can determine whether to utilize CNN module 190b to generate disparity and depth maps. If CNN module 190b is not used, method 620 can proceed to step 632. In step 632, processor 102 can perform various 3D reconstruction calculations without relying on hardware acceleration provided by CNN module 190b (e.g., without relying on hardware engines 204a-204n and / or without adding a slower computation path of adaptive offset value 354, which could result in inaccuracies 552 shown in disparity map 550 and / or inaccuracies 562 shown in depth map 560). Next, method 620 can proceed to step 648. If CNN module 190b was used in decision step 630, method 620 can proceed to step 634.
[0199] In step 634, CNN module 190b can perform partial block summation and averaging on the video frame. In the example, partial block summation and averaging enable the generation of a convolution result 308 for either the source or reference image. Next, in step 636, CNN module 190b can compare the convolution result 308 with an ideal speckle value 352 to generate an offset value 354. In step 638, CNN module 190b can perform a convolution operation to add the offset value 354 to the video frame (e.g., the source or reference image) to generate an adaptive result 412. Next, in step 640, CNN module 190b can compare the video frame (e.g., if the adaptive result 412 comes from the source image, it is the source image in the signal VIDEO; or if the adaptive result 412 comes from the real reference image, it is the reference image in the signal VIDEO) with the adaptive result 412 to generate a binarized result (e.g., the signal BINVID). In step 642, the CNN module 190b can be configured to remove error points (e.g., isolated points 452, connected points 454, and / or glitch points 456) from the binarized result BINVID. The result of removing error points can be binary data (e.g., the signal ERR, which can correspond to source binary data when the video frame signal VIDEO includes the source image, and to real binary data when the video frame signal VIDEO includes the real reference image). Next, method 620 can proceed to step 644.
[0200] In step 644, CNN module 190b can perform preprocessing on the video frame (e.g., a source image or a real reference image). Next, in step 646, processor 102 can generate a disparity map 580 and a depth map 600. The disparity map 580 and depth map 600 can be generated from binary data of the source and reference images (e.g., inputs for the matching operation method). Next, method 620 can move to step 648. Step 648 can end method 620.
[0201] refer to Figure 16 The diagram illustrates method (or process) 680. Method 680 can use a convolutional neural network to perform partial block summation and averaging. Method 680 generally includes steps (or states) 682, 684, decision steps (or states) 686, 688, 690, 692, 694, 696, 698, and 700.
[0202] Step 682 can initiate method 680. In step 684, processor 102 can generate video frames capturing the structured light pattern SLP. Next, method 680 can move to decision step 686. In decision step 686, processor 102 can determine whether the captured scene has been previously determined. If distance information and / or 3D information is known in advance, the scene can be determined in advance. In this example, the scene can be determined in advance for a real reference video frame. If the scene has been previously determined, method 680 can move to step 688. In step 688, processor 102 can determine the offset value 354 in real time (e.g., analysis of the source image). Next, method 680 can move to step 692. In decision step 686, if the scene has been previously determined, method 680 can move to step 690. In step 690, processor 102 can determine the offset value 354 offline (e.g., analysis of the reference image). Next, method 680 can move to step 692.
[0203] In step 692, the hardware-implemented CNN module 190b can receive a single-channel video frame (e.g., a source image or a real reference image from a monocular camera 104). Next, in step 694, the CNN module 190b can generate a partial block sum 304. The partial block sum 304 can be generated from the source image or the reference image. In step 694, the CNN module 190b (e.g., using one or more hardware engines 204a-204n) can divide the partial block sum 304 by 81 using a 9x9 averaging block 306. Next, in step 696, the CNN module 190b can generate a 9x9 convolution result 308 (e.g., the signal CONVRES). Next, method 680 can move to step 700. Step 700 can end method 680.
[0204] refer to Figure 17 The diagram illustrates method (or process) 720. Method 720 can determine an offset value. Method 720 generally includes steps (or states) 722, 724, 726, 728, 730, 732, a decision step (or state) 734, 736, and 738.
[0205] Step 722 can initiate method 720. Next, method 720 can move to steps 724 and 726, which can be executed in parallel or substantially in parallel. In step 724, CNN module 190b can receive (or determine) the convolution result 308 (e.g., based on an averaging operation) Figure 5(As shown in association). Next, method 720 can move to step 728. In step 726, CNN module 190b can receive the signal IDEAL, which includes the ideal speckle image pattern 352. Next, method 720 can move to step 728.
[0206] In step 728, CNN module 190b can calculate the difference between the convolution result 308 and the corresponding value of the ideal speckle image pattern 352. Next, in step 730, CNN module 190b can generate a histogram of the difference results between the convolution result 308 and the ideal speckle image pattern 352. In step 732, CNN module 190b can analyze the generated histogram. Next, method 720 can move to decision step 734.
[0207] In decision step 734, CNN module 190b can determine whether a difference value in the histogram of the most counted points has been found. If the most counted point in the histogram has not yet been found, method 720 can return to step 732 (e.g., continue generating and / or analyzing the histogram). If the most counted point in the histogram has been found, method 720 can move to step 736. In step 736, CNN module 190b can use the difference value of the most counted point in the histogram as an offset value 354 (e.g., signal OFFSET). Next, method 720 can move to step 738. Step 738 can end method 720.
[0208] refer to Figure 18 The diagram illustrates method (or process) 780. Method 780 can generate adaptive results and binarized results by adding offset values. Method 780 generally includes steps (or states) 782, 784, 786, 788, 790, 792, decision steps (or states) 794, 796, 798, decision steps (or states) 800, 802, and 804.
[0209] Step 782 can begin method 780. In step 784, the hardware of CNN module 190b can receive video frames from the signal channel (e.g., signal VIDEO) and offset value 354 (e.g., as shown in the image). Figure 6(As determined in association). Next, in step 786, CNN module 190b can perform partial block summation 410 on video frames (e.g., source images or real reference images) with 9x9 block sizes. In step 788, CNN module 190b can divide the partial block summation 410 by 81 using 9x9 averaging blocks 306. Next, method 780 can move to step 790.
[0210] In step 790, CNN module 190b may add offset value 354 to the average result determined in step 788 to generate adaptive result 412. Adaptive result 412 may be determined in response to convolution operation. Next, in step 792, binarization module 402 (e.g., one of hardware engines 204a-204n implemented for CNN module 190b) may compare adaptive result 412 (e.g., signal ADVRES) with source video frames (e.g., source image or real reference image in signal VIDEO). Next, method 780 may move to decision step 794.
[0211] In decision step 794, binarization module 402 determines whether the comparison result of adaptive result 412 and video frame (e.g., the comparison result of the corresponding value of adaptive result 412 and video frame) is greater than or equal to the value 1. If the result is greater than or equal to 1, method 780 can move to step 796. In step 796, binarization module 402 outputs the value of the corresponding 0 point of the binarization result. Next, method 780 can move to decision step 794. In decision step 794, if the result is not greater than or equal to 1 (e.g., less than 1), method 780 can move to step 798. In step 798, binarization module 402 outputs the value of the corresponding 1 point of the binarization result. Next, method 780 can move to decision step 800.
[0212] In decision step 800, CNN module 190b determines whether there are more values to compare between the adaptive result 412 and the video frame. While method 780 may describe the comparison as performed sequentially, CNN module 190b, hardware engines 204a-204n, and / or binarization module 402 can be configured to compare the corresponding values of the adaptive result 412 and the video frame in parallel computation or substantially parallel execution. If there are more values to compare, method 780 can return to step 792. If there are no more values to compare, method 780 can move to step 802.
[0213] In step 802, binarization module 402 may generate a binarization result (e.g., the signal BINVID). In some embodiments, the binarization result may be a binarization result for the source image (e.g., if the signal VIDEO includes the source image). In some embodiments, the binarization result may be a binarization result for a real reference image (e.g., if the signal VIDEO includes the real reference image). Next, method 780 may proceed to step 804. Step 804 may end method 780.
[0214] refer to Figure 19 The diagram illustrates method (or process) 820. Method 820 can remove error points to generate binary data. Method 820 typically includes steps (or states) 822, 824, 826, 828, decision step (or state) 830, 832, 834, 836, 838, 840, 842, 844, 846, 848, and 850.
[0215] Step 822 can initiate method 820. Next, in step 824, the quadruple domain module 404 (e.g., one of hardware engines 204a-204n implemented for CNN module 190b) can receive the binarization result (e.g., the signal BINVID). The binarization result can be for either the source image or a reference image. In step 826, the quadruple domain module 404 can perform a quadruple domain operation on the binarization result BINVID. Next, in step 828, the quadruple domain module 404 can detect isolated points 452, connected points 454, and spur points 456 in the binarization result BINVID. Next, method 820 can proceed to decision step 830.
[0216] In decision step 830, the quadruple domain module 404 can determine whether the detected error points 452-456 are isolated points 460. If error points 452-456 are isolated points 460, then method 820 can move to step 832. In step 832, the quadruple domain module 404 can remove error points with a value of 0. Next, method 820 can move to decision step 842. In decision step 830, if error points 452-456 are not isolated points 460, then method 820 can move to decision step 834.
[0217] In decision step 834, the quadruple domain module 404 can determine whether the detected error points 452-456 are glitch points 470. If error points 452-456 are glitch points 470, then method 820 can proceed to step 836. In step 836, the quadruple domain module 404 can remove error points with a value of 1. Next, method 820 can proceed to decision step 842. In decision step 834, if error points 452-456 are not glitch points 470, then method 820 can proceed to decision step 838.
[0218] In decision step 838, the quadruple domain module 404 can determine whether the detected error points 452-456 are connection points 464. If error points 452-456 are connection points 464, then method 820 can proceed to step 840. In step 840, the quadruple domain module 404 can remove error points with a value of 2. Next, method 820 can proceed to decision step 842. In decision step 838, if error points 452-456 are not connection points 464, then method 820 can proceed to decision step 842.
[0219] In decision step 842, the quadruple domain module 404 can determine whether there are more error points 452-456. Although method 820 may present the detection of error points 452-456 as performed sequentially, CNN module 190b, hardware engines 204a-204n, and / or quadruple domain module 404 can be configured to analyze, detect, and / or remove error points 452-456 in parallel or substantially parallel operations. If more error points 452-456 exist, method 820 can return to decision step 830. If no more error points 452-456 exist, method 820 can move to step 844.
[0220] In step 844, the quadruple domain module 404 can determine that all error points 452-456 have been removed from the binarized result BINVID. Next, in step 846, the quadruple domain module 404 can generate a binarized result with error points removed (e.g., binary data in the signal ERR). In some embodiments, the binary data in the signal ERR can be source binary data determined from the source image. In some embodiments, the binary data in the signal ERR can be reference binary data determined from a real reference image. In step 848, preprocessing performed by the CNN module 190b can be completed. For example, the source binary data and / or reference binary data in the signal ERR can be used by other upcoming operations performed by the processor 102. The processor 102 can use the source binary data and / or reference binary data to generate a disparity map 580 and / or a depth map 600. Next, method 820 can move to step 850. Step 850 can end method 820.
[0221] One or more of the following can be used to achieve the effect of Figures 1-19 The functions performed by the diagram include: general-purpose processors, digital computers, microprocessors, microcontrollers, RISC (Reduced Instruction Set Computer) processors, CISC (Complex Instruction Set Computer) processors, SIMD (Single Instruction Multiple Data) processors, signal processors, central processing units (CPUs), arithmetic logic units (ALUs), video digital signal processors (VDSPs), and / or similar operating machines, programmed according to the instructions in this manual, as will be obvious to those skilled in the art. Appropriate software, firmware, codes, routines, instructions, opcodes, microcode, and / or program modules can be readily prepared by a skilled programmer based on the teachings of this disclosure, as will also be obvious to those skilled in the art. The software is typically implemented by one or more processors and executed from one or more media.
[0222] The present invention can also be implemented by preparing an ASIC (Application-Specific Integrated Circuit), a platform ASIC, an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), a CPLD (Complex Programmable Logic Device), a gate sea, an RFIC (Radio Frequency Integrated Circuit), an ASSP (Application-Specific Standard Product), one or more monolithic integrated circuits, one or more chips or dies arranged as flip-chip modules and / or multi-chip modules, or by interconnecting appropriate conventional component circuit networks as described herein, modifications of which will be apparent to those skilled in the art.
[0223] Therefore, the present invention may also include a computer product, which may be a storage medium or medium and / or a transmission medium or medium, including instructions that can be used to program a machine to execute one or more processes or methods according to the present invention. The execution of the instructions contained in the computer product and the operation of surrounding circuitry by the machine can convert input data into one or more files on the storage medium and / or one or more output signals representing physical objects or substances (e.g., audio and / or visual descriptions). The storage medium may include, but is not limited to, any type of disk, including floppy disks, hard disks, magnetic disks, optical disks, CD-ROMs, DVDs, and magneto-optical disks, as well as circuitry such as ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), UVPROM (Ultraviolet Erasable Programmable ROM), flash memory, magnetic cards, optical cards, and / or any type of medium suitable for storing electronic instructions.
[0224] The elements of this invention can form part or all of one or more devices, units, components, systems, machines, and / or apparatuses. These devices may include, but are not limited to, servers, workstations, storage array controllers, storage systems, personal computers, laptop computers, notebook computers, handheld computers, cloud servers, personal digital assistants, portable electronic devices, battery-powered devices, set-top boxes, encoders, decoders, transcoders, compressors, decompressors, preprocessors, post-processors, transmitters, receivers, transceivers, cryptographic circuits, cellular phones, digital cameras, positioning and / or navigation systems, medical devices, head-up displays, wireless devices, audio recording, audio storage and / or audio playback devices, video recording, video storage and / or video playback devices, gaming platforms, peripheral devices, and / or multi-chip modules. Those skilled in the art will understand that the elements of this invention can be implemented in other types of devices to meet specific application standards.
[0225] The terms “may” and “generally” are used herein in conjunction with “is” and verbs to convey the intention that the description is exemplary and considered broad enough to encompass both the specific examples presented in this disclosure and the alternative examples that may be derived based on this disclosure. The terms “may” and “generally” as used herein should not be construed as necessarily implying the desirability or possibility of omitting the corresponding element.
[0226] The designations “a” through “n” for various components, modules, and / or circuits, when used herein, disclose a single component, module, and / or circuit or multiple such components, modules, and / or circuits, wherein the designation “n” is used to indicate any particular integer. Each distinct component, module, and / or circuit having an instance (or event) designated as “a” through “n” may indicate that the distinct component, module, and / or circuit may have a matching number of instances or a different number of instances. An instance designated as “a” may represent the first of multiple instances, and an instance “n” may refer to the last of multiple instances without implying a specific number of instances.
[0227] Although the invention has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention.
Claims
1. An apparatus for preprocessing video frames, comprising: An interface configured to receive pixel data; A structured light projector configured to generate structured light patterns; as well as A processor configured to: (i) process the pixel data arranged as the video frame, (ii) perform operations using a convolutional neural network to determine a binarization result and an offset value, and (iii) generate a disparity map and a depth map in response to (a) the video frame, (b) the structured light pattern, (c) the binarization result, (d) the offset value, and (e) the removal of error points, wherein the convolutional neural network: (A) Perform partial block summation and averaging on the video frames to generate the convolution result. (B) The convolution result is compared with the ideal speckle value to determine the offset value. (C) In response to performing a convolution operation to add the offset value to the video frame, an adaptive result is generated. (D) Compare the video frame with the adaptive result to generate a binarized result for the video frame, and (E) Remove the error points from the binarization result.
2. The apparatus according to claim 1, wherein, The error points were removed using the quadruple field method.
3. The apparatus according to claim 2, wherein, The error points include at least one of the following: isolated points, connection points, and burr points.
4. The apparatus according to claim 3, wherein, The quadruple domain method is configured to remove (i) isolated points with a value of 0, (ii) spurious points with a value of 1, and (iii) connected points with a value of 2.
5. The apparatus according to claim 1, wherein, The offset value is configured to (i) separate the structured light pattern from the background image and (ii) reduce the number of error points.
6. The apparatus according to claim 1, wherein, The convolutional neural network is configured to remove the error points after generating the binarized result in order to reduce the probability of mismatches in the upcoming operation.
7. The apparatus according to claim 1, wherein, The convolutional neural network is configured to generate the binarized result so that convolution operations can be used to perform area summation and add offset operations.
8. The apparatus according to claim 1, wherein, The partial block summation is implemented with a block size of 9×9, so that 9×9 convolution can replace the averaging operation.
9. The apparatus according to claim 8, wherein, Each value in the 9×9 convolution is 81, and the stride of the 9×9 convolution is 1.
10. The apparatus according to claim 1, wherein, The ideal speckle value is determined in response to video data captured at the maximum ideal distance of the structured light projector projecting the structured light pattern onto a white wall.
11. The apparatus according to claim 1, wherein, The partial block summation and determination of the offset value are performed during offline training of the device.
12. The apparatus according to claim 1, wherein, The partial block summation and determination of the offset value are performed during the real-time operation of the device.
13. The apparatus according to claim 1, wherein, The offset value is determined in response to: (i) calculating the difference between each corresponding value of the structured light pattern captured in the video frame and the ideal speckle value, (ii) using a histogram to determine the difference between the corresponding value of the structured light pattern captured in the video frame and the ideal speckle value, and (iii) using the difference of the most counted points in the histogram as the offset value.
14. The apparatus according to claim 1, wherein, The video frame includes an image of a scene having the structured light pattern.
15. The apparatus according to claim 1, wherein, (i) The convolutional neural network is configured to generate the binarization result, wherein the error points are removed for the source image and the reference image; (ii) The processor is further configured to generate the binarization result in response to an XOR operation, combined with the result of removing the error points for the source image and the reference image.
16. The apparatus according to claim 1, wherein, The video frame includes a single-channel image captured by a monocular camera.
17. The apparatus according to claim 1, wherein, The convolutional neural network is configured to (i) output 0 for the binarization result when the comparison between the adaptive result and the video frame is greater than or equal to 1, and (ii) output 1 for the binarization result when the comparison between the adaptive result and the video frame is less than 1.
18. The apparatus according to claim 1, wherein, The binarized result generated by the convolutional neural network, after removing the error points, includes preprocessing results of the source and reference images for the upcoming operation to be performed by the processor.
19. The apparatus of claim 18, wherein (i) the upcoming operation includes generating the disparity map and the depth map in response to a matching operation, and (ii) the preprocessing results of the source image and the reference image include inputs for the matching operation.
20. The apparatus according to claim 1, wherein, The device is configured to use convolution techniques to add an adaptive offset term to a locally adaptive binarized expression.
Citation Information
Patent Citations
Memory hierarchy to transfer vector data for operators of a directed acyclic graph
US10437600B1
Using camera data to manage a vehicle parked outside in cold climates
US11001231B1
Generating training data for speed bump detection
US11586843B1
Generating detection parameters for a rental property monitoring solution using computer vision and audio analytics from a rental agreement
US11645706B1
Object-aware temperature anomalies monitoring and early warning by combining visual and thermal sensing sensing
US20220044023A1