Large zoom ratio lens calibration for electronic image stabilization

By combining inertial measurement units and digital signal processing, lens calibration values ​​are automatically determined, solving the problem of poor image stabilization under high zoom ratio lenses and achieving efficient and accurate electronic image stabilization.

CN121056735APending Publication Date: 2025-12-02AMBARELLA INT LP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202410683014.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing electronic image stabilization technologies cannot effectively distinguish between external vibrations and object movement under high zoom ratio lenses, resulting in poor image stabilization performance. Furthermore, the lens calibration process is time-consuming and relies on manual operation.

Method used

A method combining inertial measurement units and digital signal processing is adopted to determine calibration values ​​through lens projection models and motion information. The calibration process is automated, including lens geometric distortion compensation and additional compensation, and calibration values ​​are generated to achieve electronic image stabilization.

Benefits of technology

It achieves high-precision image stabilization at large zoom ratios, reduces manual intervention, and improves calibration efficiency and image stabilization accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121056735A_ABST
    Figure CN121056735A_ABST
Patent Text Reader

Abstract

An apparatus includes an interface and a processor. The interface may be configured to receive pixel data of an environment and movement information about the device. The processor may be configured to: process pixel data arranged as a video frame; measuring movement information; generating a playback record in response to video frames of the calibration target captured during a vibration mode applied to the device and movement information of the vibration mode; implementing image stabilization compensation in response to the lens projection function and the movement information; performing additional compensation in response to the calibration value; and performing a calibration operation to determine a calibration value. The playback record may include video frames generated at a plurality of predefined zoom levels of the shot, and the calibration operation may include determining coordinates of a calibration target for each of the video frames in the playback record.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to video capture, and more specifically to methods and / or apparatus for implementing high zoom ratio lens calibration for electronic image stabilization. Background Technology

[0002] Electronic image stabilization (EIS) is a crucial aspect of Internet Protocol (IP) cameras and other types of cameras. EIS is a highly complex system involving digital signal processing (DSP) that operates at frame-accurate levels. Conventional image stabilization techniques do not perform well when optical zoom reaches high levels (i.e., 10x to 40x). Generally, image stabilization performs worse as the optical zoom ratio increases.

[0003] Conventional image stabilization techniques rely on digital image stabilization (DIS) using pure image processing for lenses with large optical zoom. Without an inertial measurement unit (IMU), DIS cannot distinguish between external vibrations and relatively large objects moving within the field of view. Other such edge cases exist where DIS cannot provide an acceptable level of image stabilization.

[0004] Each lens used in a camera offers a different zoom level and has different inherent characteristics. To provide accurate EIS, each lens must be individually calibrated. Calibrating a lens can involve a significant amount of manual work by technicians / engineers (i.e., setting up the calibration scene, setting the camera to a specific zoom level, capturing images, etc.). Individually calibrating each lens at different zoom levels is a time-consuming process.

[0005] The goal is to achieve high zoom ratio lens calibration for electronic image stabilization. Summary of the Invention

[0006] This invention relates to an apparatus including an interface and a processor. The interface can be configured to receive pixel data of the environment and motion information about the apparatus. The processor can be configured to: process pixel data arranged as video frames; measure motion information; generate a playback recording in response to video frames of a calibration target captured during a vibration mode applied to the apparatus and motion information of the vibration mode; perform image stabilization compensation in response to a lens projection function and motion information; perform additional compensation in response to a calibration value; and perform a calibration operation to determine a calibration value. The playback recording may include video frames generated at multiple predefined zoom levels of the lens, and the calibration operation may include: determining the coordinates of the calibration target for each video frame in the playback recording; determining a pixel difference matrix in response to a comparison of the coordinates of the video frames determined using image stabilization compensation with the coordinates of the calibration target; and generating a calibration value in response to curve fitting performed on the pixel difference matrix. Attached Figure Description

[0007] Embodiments of the present invention will become apparent from the following detailed description, the appended claims, and the accompanying drawings.

[0008] Figure 1 This is a diagram illustrating an example of an Internet Protocol camera according to an exemplary embodiment of the present invention, which may utilize a processor configured to implement electronic image stabilization for high zoom ratio lenses.

[0009] Figure 2 This is a diagram of an example camera that demonstrates electronic image stabilization for high zoom ratio lenses using calibration values.

[0010] Figure 3 This is a block diagram showing the camera system.

[0011] Figure 4 This is a diagram showing movement information.

[0012] Figure 5 It is a graph showing the total compensation for the range of optical zoom ratios.

[0013] Figure 6 It is a graph showing the additional compensation for the range of optical zoom ratios.

[0014] Figure 7 This is a diagram illustrating the calibration of a high zoom ratio lens used for electronic image stabilization.

[0015] Figure 8 This is a diagram showing a camera system for capturing a calibrated target.

[0016] Figure 9 This is a diagram showing the predefined arrangement of the calibration targets.

[0017] Figure 10 It is a graph showing the curve representing the total compensation achieved by the capture device.

[0018] Figure 11 This is a graph showing the curve fit used to determine the calibration values ​​for additional compensation.

[0019] Figure 12 This is a flowchart illustrating a method for achieving a high zoom ratio lens for electronic image stabilization.

[0020] Figure 13 This is a flowchart illustrating a method for performing a calibration operation in response to a playback recording.

[0021] Figure 14 This is a flowchart illustrating a method for capturing a record of a calibration target for playback during vibration mode.

[0022] Figure 15 This is a flowchart illustrating a method for generating calibration values ​​for a camera during camera manufacturing. Detailed Implementation

[0023] Embodiments of the present invention include providing high zoom ratio lens calibration for electronic image stabilization, which may (i) utilize an inertial measurement unit, an image sensor, and digital signal processing at frame precision levels; (ii) provide image stabilization at high zoom ratios where pure digital image stabilization alone cannot provide accurate results; (iii) cover all extreme cases where digital image stabilization fails; (iv) utilize vibration frequency and amplitude from the inertial measurement unit; (v) implement a lens projection model; (vi) provide additional compensation with weights increasing with increasing optical zoom ratio; (vii) determine calibration values ​​based on motion information, pixel offset, image center distance, and the optical zoom ratio of the lens; (viii) calculate calibration values ​​using an on-chip system; (ix) generate a simulation frame record capturing metadata and video data of the calibration configuration; (x) automatically synchronize video data and vibration information to calibrate the lens for a range of predefined zoom levels; (xi) provide evaluation tools to ensure the accuracy of the electronic image stabilization results; and / or (xii) be implemented as one or more integrated circuits.

[0024] Embodiments of the present invention can be configured to provide a calibration system for determining calibration values ​​for electronic image stabilization (EIS) of a camera while capturing video data. The calibration system can be implemented via a system-on-a-chip (SoC) implemented by the camera device. In one example, the camera may be an Internet Protocol (IP) camera. In another example, the camera may be a handheld camera. EIS can be a complex system that combines various data sources to generate stabilized video frames in response to captured pixel data. These various data sources may include data from an inertial measurement unit (IMU), an image sensor, and digital signal processing (DSP) at frame-accuracy levels. Each camera (or camera lens) may have unique characteristics. The determined calibration values ​​enable the execution of accurate EIS that takes into account the unique characteristics of each camera.

[0025] An inertial measurement unit (IMU) can be implemented to determine motion information. For example, a gyroscope can be configured to capture the movement of an IP camera while capturing video frames, and the IMU can convert this movement into motion information. This motion information can be a source of data that can be used to perform EIS. Vibration devices (e.g., shaker devices) can provide vibration patterns to jitter the capture device during calibration. These vibration patterns can be captured by the IMU and used to determine calibration values.

[0026] EIS can be performed using techniques that consider lens projection models and motion information to determine the appropriate amount of correction. Additional compensation can be performed to provide state-of-the-art image stabilization results. In one example, when the optical zoom ratio becomes larger (e.g., from approximately 10x to 40x or greater), image stabilization using only lens projection models and motion information (e.g., a solution using only digital image stabilization) may not perform well. In another example, a solution using only digital image stabilization may fail to distinguish between external vibrations and large objects moving within the camera's field of view. Additional compensation can be configured to provide accurate image stabilization at large optical zoom ratios. Additional compensation can be configured to provide accurate image stabilization in various edge cases where digital image stabilization alone cannot provide accurate results. The combination of image stabilization using lens projection models and motion information with additional stabilization can operate similarly to optical image stabilization (OIS). The performed EIS can be compatible with IP cameras that implement computer vision and / or IP cameras that do not. The additional stabilization may rely on calibration values ​​determined by the calibration operation.

[0027] In some embodiments, external vibration compensation may depend in part on the lens projection model. In one example, when the effective focal length (EFL) is relatively small (e.g., it can be less than a relatively small EFL of about 50mm), the compensation (e.g., image stabilization) may depend largely on the lens projection model. For example, the lens projection model may be one of isometric (e.g., f-theta), stereo (e.g., custom distortion), pinhole model, etc. In some embodiments, the results of the lens projection model may be stored in a lookup table. In another example, when the EFL is relatively large (e.g., it can be greater than a relatively large EFL of about 50mm), the additional compensation may provide a higher proportion of image stabilization compared to the lens projection function. For example, the additional compensation may be determined independently of the lens projection function. Generally, the larger the optical zoom ratio, the more weighted the additional compensation can be applied (e.g., the higher the proportion of the compensation contributing to the total image stabilization).

[0028] Embodiments of the present invention can be configured to determine EIS, including at least two types of compensation. A first amount of compensation (e.g., contribution) can be determined as a function of lens geometry distortion projection and motion information. A second amount of compensation (e.g., contribution) can be determined as a function of various calibration values. Calibration values ​​can include a set of values ​​determined for a specific zoom ratio. For example, each zoom ratio can include a separate set of calibration values. Each of the two types of compensation can be applied to all zoom ratios. However, the first contribution (e.g., using lens geometry distortion projection and motion information) can have a higher weight at lower optical zoom ratios, and the second contribution can have a higher weight at higher optical zoom ratios. For example, the compensation amount from each type of compensation can be a variable ratio, and this ratio can be different at different zoom values ​​and different distances.

[0029] Lens geometric distortion can be determined based on various lens optical projection functions (e.g., design). In some embodiments, the processor can be configured to calculate the lens optical projection function. In some embodiments, a lookup table can be implemented to describe geometric distortion compensation for the lens optical projection at different angles and / or distances from the point to the lens center. Motion information can include vibration modes. Vibration modes can include the vibration frequency and amplitude in each rotational direction (e.g., pitch, yaw, and roll). Vibration modes can be applied during calibration.

[0030] Additional compensation can be determined based on the lens's inherent behavior and / or properties and / or motion information. Calibration values ​​can be determined based on zoom value, pixel offset, distance from the point to the lens center, and / or motion information. Pixel offset may include additional pixel offset relative to ideal geometric distortion projection due to the optical path affected by zoom. Specific results for each calibration value can be determined in response to multiple calibration operations. The contribution of each calibration value can vary depending on the design criteria of a particular implementation.

[0031] Embodiments of the present invention can be configured to determine calibration values ​​for each camera at the time of manufacture. The camera may include a processor and / or memory (e.g., a system-on-a-chip), which can be configured to implement emulation framing and / or calibration operations. The emulation framing can be configured to generate emulation framing records (e.g., playback records). The emulation framing can be configured to synchronize captured video frames with metadata. The metadata may include at least motion data captured by the IMU. The emulation framing can generate playback records in a format compatible with calibration equipment.

[0032] The calibration system may include multiple calibration targets. Calibration targets may include multiple test targets with predefined image patterns. In an example, the predefined image pattern may include a checkerboard (or checkerboard grid) pattern. The calibration targets may be arranged in a specific layout within the field of view of a camera at various optical zoom levels. The calibration system may include a vibration device that applies vibration to the camera while the camera captures a large number of video frames and / or images of the calibration targets. A simulation framework can synchronize the video frames captured from the calibration targets with the motion data and / or other metadata captured by the vibration patterns.

[0033] A camera-based simulation framework (e.g., a camera-based SoC) can be configured to partially automate the calibration process. Typically, calibration involves significant manual work by technicians / engineers. For example, a calibration process might include setting a calibration target in a predefined arrangement, setting the camera to a specific zoom level, capturing video frames at each zoom level, and / or calculating results for each zoom level. A simulation framework can save technicians / engineers time and / or effort. The simulation framework can be configured to automatically trigger the generation of playback recordings within a range of zoom levels. For example, a technician / engineer can set a calibration target in a predefined arrangement for various zoom levels and trigger the simulation framework, which can then automatically record data for a predefined range of optical zoom levels.

[0034] The calibration operation can be implemented by a camera (e.g., a SoC implemented by the camera). The calibration operation can use playback recordings to determine calibration values ​​for additional compensation. The calibration operation can be configured to apply a first contribution to image stabilization (e.g., based on a lens optical projection function) and analyze the playback recording without additional compensation. The calibration operation can be configured to perform computer vision operations to detect patterns on the calibration target. The calibration operation can generate a matrix that includes values ​​at specific locations on the calibration target. For example, the playback recording may include hundreds of video frames of the calibration target, and the calibration operation can generate a location-specific matrix on the calibration target for each of the video frames in the playback recording.

[0035] The calibration operation can be configured to generate a pixel difference matrix based on a specific location on the detected calibration target and a playback recording. The pixel difference matrix can include a grid of values ​​indicating the actual real-world location of a specific location (e.g., a grid location) on the calibration target, compared to the location detected when image stabilization is performed using only an image projection function. A pixel difference matrix can be generated for each zoom level. The calibration operation can be configured to perform curve fitting to determine a set of calibration values ​​that can be applied to provide additional compensation at a specific zoom ratio. The camera device can use these calibration values ​​to perform EIS (Enhanced Image Synthesis).

[0036] After the camera device determines a set of calibration values ​​for a specific zoom ratio, the camera system can apply the calibration values ​​to perform EIS. For example, the total amount of image compensation for EIS may include a first contribution (e.g., based on a lens projection function) and additional compensation. In some embodiments, the camera system can generate a new set of video frames that can be corrected using EIS based on the calibration values. The corrected image can be evaluated using an EIS evaluation tool implemented by the camera system's SoC. In some embodiments, the calibration values ​​can be applied to playback recordings, and the EIS evaluation tool can evaluate the output of the playback recordings with the calibration values ​​applied. The EIS evaluation tool can be configured to evaluate the accuracy of the result of total compensation applied to video frames corrected using total compensation based on the determined calibration values. If the EIS evaluation tool determines that the total compensation does not provide an accurate result, calibration can be repeated to generate a new set of calibration values. If the EIS evaluation tool determines that the total compensation does provide an accurate result, this set of calibration values ​​can be stored as a set of calibration values ​​for a specific zoom ratio, and the calibration operation can determine the next set of calibration values ​​for another zoom ratio.

[0037] refer to Figure 1 The diagram illustrates an example of an Internet Protocol (IP) camera according to an exemplary embodiment of the present invention, which may utilize a processor configured to implement electronic image stabilization for high zoom ratio lenses. A top view of zone 50 is shown. In the illustrated example, zone 50 may be an outdoor location. Streets, vehicles, and buildings are shown.

[0038] Devices 100a-100n are shown at various locations within zone 50. Devices 100a-100n can each implement an edge device. Edge devices 100a-100n may include intelligent IP cameras (e.g., camera systems). Edge devices 100a-100n may include low-power technologies designed for deployment in embedded platforms at the network edge (e.g., microprocessors running on sensors, cameras, or other battery-powered devices), where power consumption is a critical issue. In the examples, edge devices 100a-100n may include various traffic cameras and Intelligent Transportation System (ITS) solutions.

[0039] Edge devices 100a-100n can be implemented for a variety of applications. In the illustrated example, edge devices 100a-100n may include an Automatic License Plate Recognition (ANPR) camera 100a, a traffic camera 100b, a vehicle camera 100c, an access control camera 100d, an ATM camera 100e, a bullet camera 100f, a dome camera 100n, etc. In the example, edge devices 100a-100n can be implemented as a traffic camera and Intelligent Transportation System (ITS) solution designed to enhance road safety by utilizing a combination of people and vehicle detection, vehicle brand / model recognition, and Automatic License Plate Recognition (ANPR) functions.

[0040] In the illustrated example, zone 50 may be an outdoor location. In some embodiments, edge devices 100a-100n may be implemented at various indoor locations. In the example, edge devices 100a-100n may incorporate convolutional neural networks for use in security (surveillance) applications and / or access control applications. In the example, edge devices 100a-100n implemented as security cameras and access control applications may include battery-powered cameras, doorbell cameras, outdoor cameras, indoor cameras, etc. According to embodiments of the invention, security cameras and access control applications can realize performance benefits from the application of convolutional neural networks. In the example, edge devices utilizing convolutional neural networks according to embodiments of the invention can acquire large amounts of image data and perform on-device inference to obtain useful information (e.g., multiple temporal instances of images performed by each network), thereby reducing bandwidth and / or power consumption. In another example, security (surveillance) applications and / or location monitoring applications (e.g., tracking cameras) may benefit from large optical zoom. The design, type, and / or application performed by edge devices 100a-100n may vary depending on the design criteria of a particular implementation.

[0041] refer to Figure 2A diagram illustrating an example camera demonstrating electronic image stabilization for a high zoom ratio lens using calibration values ​​is shown. Camera systems 100a-100n are shown. Each camera device 100a-100n can have a different style and / or use case. For example, camera 100a can be an action camera, camera 100b can be a ceiling-mounted security camera, camera 100n can be a network camera, etc. Other types of cameras can be implemented (e.g., home security cameras, battery-powered cameras, doorbell cameras, stereo cameras, etc.). In some embodiments, camera systems 100a-100n can be fixed cameras (e.g., mounted and / or fixed in a single location). In some embodiments, camera systems 100a-100n can be handheld cameras. In some embodiments, camera systems 100a-100n can be configured for cross-area panning, can be attached to a mount, gimbal, camera stand, etc. The design / style of cameras 100a-100n can vary depending on the design criteria of a particular implementation.

[0042] Each of the camera systems 100a-100n may include block (or circuit) 102, block (or circuit) 104, and / or block (or circuit) 106. Circuit 102 may implement a processor. Circuit 104 may implement a capture device. Circuit 106 may implement an inertial measurement unit (IMU). Camera systems 100a-100n may include other components (not shown). They can be used with... Figure 3 The components of the cameras 100a-100n are described in detail.

[0043] Processor 102 can be configured to implement an artificial neural network (ANN). In an example, the ANN may include a convolutional neural network (CNN). Processor 102 can be configured to implement a video encoder. Processor 102 can be configured to process pixel data arranged as video frames. Capture device 104 can be configured to capture pixel data that can be used by processor 102 to generate video frames. IMU 106 can be configured to generate motion data (e.g., vibration information, camera jitter, translation direction, etc.). In some embodiments, a structured light projector can be implemented to project a speckle pattern onto the environment. Capture device 104 can capture pixel data including a background image (e.g., the environment) with a speckle pattern. Although each of the cameras 100a-100n is shown as not implementing a structured light projector, some of the cameras 100a-100n may utilize a structured light projector (e.g., a camera implementing a sensor for capturing IR light).

[0044] Cameras 100a-100n can be edge devices. A processor 102 implemented by each of the cameras 100a-100n enables the cameras 100a-100n to perform various functions internally (e.g., at the local level). For example, processor 102 can be configured to perform object / event detection (e.g., computer vision operations), 3D reconstruction, presence detection, depth map generation, video encoding, electronic image stabilization, and / or video transcoding on the device. For example, processor 102 can even perform advanced processes such as computer vision and 3D reconstruction without uploading video data to a cloud service, thereby offloading computationally intensive functions (e.g., computer vision, video encoding, video transcoding, etc.).

[0045] In some embodiments, multiple camera systems may be implemented (e.g., camera systems 100a-100n may operate independently of each other). For example, each camera in 100a-100n may individually analyze the captured pixel data and perform event / object detection locally. In some embodiments, cameras 100a-100n may be configured as a camera network (e.g., security cameras that send video data to a central source such as a network-attached storage device and / or cloud service). The location and / or configuration of cameras 100a-100n may vary according to the design criteria of a particular implementation.

[0046] The capture device 104 of each camera in the camera systems 100a-100n may include a single lens (e.g., a monocular camera). The processor 102 may be configured to accelerate the preprocessing of speckle structured light for monocular 3D reconstruction. Monocular 3D reconstruction can be performed to generate depth maps and / or disparity maps without using a stereo camera.

[0047] refer to Figure 3 The diagram illustrates a block diagram of a camera system. The camera system (or device) 100 can be... Figure 2 Representative examples of cameras 100a-100n are shown in association. Camera system 100 may include processor / SoC 102, capture device 104, and IMU 106.

[0048] The camera system 100 may further include blocks (or circuits) 150, 152, 154, 156, 158, 160, 164, and / or 166. Circuit 150 may implement a memory. Circuit 152 may implement a battery. Circuit 154 may implement a communication device. Circuit 156 may implement a wireless interface. Circuit 158 ​​may implement a general-purpose processor. Block 160 may implement an optical lens. Circuit 164 may implement one or more sensors. Circuit 166 may implement a human-machine interface (HID) device. In some embodiments, the camera system 100 may include a processor / SoC 102, a capture device 104, an IMU 106, a memory 150, a lens 160, a sensor 164, a battery 152, a communication module 154, a wireless interface 156, and a processor 158. In another example, camera system 100 may include a processor / SoC 102, a capture device 104, an IMU 106, a processor 158, a lens 160, and a sensor 164 as a single device, while memory 150, battery 152, communication module 154, and wireless interface 156 may be components of separate devices. Camera system 100 may include other components (not shown). The number, type, and / or arrangement of components in camera system 100 may vary depending on the design criteria of a particular implementation.

[0049] In some embodiments, processor 102 may be implemented as a video processor. In an example, processor 102 may be configured to receive three-sensor video input using a high-speed SLVS / MIPI-CSI / LVCMOS interface. In some embodiments, processor 102 may be configured to perform depth sensing in addition to generating video frames. In an example, depth sensing may be performed in response to depth information and / or vector light data captured in the video frames. In some embodiments, processor 102 may be implemented as a data stream vector processor. In an example, processor 102 may include a highly parallel architecture configured to perform image / video processing and / or radar signal processing.

[0050] Memory 150 can store data. Memory 150 can be implemented in various types of memory, including but not limited to cache, flash memory, memory card, random access memory (RAM), dynamic RAM (DRAM), etc. The type and / or size of memory 150 can vary according to the design criteria of a particular implementation. The data stored in memory 150 may correspond to video files, motion information (e.g., readings from sensor 164), video fusion parameters, image stabilization parameters, user input, computer vision models, feature sets, radar data cubes, radar detection and / or metadata information. In some embodiments, memory 150 can store reference images. Reference images can be used for computer vision operations, 3D reconstruction, automatic exposure, etc. In some embodiments, reference images may include reference structured light images.

[0051] Processor / SoC 102 can be configured to execute computer-readable code and / or procedural information. In various embodiments, the computer-readable code can be stored within processor / SoC 102 (e.g., microcode, etc.) and / or in memory 150. In one example, processor / SoC 102 can be configured to execute one or more artificial neural network models (e.g., face recognition CNN, object detection CNN, object classification CNN, 3D reconstruction CNN, presence detection CNN, etc.) stored in memory 150. In another example, memory 150 can store one or more directed acyclic graphs (DAGs) and one or more sets of weights or biases defining one or more artificial neural network models. In yet another example, memory 150 can store instructions for performing transformation operations (e.g., discrete cosine transform, discrete Fourier transform, fast Fourier transform, etc.). Processor / SoC 102 can be configured to receive input from memory 150 and / or present output to memory 150. Processor / SoC 102 can be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of the processor / SoC 102 can vary depending on the design criteria of a particular implementation. The processor / SoC 102 can be configured for low-power (e.g., battery) operation.

[0052] Battery 152 can be configured to store and / or supply power to components of camera system 100. The dynamic actuator mechanism for the rolling shutter sensor 130 can be configured to conserve power. Reduced power consumption allows camera system 100 to operate for extended periods using battery 152 without recharging. Battery 152 can be rechargeable. Battery 152 can be built-in (e.g., non-replaceable) or replaceable. Battery 152 can have an input for connection to an external power source (e.g., for charging). In some embodiments, device 100 can be powered by an external power source (e.g., battery 152 may not be implemented or may be implemented as a backup power source). Battery 152 can be implemented using various battery technologies and / or chemistry. The type of battery 152 implemented can vary depending on the design criteria of a particular implementation.

[0053] The communication module 154 can be configured to implement one or more communication protocols. For example, the communication module 154 and the wireless interface 156 can be configured to implement one or more of the following: IEEE 102.11, IEEE 102.15, IEEE 102.15.1, IEEE 102.15.2, IEEE 102.15.3, IEEE 102.15.4, IEEE 102.15.5, IEEE 102.20, and / or In some embodiments, the communication module 154 may be a hardwired data port (e.g., a USB port, a mini-USB port, a USB-C connector, an HDMI port, an Ethernet port, a DisplayPort interface, a Lightning port, etc.). In some embodiments, the wireless interface 156 may also implement one or more protocols associated with a cellular communication network (e.g., GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, SMS, etc.). In embodiments where the camera system 100 is implemented as a wireless camera, the protocols implemented by the communication module 154 and the wireless interface 156 may be wireless communication protocols. The type of communication protocol implemented by the communication module 154 may vary depending on the design criteria of the specific implementation.

[0054] Communication module 154 and / or wireless interface 156 can be configured to generate broadcast signals as output from camera system 100. The broadcast signals can transmit video data, parallax data, and / or (multiple) control signals to external devices. For example, the broadcast signals can be sent to cloud storage services (e.g., storage services capable of on-demand scaling). In some embodiments, communication module 154 may not transmit data until the processor / SoC 102 has performed video analysis and / or radar signal processing to determine that an object is in the field of view of camera system 100.

[0055] In some embodiments, the communication module 154 can be configured to generate a manual control signal. The manual control signal can be generated in response to a signal received from a user by the communication module 154. The manual control signal can be configured to activate the processor / SoC 102. The processor / SoC 102 can be activated in response to the manual control signal, regardless of the power state of the camera system 100.

[0056] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive a feature set. The received feature set may be used to detect events and / or objects. For example, the feature set may be used to perform computer vision operations. The feature set information may include instructions for the processor 102 to determine which types of objects correspond to objects of interest and / or events.

[0057] In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to receive user input. User input allows the user to adjust operating parameters for various features implemented by the processor 102. In some embodiments, the communication module 154 and / or the wireless interface 156 may be configured to interact with an application (e.g., an app) (e.g., using an application programming interface (API)). For example, the application may be implemented on a smartphone to allow the end user to adjust various settings and / or parameters for various features implemented by the processor 102 (e.g., setting video resolution, selecting frame rate, selecting output format, setting tolerance parameters for 3D reconstruction, etc.).

[0058] Processor 158 can be implemented using general-purpose processor circuitry. Processor 158 may be operable to interact with video processing circuitry 102 and memory 150 to perform various processing tasks. Processor 158 may be configured to execute computer-readable instructions. In one example, the computer-readable instructions may be stored in memory 150. In some embodiments, the computer-readable instructions may include controller operations. Typically, input from sensor 164 and / or human-machine interface device 166 is shown to be received by processor 102. In some embodiments, general-purpose processor 158 may be configured to receive and / or analyze data from sensor 164 and / or HID 166 and make decisions in response to input. In some embodiments, processor 158 may send data to and / or receive data from other components of camera system 100, such as battery 152, communication module 154, and / or wireless interface 156. In some embodiments, processor 158 may implement an integrated digital signal processor (IDSP). For example, IDSP 158 may be configured to implement a warp engine. Which functions of the camera system 100 are executed by processor 102 and general-purpose processor 158 can vary depending on the design criteria of a particular implementation.

[0059] Lens 160 may be attached to capture device 104. Capture device 104 may be configured to receive an input signal (e.g., LIN) via lens 160. The signal LIN may be an optical input (e.g., an analog image). Lens 160 may be implemented as an optical lens. Lens 160 may provide zoom and / or focus features. In one example, capture device 104 and / or lens 160 may be implemented as a single lens assembly. In another example, lens 160 may be implemented separately from capture device 104.

[0060] Capture device 104 can be configured to convert input light LIN into computer-readable data. Capture device 104 can capture data received through lens 160 to generate raw pixel data. In some embodiments, capture device 104 can capture data received through lens 160 to generate a bitstream (e.g., generate video frames). For example, capture device 104 can receive focused light from lens 160. Lens 160 can be oriented, tilted, panned, scaled, and / or rotated to provide a target view from camera system 100 (e.g., a view of video frames, a view of panoramic video frames captured using multiple camera systems 100a-100n, a target image and reference image view for stereo vision, etc.). Capture device 104 can generate a signal (e.g., VIDEO). The signal VIDEO can be pixel data (e.g., a sequence of pixels that can be used to generate video frames). In some embodiments, the signal VIDEO can be video data (e.g., a sequence of video frames). The signal VIDEO can be presented to one of the inputs of processor 102. In some embodiments, the pixel data generated by the capture device 104 may be uncompressed and / or raw data generated in response to focused light from the lens 160. In some embodiments, the output of the capture device 104 may be a digital video signal.

[0061] In this example, capture device 104 may include block (or circuitry) 180, block (or circuitry) 182, and block (or circuitry) 184. Circuitry 180 may be an image sensor. Circuitry 182 may be a processor and / or logic unit. Circuitry 184 may be memory circuitry (e.g., a frame buffer). Lens 160 (e.g., a camera lens) may be oriented to provide a view of the environment surrounding camera system 100. Lens 160 may be designed to capture environmental data (e.g., light input LIN). Lens 160 may be a wide-angle lens and / or a fisheye lens (e.g., a lens capable of capturing a wide field of view). Lens 160 may be configured to capture and / or focus light for capture device 104. Typically, image sensor 180 is located behind lens 160. Based on the light captured from lens 160, capture device 104 may generate bitstream and / or video data (e.g., a signal VIDEO).

[0062] Capture device 104 can be configured to capture video image data (e.g., light collected and focused by lens 160). Capture device 104 can capture data received through lens 160 to generate a video bitstream (e.g., pixel data for a sequence of video frames). In various embodiments, lens 160 can be implemented as a fixed-focus lens. Fixed-focus lenses are generally advantageous for smaller size and lower power consumption. In examples, fixed-focus lenses can be used in battery-powered, doorbell, and other low-power camera applications. In some embodiments, lens 160 can be oriented, tilted, panned, zoomed, and / or rotated to capture the environment around camera system 100 (e.g., capturing data from the field of view). In examples, professional camera models can utilize active lens systems for enhanced functionality, remote control, etc.

[0063] The capture device 104 can convert received light into a digital data stream. In some embodiments, the capture device 104 can perform analog-to-digital conversion. For example, the image sensor 180 can perform photoelectric conversion on the light received by the lens 160. The processor / logic unit 182 can convert the digital data stream into a video data stream (or bitstream), a video file, and / or multiple video frames. In this example, the capture device 104 can present the video data as a digital video signal (e.g., VIDEO). The digital video signal can include video frames (e.g., continuous digital images and / or audio). In some embodiments, the capture device 104 can include a microphone for capturing audio. In some embodiments, the microphone can be implemented as a separate component (e.g., one of the sensors 164).

[0064] Video data captured by capture device 104 can be represented as a signal / bitstream / data VIDEO (e.g., a digital video signal). Capture device 104 can present the signal VIDEO to processor / SoC 102. The signal VIDEO can represent video frames / video data. The signal VIDEO can be a video stream captured by capture device 104. In some embodiments, the signal VIDEO may include pixel data operable by processor 102 (e.g., a video processing pipeline, image signal processor (ISP), etc.). Processor 102 can generate video frames in response to the pixel data in the signal VIDEO.

[0065] The signal VIDEO may include pixel data arranged as video frames. In some embodiments, the signal VIDEO may be an image including a background (e.g., the captured environment and / or objects) and a speckle pattern generated by a structured light projector. The signal VIDEO may include a single-channel source image. A single-channel source image may be generated in response to capturing pixel data using a monocular lens 160.

[0066] Image sensor 180 can receive input light LIN from lens 160 and convert the light LIN into digital data (e.g., a bitstream). For example, image sensor 180 can perform photoelectric conversion on light from lens 160. In some embodiments, image sensor 180 may have an additional margin that is not used as part of the image output. In some embodiments, image sensor 180 may not have an additional margin. In various embodiments, image sensor 180 can be implemented as an RGB sensor, RGB-IR sensor, RCCB sensor, monocular image sensor, stereo image sensor, thermal sensor, event-based sensor, etc. For example, image sensor 180 can be any type of sensor configured to provide sufficient output for a computer vision operation (e.g., neural network-based detection) to be performed on the output data. In the context of the illustrated embodiments, image sensor 180 can be configured to generate an RGB-IR video signal. In a field of view illuminated only by infrared light, image sensor 180 can generate a monochrome (B / W) video signal. In a field of view illuminated by both IR and visible light, image sensor 180 can be configured to generate color information in addition to generating a monochrome video signal. In various embodiments, the image sensor 180 may be configured to generate video signals in response to visible light and / or infrared (IR) light.

[0067] In some embodiments, camera sensor 180 may include a rolling shutter sensor or a global shutter sensor. In an example, the rolling shutter sensor 180 may implement an RGB-IR sensor. In some embodiments, capture device 104 may include a rolling shutter IR sensor and an RGB sensor (e.g., implemented as separate components). In an example, the rolling shutter sensor 180 may be implemented as an RGB-IR rolling shutter complementary metal-oxide-semiconductor (CMOS) image sensor. In one example, the rolling shutter sensor 180 may be configured to assert a signal indicating the exposure time of the first row. In one example, the rolling shutter sensor 180 may apply a mask to a monochrome sensor. In an example, the mask may include multiple cells containing a red pixel, a green pixel, a blue pixel, and an IR pixel. The IR pixel may contain red, green, and blue filter materials that efficiently absorb all light in the visible spectrum while allowing longer infrared wavelengths to pass through with minimal loss. In the case of a rolling shutter, as each row (or line) of the sensor begins exposure, all pixels in that row (or line) may begin exposure simultaneously.

[0068] Processor / logic unit 182 can convert the bitstream into human-readable content (e.g., video data that an average person can understand regardless of image quality, such as video frames and / or pixel data that can be converted into video frames by processor 102). For example, processor / logic unit 182 can receive raw (e.g., raw) data from image sensor 180 and generate (e.g., encode) video data (e.g., bitstream) based on the raw data. Capture device 104 may have memory 184 to store raw data and / or processed bitstreams. For example, capture device 104 may implement frame memory and / or buffer 184 to store (e.g., provide temporary storage and / or cache) one or more video frames (e.g., digital video signals). In some embodiments, processor / logic unit 182 can perform analysis and / or correction on the video frames stored in memory / buffer 184 of capture device 104. Processor / logic unit 182 can provide status information about the captured video frames.

[0069] IMU 106 can be configured to detect motion and / or movement of camera system 100. IMU 106 is shown as receiving signals (e.g., MTN). Signal MTN may include a combination of forces acting on camera system 100. Signal MTN may include movement, vibration, jitter, translation, abrupt changes, etc. Signal MTN may represent movement in three-dimensional space (e.g., movement in the X, Y, and Z directions). The type and / or amount of motion received by IMU 106 may vary depending on the design criteria of a particular implementation.

[0070] IMU 106 may include block (or circuitry) 186. Circuitry 186 may implement a motion sensor. In one example, motion sensor 186 may be a gyroscope. Gyroscope 186 may be configured to measure the amount of movement. For example, gyroscope 186 may be configured to detect the amount and / or direction of movement of signal MTN and convert that movement into electrical data. IMU 106 may be configured to determine the amount and / or direction of movement measured by gyroscope 186. IMU 106 may convert the electrical data from gyroscope 186 into a format readable by processor 102. IMU 106 may be configured to generate a signal (e.g., M_INFO). Signal M_INFO may include measurement information in a format readable by processor 102. IMU 106 may present signal M_INFO to processor 102. The number, type, and / or arrangement of components of IMU 106 and / or the number, type, and / or functionality of signals communicated by IMU 106 may vary according to design criteria for a particular implementation.

[0071] Sensor 164 can implement multiple sensors, including but not limited to motion sensors, ambient light sensors, proximity sensors (e.g., ultrasonic, radar, passive infrared, lidar, etc.), audio sensors (e.g., microphones), etc. In embodiments implementing a motion sensor, sensor 164 can be configured to detect motion anywhere (or some location outside the field of view) monitored by camera system 100. In various embodiments, motion detection can be used as a threshold for activating capture device 104. Sensor 164 can be implemented as an internal component of camera system 100 and / or an external component of camera system 100. In one example, sensor 164 can be implemented as a passive infrared (PIR) sensor. In another example, sensor 164 can be implemented as a smart motion sensor. In yet another example, sensor 164 can be implemented as a microphone. In embodiments implementing a smart motion sensor, sensor 164 may include a low-resolution image sensor configured to detect motion and / or people.

[0072] In various embodiments, sensor 164 may generate signals (e.g., SENS). The SENS may include various data (or information) collected by sensor 164. In an example, the SENS may include data collected in response to motion detected in a monitored field of view, ambient light levels in the monitored field of view, and / or sound picked up in the monitored field of view. However, other types of data may be collected and / or generated based on application-specific design criteria. The SENS may be presented to processor / SoC 102. In an example, sensor 164 may generate (assert) the SENS when motion is detected in a field of view monitored by camera system 100. In another example, sensor 164 may generate (assert) the SENS when audio is triggered in a field of view monitored by camera system 100. In yet another example, sensor 164 may be configured to provide directional information about motion and / or sound detected in the field of view. This directional information may also be communicated to processor / SoC 102 via the SENS.

[0073] HID 166 can implement an input device. For example, HID 166 can be configured to receive human input. In one example, HID 166 can be configured to receive password input from a user. In another example, HID 166 can be configured to receive user input to provide various parameters and / or settings to processor 102 and / or memory 150. In some embodiments, camera system 100 may include a keypad, touchpad (or screen), doorbell switch, and / or other human-machine interface device (HID) 166. In an example, sensor 164 can be configured to determine when an object approaches HID 166. In an example where camera system 100 is implemented as part of an access control application, capture device 104 can be activated to provide images for identifying people attempting access, and the lock zone and / or illumination for access touchpad 166 can be turned on. For example, a combination of input from HID 166 (e.g., a password or PIN code) can be combined with presence determination and / or depth analysis performed by processor 102 to achieve two-factor authentication. HID 166 can present a signal (e.g., USR) to processor 102. The signal USR may include input received by HID 166.

[0074] In an embodiment of the camera system 100 implementing a structured light projector, the structured light projector may include a structured light pattern lens and / or a structured light source. The structured light source may be configured to generate a structured light pattern signal (e.g., a speckle pattern) that can be projected onto an environment near the camera system 100. The structured light pattern may be captured by a capture device 104 as part of a light input LIN. The structured light pattern lens may be configured to emit structured light generated by the structured light source of the structured light projector while protecting the structured light source. The structured light pattern lens may be configured to decompose the laser pattern generated by the structured light source into a pattern array (e.g., a dense dot pattern array for a speckle pattern).

[0075] In the example, the structured light source can be implemented as an array of vertical-cavity surface-emitting lasers (VCSELs) and a lens. However, other types of structured light sources can be implemented to meet the design criteria of specific applications. In the example, the VCSEL array is generally configured to generate laser patterns (e.g., signal SLP). The lens is generally configured to decompose the laser pattern into an array of dense point patterns. In the example, the structured light source can be implemented as a near-infrared (NIR) source. In various embodiments, the light source of the structured light source can be configured to emit light at a wavelength of approximately 940 nanometers (nm), which is invisible to the human eye. However, other wavelengths can be utilized. In the example, wavelengths in the range of approximately 800 nm to 1000 nm can be utilized.

[0076] Processor / SoC 102 can receive signals VIDEO, M_INFO, SENS, and USR. Processor / SoC 102 can generate one or more video output signals (e.g., VIDOUT), one or more control signals (e.g., CTRL), one or more depth data signals (e.g., DIMAGES), and / or one or more twist table data signals (e.g., WT) based on signals VIDEO, M_INFO, SENS, USR, and / or other inputs. In some embodiments, signals VIDOUT, DIMAGES, WT, and CTRL can be generated based on analysis of signals VIDEO and / or objects detected in signals VIDEO. In some embodiments, signals VIDOUT, DIMAGES, WT, and CTRL can be generated based on analysis of signals VIDEO, motion information captured by IMU 106, and / or inherent properties of lens 160 and / or capture device 104.

[0077] In various embodiments, the processor / SoC 102 may be configured to perform one or more of the following: feature extraction, object detection, object tracking, electronic image stabilization, 3D reconstruction, presence detection, and object identification. For example, the processor / SoC 102 may determine motion information and / or depth information by analyzing frames from a signal VIDEO and comparing those frames with previous frames. The comparison may be used to perform digital motion estimation. In some embodiments, the processor / SoC 102 may be configured to generate a video output signal VIDOUT, a distortion table data signal WT, and / or a depth data signal DIMAGES, including a disparity map and a depth map, based on the signal VIDEO. The video output signal VIDOUT, the distortion table data signal WT, and / or the depth data signal DIMAGES may be presented to memory 150, communication module 154, and / or wireless interface 156. In some embodiments, the video signal VIDOUT, the distortion table data signal WT, and / or the depth data signal DIMAGES may be used internally by the processor 102 (e.g., not presented as output). In one example, the warped table data signal WT can be used by a warping engine implemented by a digital signal processor (e.g., processor 158).

[0078] The signal VIDOUT can be presented to the communication device 156. In some embodiments, the signal VIDOUT may include encoded video frames generated by the processor 102. In some embodiments, the encoded video frames may include a complete video stream (e.g., encoded video frames representing all video captured by the capture device 104). The encoded video frames may be encoded, cropped, stitched, stabilized, and / or enhanced versions of pixel data received from the signal VIDEO. In the example, the encoded video frames may be high-resolution, digital, encoded, de-distorted, stabilized, cropped, blended, stitched, and / or rolled shutter effect corrected versions of the signal VIDEO.

[0079] In some embodiments, the signal VIDOUT can be generated based on video analysis (e.g., computer vision operations) performed by processor 102 on generated video frames. Processor 102 can be configured to perform computer vision operations to detect objects and / or events in the video frames, and then convert the detected objects and / or events into statistics and / or parameters. In one example, the data determined by the computer vision operations can be converted by processor 102 into a human-readable format. The data from the computer vision operations can be used to detect objects and / or events. The computer vision operations can be performed locally by processor 102 (e.g., without needing to communicate with external devices to offload computational operations). Similarly, other video processing and / or encoding operations (e.g., stabilization, compression, stitching, cropping, rolling shutter effect correction, etc.) can be performed locally by processor 102. For example, locally executed computer vision operations allow the computer vision operations to be performed by processor 102 and avoid heavy video processing running on a backend server. Avoiding video processing running on a backend (e.g., remotely located) server can protect privacy.

[0080] In some embodiments, the signal VIDOUT can be data generated by processor 102 (e.g., video analysis results, audio / speech analysis results, stabilized video frames, etc.) that can be communicated to cloud computing services to aggregate information and / or provide training data for machine learning (e.g., to improve object detection, improve audio detection, improve presence detection, etc.). In some embodiments, the signal VIDOUT can be provided to cloud services for mass storage (e.g., to enable users to retrieve encoded video using smartphones and / or desktop computers). In some embodiments, the signal VIDOUT can include data extracted from video frames (e.g., computer vision results) and can communicate the results to another device (e.g., a remote server, cloud computing system, etc.) to offload the analysis of the results to another device (e.g., offloading the analysis of the results to a cloud computing service instead of performing all analysis locally). The type of information communicated by the signal VIDOUT can vary depending on the design criteria of a particular implementation.

[0081] The CTRL signal can be configured to provide a control signal. The CTRL signal can be generated in response to a decision made by the processor 102. In one example, the CTRL signal can be generated in response to a detected object and / or features extracted from a video frame. The CTRL signal can be configured to enable, disable, or change the operating mode of another device. In one example, the CTRL signal can be used to lock / unlock a door controlled by an electronic lock. In another example, the CTRL signal can be used to put a device into sleep mode (e.g., low power mode) and / or activate it from sleep mode. In yet another example, the CTRL signal can be used to generate an alarm and / or notification. The type of device controlled by the CTRL signal and / or the response performed by the device in response to the CTRL signal can vary depending on the design criteria of the specific implementation.

[0082] A signal CTRL can be generated based on data received by sensor 164 (e.g., temperature readings, motion sensor readings, etc.). A signal CTRL can be generated based on input from HID 166. A signal CTRL can be generated based on human behavior detected by processor 102 in a video frame. A signal CTRL can be generated based on the type of object detected (e.g., person, animal, vehicle, etc.). A signal CTRL can be generated in response to detecting a specific type of object at a specific location. A signal CTRL can be generated in response to user input to provide various parameters and / or settings to processor 102 and / or memory 150. Processor 102 can be configured to generate a signal CTRL in response to sensor fusion operations (e.g., aggregating information received from different sources). Processor 102 can be configured to generate a signal CTRL in response to the result of an presence detection performed by processor 102. The conditions used to generate the signal CTRL can vary depending on the design criteria of a particular implementation.

[0083] Signals DIMAGES may include one or more depth maps and / or disparity maps generated by processor 102. Signals DIMAGES may be generated in response to 3D reconstruction performed on a monocular single-channel image. Signals DIMAGES may be generated in response to analysis of captured video data and structured light patterns.

[0084] A multi-step approach can be implemented to activate and / or disable the capture device 104 based on the output of the motion sensor 164 and / or any other power consumption characteristics of the camera system 100, thereby reducing the power consumption of the camera system 100 and extending the lifespan of the battery 152. The motion sensor in sensor 164 may have low power consumption on the battery 152 (e.g., less than 10W). In the example, the motion sensor in sensor 164 may be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 102. Video analysis performed by the processor / SoC 102 may have relatively high power consumption on the battery 152 (e.g., greater than that of the motion sensor 164). In the example, the processor / SoC 102 may be in a low-power state (or powered down) until some motion is detected by the motion sensor in sensor 164.

[0085] Camera system 100 can be configured to operate using various power states. For example, in a power-down state (e.g., sleep state, low power state), the motion sensor in sensor 164 and processor / SoC 102 can be turned on, and other components of camera system 100 (e.g., image capture device 106, memory 150, communication module 154, etc.) can be turned off. In another example, camera system 100 can operate in an intermediate state. In the intermediate state, image capture device 106 can be turned on, and memory 150 and / or communication module 154 can be turned off. In yet another example, camera system 100 can operate in a power-on (or high-power) state. In the power-on state, sensor 164, processor / SoC 102, capture device 104, memory 150, and / or communication module 154 can be turned on. Camera system 100 can consume some power from battery 152 (e.g., a relatively small and / or minimal amount of power) in the power-down state. In the power-on state, camera system 100 can consume more power from battery 152. The number of components of the camera system 100 that are turned on when the camera system 100 is in power state and / or operating in each power state can vary according to the design criteria of a particular implementation.

[0086] In some embodiments, the camera system 100 may be implemented as a system-on-a-chip (SoC). For example, the camera system 100 may be implemented as a printed circuit board including one or more components. The camera system 100 may be configured to perform intelligent video analysis on video frames of a video. The camera system 100 may be configured to crop and / or enhance the video.

[0087] In some embodiments, the video frame may be a view (or a derivative of a view) captured by the capture device 104. Pixel data signals may be enhanced by the processor 102 (e.g., color conversion, noise filtering, automatic exposure, automatic white balance, automatic focus, etc.). In some embodiments, the video frame may provide a series of cropped and / or enhanced video frames that improve the view from the perspective of the camera system 100 (e.g., providing night vision, providing high dynamic range (HDR) imaging, providing more viewing area, highlighting detected objects, providing additional data (e.g., digital distance to the detected object), etc.) to enable the processor 102 to see the location better than a human can see with human vision.

[0088] Encoded video frames can be processed locally. In one example, the encoded video can be stored locally by memory 150, enabling processor 102 to facilitate computer vision analysis internally (e.g., without first uploading the video frames to a cloud service). Processor 102 can be configured to select video frames to be encapsulated into a video stream that can be transmitted over a network (e.g., a bandwidth-limited network).

[0089] In some embodiments, processor 102 may be configured to perform sensor fusion operations. The sensor fusion operations performed by processor 102 may be configured to analyze information from multiple sources (e.g., capture device 104, IMU 106, sensor 164, and HID 166). By analyzing various data from different sources, sensor fusion operations may be able to make inferences about the data that might not be possible from just one data source. For example, sensor fusion operations implemented by processor 102 may analyze video data (e.g., mouth movements) as well as speech patterns from directional audio. Different sources can be used to develop scene models to support decision-making. For example, processor 102 may be configured to compare the synchronization of detected speech patterns with mouth movements in video frames to determine which person is speaking in the video frame. Sensor fusion operations may also provide temporal correlation, spatial correlation, and / or reliability between the received data.

[0090] In some embodiments, processor 102 may implement convolutional neural network (CNN) capabilities. CNN capabilities can be implemented using deep learning techniques for computer vision. CNN capabilities can be configured to perform pattern and / or image recognition using a training process involving multi-layer feature detection. Computer vision and / or CNN capabilities can be executed locally by processor 102. In some embodiments, processor 102 may receive training data and / or feature set information from external sources. For example, external devices (e.g., cloud services) may access various data sources to provide training data that camera system 100 may not be able to obtain. However, computer vision operations performed using feature sets can be performed using the computational resources of processor 102 within camera system 100.

[0091] The video pipeline of processor 102 can be configured to perform local dewarping, cropping, enhancement, rolling shutter correction, stabilization, downsizing, encapsulation, compression, conversion, mixing, synchronization, and / or other video operations. The video pipeline of processor 102 can support multiple streams (e.g., generating multiple bitstreams in parallel, each including a different bitrate). In the example, the video pipeline of processor 102 can implement an image signal processor (ISP) with an input pixel rate of 320 Mbps. The architecture of the video pipeline of processor 102 enables real-time and / or near-real-time video operations on high-resolution video and / or high-bitrate video data. The video pipeline of processor 102 can perform computer vision processing, stereo vision processing, object detection, 3D noise reduction, fisheye lens correction (e.g., real-time 360-degree dewarping and lens distortion correction), oversampling, and / or high dynamic range processing on 4K resolution video data. In one example, the architecture of the video pipeline can achieve 4K ultra-high resolution with H.264 encoding at twice the real-time speed (e.g., 60fps) and 4K ultra-high resolution with H.265 / HEVC and / or 4K AVC encoding at 30fps (e.g., 4KP30 AVC and HEVC encoding with multi-stream support). The type of video operation and / or the type of video data operated on by the processor 102 can vary according to the design criteria of the specific implementation.

[0092] The camera sensor 180 can be a high-resolution sensor. Using the high-resolution sensor 180, the processor 102 can combine oversampling of the image sensor 180 with digital scaling within the cropped area. Oversampling and digital scaling can each be one of the video operations performed by the processor 102. Oversampling and digital scaling can be implemented to provide a higher resolution image within the total size limit of the cropped area.

[0093] In some embodiments, lens 160 may be a fisheye lens. One of the video operations implemented by processor 102 may be a dedistortion operation. Processor 102 may be configured to dedistort generated video frames. Dedistortion may be configured to reduce and / or remove severe distortion caused by fisheye lens and / or other lens characteristics. For example, dedistortion may reduce and / or eliminate bulging effects to provide a linear image.

[0094] Processor 102 can be configured to crop (e.g., trim) a region of interest from a full video frame (e.g., generate a region of interest video frame). Processor 102 can generate video frames and select regions. In the example, cropping the region of interest can generate a second image. The cropped image (e.g., the region of interest video frame) can be smaller than the original video frame (e.g., the cropped image can be a portion of the captured video).

[0095] The region of interest (ROI) can be dynamically adjusted based on the location of the audio source. For example, the detected audio source may be moving, and its location may shift as video frames are captured. Processor 102 can update the coordinates of the selected ROI and dynamically update the cropped portion (e.g., a directional microphone implemented as one or more sensors in sensor 164 can dynamically update its location based on captured directional audio). The cropped portion may correspond to the selected ROI. As the ROI changes, the cropped portion may change. For example, the selected coordinates of the ROI may change from frame to frame, and processor 102 may be configured to crop the selected region in each frame.

[0096] Processor 102 can be configured to oversample image sensor 180. Oversampling of image sensor 180 can produce a higher resolution image. Processor 102 can be configured to digitally upscale regions of video frames. For example, processor 102 can digitally upscale a cropped region of interest. For example, processor 102 can establish a region of interest based on directional audio, crop the region of interest, and then digitally upscale the cropped region of interest video frame.

[0097] The dedistortion operation performed by processor 102 can adjust the visual content of the video data. The adjustment performed by processor 102 can make the visual content appear natural (e.g., as if seen by a person viewing a position corresponding to the field of view of capture device 104). In the example, dedistortion can alter the video data to generate linear video frames (e.g., correcting artifacts caused by lens characteristics of lens 160). Dedistortion operations can be implemented to correct distortion caused by lens 160. Adjusted visual content can be generated to achieve more accurate and / or reliable object detection.

[0098] Various features (e.g., de-distortion, digital scaling, cropping, etc.) can be implemented as hardware modules in processor 102. Implementing hardware modules can increase the video processing speed of processor 102 (e.g., faster than software implementations). Hardware implementations enable video processing while reducing latency. The hardware components used can vary depending on the design standards of the specific implementation.

[0099] In some embodiments, processor 102 may implement one or more coprocessors, cores, and / or chiplets. For example, processor 102 may implement one coprocessor configured as a general-purpose processor and another coprocessor configured as a video processor. In some embodiments, processor 102 may be a dedicated hardware module designed to perform a specific task. In one example, processor 102 may implement an AI accelerator. In another example, processor 102 may implement a radar processor. In yet another example, processor 102 may implement a dataflow vector processor. In some embodiments, other processors implemented by device 100 may be general-purpose processors and / or video processors (e.g., coprocessors that are physically different from processor 102 in chipset and / or silicon). In one example, processor 102 may implement the x86-64 instruction set. In another example, processor 102 may implement the ARM instruction set. In yet another example, processor 102 may implement the RISC-V instruction set. The number of cores, coprocessors, design optimizations, and / or instruction sets implemented by processor 102 may vary depending on the design criteria of a particular implementation.

[0100] The processor 102 shown includes multiple blocks (or circuits) 190a-190n. Blocks 190a-190n can implement various hardware modules implemented by the processor 102. Hardware modules 190a-190n can be configured to provide various hardware components to implement video processing pipelines, radar signal processing pipelines, and / or AI processing pipelines. Circuits 190a-190n can be configured to receive pixel data VIDEO, generate video frames based on the pixel data, perform various operations on the video frames (e.g., de-distortion, rolling shutter correction, cropping, magnification, image stabilization, 3D reconstruction, presence detection, automatic exposure, etc.), prepare video frames for communication with external hardware (e.g., encoding, encapsulation, color correction, etc.), parse feature sets, and implement various computer vision operations (e.g., object detection, segmentation, classification, etc.). Hardware modules 190a-190n can be configured to implement various security features (e.g., secure boot, I / O virtualization, etc.). Various implementations of processor 102 may not necessarily utilize all the features of hardware modules 190a-190n. The features and / or functions of hardware modules 190a-190n may vary depending on the design criteria of a particular implementation. Details of hardware modules 190a-190n may be described in connection with U.S. Patent Application No. 16 / 831,549, filed April 16, 2020; U.S. Patent Application No. 16 / 288,922, filed February 28, 2019; U.S. Patent Application No. 15 / 593,493 (now U.S. Patent No. 10,437,600), filed May 12, 2017; U.S. Patent Application No. 15 / 931,942, filed May 14, 2020; U.S. Patent Application No. 16 / 991,344, filed August 12, 2020; and U.S. Patent Application No. 17 / 479,034, filed September 20, 2021, the appropriate portions of which are incorporated herein by reference in their entirety.

[0101] Hardware modules 190a-190n can be implemented as dedicated hardware modules. Compared to software implementation, using dedicated hardware modules 190a-190n to implement various functions of processor 102 allows processor 102 to be highly optimized and / or customized to limit power consumption, reduce heat generation, and / or increase processing speed. Hardware modules 190a-190n can be customizable and / or programmable to implement multiple types of operations. Implementing dedicated hardware modules 190a-190n allows the hardware used to perform each type of computation to be optimized for speed and / or efficiency. For example, hardware modules 190a-190n can implement several relatively simple operations frequently used in computer vision operations, which together enable computer vision operations to be performed in real time. The video pipeline can be configured to identify objects. Objects can be identified by interpreting numerical and / or symbolic information to determine the specific type and / or characteristics of the object represented by the visual data. For example, the number of pixels and / or the color of pixels in video data can be used to identify portions of video data as objects. Hardware modules 190a-190n enable computationally intensive operations (e.g., computer vision operations, video encoding, video transcoding, 3D reconstruction, depth map generation, presence detection, etc.) to be performed locally by the camera system 100.

[0102] One of the hardware modules 190a-190n (e.g., 190a) can implement a scheduler circuit. Scheduler circuit 190a can be configured to store a directed acyclic graph (DAG). In the example, scheduler circuit 190a can be configured to generate and store a DAG in response to received (e.g., loaded) feature set information. The DAG can define video operations to be performed to extract data from video frames. For example, the DAG can define various mathematical weights (e.g., neural network weights and / or biases) to be applied when performing computer vision operations to classify various groups of pixels into specific objects.

[0103] Scheduler circuit 190a can be configured to parse acyclic graphs to generate various operators. Operators can be scheduled by scheduler circuit 190a in one or more of the other hardware modules 190a-190n. For example, one or more hardware modules 190a-190n can implement a hardware engine configured to perform a specific task (e.g., a hardware engine designed to perform repetitive specific mathematical operations used for performing computer vision operations). Scheduler circuit 190a can schedule operators based on when they are ready to be processed by hardware engines 190a-190n.

[0104] Scheduler circuit 190a can time-multiplex tasks to hardware modules 190a-190n for execution based on the availability of hardware modules 190a-190n. Scheduler circuit 190a can parse a directed acyclic graph into one or more data streams. Each data stream can include one or more operators. Once the directed acyclic graph is parsed, scheduler circuit 190a can allocate data streams / operators to hardware engines 190a-190n and send relevant operator configuration information to initiate the operators.

[0105] The binary representation of each directed acyclic graph can be an ordered traversal of the directed acyclic graph, where descriptors and operators are interleaved based on data dependencies. Descriptors typically provide registers that link data buffers to specific operands in a dependent operator. In various embodiments, operators may not appear in the directed acyclic graph representation until all dependent descriptors have been declared for operands.

[0106] One of the hardware modules 190a-190n (e.g., 190b) can implement an Artificial Neural Network (ANN) module. The ANN module can be implemented as a fully connected neural network or a Convolutional Neural Network (CNN). In this example, the fully connected network is "structure-agnostic" because no special assumptions need to be made about the input. A fully connected neural network consists of a series of fully connected layers that connect each neuron in one layer to each neuron in another layer. In a fully connected layer, for n inputs and m outputs, there are n*m ​​weights. There are also bias values ​​for each output node, resulting in a total of (n+1)*m parameters. In a trained neural network, (n+1)*m parameters have been determined during the training process. A trained neural network typically includes an architectural specification and a set of parameters (weights and biases) determined during training. In another example, a CNN architecture may explicitly assume that the input is an image, enabling the encoding of specific attributes into the model architecture. A CNN architecture may include a sequence of layers, where each layer transforms one activation into another through a differentiable function.

[0107] In the example shown, the artificial neural network 190b can implement a convolutional neural network (CNN) module. The CNN module 190b can be configured to perform computer vision operations on video frames. The CNN module 190b can be configured to achieve object recognition through multi-layer feature detection. The CNN module 190b can be configured to compute descriptors based on the performed feature detections. The descriptors enable the processor 102 to determine the probability that a pixel in a video frame corresponds to a specific object (e.g., a specific manufacture / model / year of a vehicle, identifying a person as a specific individual, detecting the type of animal, detecting facial features, etc.).

[0108] CNN module 190b can be configured to implement convolutional neural network capabilities. CNN module 190b can be configured to use deep learning techniques to implement computer vision. CNN module 190b can be configured to use a training process through multi-layer feature detection to implement pattern and / or image recognition. CNN module 190b can be configured to perform inference against machine learning models.

[0109] The CNN module 190b can be configured to perform feature extraction and / or matching solely in hardware. Feature points typically represent regions of interest (e.g., corners, edges, etc.) in a video frame. By temporarily tracking feature points, estimates of the self-motion of the capture platform or motion models of objects observed in a scene can be generated. To track feature points, a matching operation is typically incorporated into the CNN module 190b by hardware to find the most probable correspondence between feature points in a reference video frame and a target video frame. During the matching of reference and target feature points, each feature point can be represented by a descriptor (e.g., image patch, SIFT, BRIEF, ORB, FREAK, etc.). Implementing the CNN module 190b using dedicated hardware circuitry allows for real-time computation of descriptor matching distances.

[0110] The CNN module 190b can be configured to perform face detection, face recognition, and / or presence determination. For example, face detection, face recognition, and / or presence determination can be performed based on a trained neural network implemented by the CNN module 190b. In some embodiments, the CNN module 190b can be configured to generate a depth image based on a structured light pattern. The CNN module 190b can be configured to perform various detection and / or recognition operations and / or perform 3D recognition operations.

[0111] CNN module 190b can be a dedicated hardware module configured to perform feature detection on video frames. Features detected by CNN module 190b can be used to compute descriptors. CNN module 190b can determine, in response to the descriptors, the probability that a pixel in a video frame belongs to a specific object and / or multiple objects. For example, using the descriptors, CNN module 190b can determine the probability that a pixel corresponds to a specific object (e.g., a person, furniture item, pet, vehicle, etc.) and / or characteristics of the object (e.g., the shape of the eyes, the distance between facial features, the hood of a vehicle, body parts, a vehicle's license plate, a person's face, clothing worn by a person, etc.). Implementing CNN module 190b as a dedicated hardware module of processor 102 allows device 100 to perform computer vision operations locally (e.g., on-chip) without relying on the processing power of remote devices (e.g., communicating data to a cloud computing service).

[0112] The computer vision operations performed by the CNN module 190b can be configured to perform feature detection on video frames to generate descriptors. The CNN module 190b can perform object detection to determine regions in the video frames with a high probability of matching a specific object. In one example, an open operand stack can be used to customize the type of the objects(s) to be matched (e.g., reference objects) (thus enabling the programmability of processor 102 to implement various artificial neural networks defined by directed acyclic graphs, each providing instructions for performing various types of object detection). The CNN module 190b can be configured to perform local masking on regions with a high probability of matching a specific object(s) to detect the object.

[0113] In some embodiments, the CNN module 190b can determine the location (e.g., 3D coordinates and / or positional coordinates) of various features (e.g., characteristics) of a detected object. In one example, 3D coordinates can be used to determine the position of a person's arms, legs, chest, and / or eyes. A positional coordinate for the vertical positioning of a body part in 3D space on a first axis and another coordinate for the horizontal positioning of a body part in 3D space on a second axis can be stored. In some embodiments, the distance from the lens 160 can represent a coordinate for the depth positioning of a body part in 3D space (e.g., a positional coordinate on a third axis). Using the positioning of the various body parts in 3D space, the processor 102 can determine the body position and / or the detected human body characteristics.

[0114] The CNN module 190b can be pre-trained (e.g., configured to perform computer vision to detect objects based on received training data). For example, the results of the training data (e.g., a machine learning model) can be pre-programmed and / or loaded into processor 102. The CNN module 190b can perform inference against the machine learning model (e.g., to perform object detection). Training may include determining weights for each layer of the neural network model. For example, weights can be determined for each layer used for feature extraction (e.g., convolutional layers) and / or for classification (e.g., fully connected layers). The weights learned by the CNN module 190b can vary depending on the design criteria of a particular implementation.

[0115] The CNN module 190b can perform feature extraction and / or object detection by performing convolution operations. These convolution operations can be hardware-accelerated for fast (e.g., real-time) computation that can be performed with low power consumption. In some embodiments, the convolution operations performed by the CNN module 190b can be used to perform computer vision operations. In some embodiments, the convolution operations performed by the CNN module 190b can be used for any function (e.g., 3D reconstruction) that may involve computing convolution operations and is performed by the processor 102.

[0116] Convolution operations can include sliding a feature detection window along a layer while performing computations (e.g., matrix operations). The feature detection window can apply filters to pixels and / or extract features associated with each layer. The feature detection window can be applied to a pixel and multiple surrounding pixels. In the example, a layer can be represented as a matrix representing the values ​​of features from one of the pixels and / or layers, and the filters applied by the feature detection window can be represented as matrices. Convolution operations can apply matrix multiplication between regions of the current layer covered by the feature detection window. Convolution operations can cause the feature detection window to slide along a region of the layer to generate a result representing each region. The size of the regions, the type of operation applied by the filters, and / or the number of layers can vary depending on the design criteria of a particular implementation.

[0117] Using convolutional operations, the CNN module 190b can compute multiple features for pixels of the input image at each extraction step. For example, each layer in a layer can receive input from a set of features located in a small neighborhood (e.g., a region) of a previous layer (e.g., a local receptive field). Convolutional operations can extract basic visual features (e.g., oriented edges, endpoints, corners, etc.), which are then combined by higher layers. Since the feature extraction window operates on pixels and nearby pixels (or subpixels), the result of the operation can have location invariance. Layers can include convolutional layers, pooling layers, non-linear layers, and / or fully connected layers. In the example, convolutional operations can learn to detect edges from raw pixels (e.g., the first layer), then use features from previous layers (e.g., detected edges) to detect shapes in the next layer, and then use the shapes to detect higher-level features in higher layers (e.g., facial features, pets, vehicles, vehicle components, furniture, etc.), and the last layer can be a classifier using the higher-level features.

[0118] The CNN module 190b can perform data streams directed to feature extraction and matching, including two-stage detection, warping operators, component operators manipulating component lists (e.g., components can be regions of vectors sharing common attributes and can be combined with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, min-max reduction operators, pooling operators, non-minimum non-maximum suppression operators, scan-window-based non-maximum suppression operators, aggregation operators, scattering operators, statistical operators, classifier operators, integral image operators, comparison operators, indexing operators, pattern matching operators, feature extraction operators, feature detection operators, two-stage object detection operators, score generation operators, block reduction operators, and upsampling operators. The types of operations performed by the CNN module 190b to extract features from training data can vary depending on the design criteria of a particular implementation.

[0119] One or more of the hardware modules 190a-190n can be configured to implement other types of AI models. In one example, hardware modules 190a-190n can be configured to implement image-to-text AI models and / or video-to-text AI models. In another example, hardware modules 190a-190n can be configured to implement large language models (LLMs). Implementing AI models using hardware modules 190a-190n can provide AI acceleration, enabling the execution of complex AI tasks on edge devices such as edge devices 100a-100n.

[0120] One of the hardware modules 190a-190n can be configured to perform virtual aperture imaging. One of the hardware modules 190a-190n can be configured to perform transform operations (e.g., FFT, DCT, DFT, etc.). The number, type, and / or operations performed by the hardware modules 190a-190n can vary according to the design criteria of a particular implementation.

[0121] Each of the hardware modules 190a-190n can implement a processing resource (or hardware resource or hardware engine). Hardware engines 190a-190n can be operable to perform specific processing tasks. In some configurations, hardware engines 190a-190n can operate in parallel and independently of each other. In other configurations, hardware engines 190a-190n can operate collaboratively to perform assigned tasks. One or more of the hardware engines 190a-190n can be homogeneous processing resources (all circuits 190a-190n can have the same capabilities) or heterogeneous processing resources (two or more circuits 190a-190n can have different capabilities).

[0122] refer to Figure 4A diagram illustrating movement information is shown. A coordinate system 200 is shown. A camera system 100 is shown on coordinate system 200. The IMU 106 and lens 160 of the camera 100 are shown.

[0123] The camera system 100 can implement optical zoom 202. For example, lens 160 can provide optical zoom 202. In some embodiments, lens 160 can implement optical zoom 202, and processor 102 can also implement digital zoom. Optical zoom 202 can make the captured environment and / or objects appear physically closer before pixel data is captured. For example, digital zoom can be a form of post-processing performed by processor 102, and optical zoom 202 can be a physical process performed by lens 160 that can increase magnification. Optical zoom 202 can be implemented in response to components moving within camera lens 160. For example, lens 160 can be adjusted to increase focal length. Overall, optical zoom 202 can magnify the subject of a video frame while maintaining image quality. In one example, optical zoom 202 can be 5x zoom. In another example, optical zoom 202 can be 10x zoom. In yet another example, optical zoom 202 can be 40x zoom. In some embodiments, the camera system 100 may implement an optional optical zoom 202 ranging from 1x to 83x. The amount of optical zoom 202 implemented may vary depending on the design criteria of a particular implementation.

[0124] Block (or circuitry) 204 is shown. Circuitry 204 may implement a vibration device (or calibration shaker). Vibration device 204 may be configured to generate a signal (e.g., VPTN). In the example shown, the signal VPTN is shown as being presented to camera system 100 for illustrative purposes. Typically, the signal VPTN may be used internally by vibration device 204. In some embodiments, camera system 100 may implement an actuator that can vibrate in response to signal VPTN. In some embodiments, vibration device 204 may present signal VPTN to a vibration platform. In the example, vibration device 204 may be a separate device including a platform on which camera system 100 is mounted. The implementation of vibration device 204 may vary depending on the design criteria of a particular implementation.

[0125] The VPTN signal can enable a vibration mode. The vibration mode can be configured to jitter, vibrate, and / or actuate the camera system 100. In some embodiments, the vibration mode generated by the vibration device 204 can be a predefined mode (e.g., a mode stored in memory that may be known in advance). In some embodiments, the vibration mode generated by the vibration device 204 can be random and / or pseudo-random. The vibration mode can be implemented to cause instability when the capturing device captures pixel data at various levels of the optical zoom 202 of the lens 160. For example, electronic image stabilization implemented by the camera system 100 can be configured to counteract and / or correct the vibration mode generated by the vibration device 204. Calibration operations performed by the processor 102 and / or memory 150 can be configured to determine calibration values ​​in response to measurements performed when the vibration device 204 generates the vibration mode VPTN. The camera system 100 may not know the vibration mode VPTN in advance. The type and / or specific mode of the vibration mode VPTN applied to the camera system 100 can vary according to the design criteria of a particular implementation.

[0126] Coordinate system 200 is shown as dashed arrows 210, 212, and 214. Arrow 210 may represent the X-axis. Arrow 212 may represent the Y-axis. Arrow 214 may represent the Z-axis. Camera system 100 is shown at the origin of coordinate system 200. Motion can be applied to camera system 100, which can result in motion MTN. For example, forces of various magnitudes can be applied to camera system 100 along axes 210-214.

[0127] Curved arrow 220 is shown as rotation about X-axis 210. Curved arrow 220 can represent roll rotation. Curved arrow 222 is shown as rotation about Y-axis 212. Curved arrow 222 can represent pitch rotation. Curved arrow 224 is shown as rotation about Z-axis 214. Curved arrow 224 can represent yaw rotation. The combination of motion MTN applied to camera system 100 can give camera system 100 roll rotation 220, pitch rotation 222, and / or yaw rotation 224. IMU 106 can be configured to detect various roll rotations 220, pitch rotations 222, and / or yaw rotations 224.

[0128] Curve 230 is shown on the X-axis 210. Curve 230 can represent vibration on the X-axis 210. Vibration 230 can be a type of motion applied to the camera system 100, which can be measured by the IMU 106. Curve 230 is shown as a sine curve with frequencies 232a-232n and amplitude 234. The frequencies 232a-232n and amplitude 234 can represent the components of movement and / or vibration that cause the rolling rotation 220.

[0129] Curve 236 is shown on the Y-axis 212. Curve 236 can represent vibration on the Y-axis 212. Vibration 236 can be a type of motion applied to the camera system 100, which can be measured by the IMU 106. Curve 236 is shown as a sine curve with frequencies 238a-238n and amplitude 240. The frequencies 238a-238n and amplitude 240 can represent the components of movement and / or vibration that cause pitch rotation 222.

[0130] Curve 242 is shown on the Z-axis 214. Curve 242 can represent vibration on the Z-axis 214. Vibration 242 can be a type of motion applied to the camera system 100, which can be measured by the IMU 106. Curve 242 is shown as a sine curve with frequencies 244a-244n and amplitude 246. The frequencies 244a-244n and amplitude 246 can represent the components of movement and / or vibration that cause yaw rotation 224.

[0131] IMU 106 can convert the frequencies 232a-232n and amplitude 234 of vibration 230 for roll rotation 220, the frequencies 238a-238n and amplitude 240 of vibration 236 for pitch rotation 222, and / or the frequencies 244a-244n and amplitude 246 of vibration 242 for yaw rotation 224 into a motion information signal M_INFO. Vibrations 230-242 can cause motion (e.g., jitter) in the captured pixel data VIDEO. Processor 102 can be configured to perform compensation to counteract this motion to generate a stabilized video frame VIDOUT. Overall, the impact of vibrations 230-242 on the amount of motion visible in the raw video data can increase with the amount of optical zoom 202. For example, at large optical zoom levels (e.g., above 10x), even small movements can appear as large movements captured in the video frame.

[0132] During calibration, vibration device 204 can apply a vibration pattern VPTN to camera system 100. For example, during calibration, the input MTN read by IMU 106 can be the vibration pattern VPTN generated by vibration device 204. IMU 106 can be configured to read and / or store the vibration pattern sensed when the vibration pattern VPTN is applied. For example, vibration device 204 can be a calibration jitter configured with various jitter amplitudes and frequencies. The vibration pattern VPTN can include combinations of frequencies 232a-232n and amplitude 234 in roll rotation 220, frequencies 238a-238n and amplitude 240 in pitch rotation 222, and / or frequencies 244a-244n and amplitude 246 in yaw rotation 224. Typically, in response to detecting the vibration pattern VPTN (or other motion input MTN), IMU 106 can detect and / or record the maximum amplitude (degrees), actual amplitude (degrees), actual angle value (degrees), vibration frequency (Hz), and / or vibration duration (seconds). The processor can be configured to associate motion information captured by IMU106 during calibration with video data captured by processor102 (e.g., based on timestamp information and / or the duration of vibration patterns).

[0133] refer to Figure 5 A graph illustrating the total compensation for a range of optical zoom ratios is shown. Graph 250 is shown. Graph 250 may include axes 252 and 254. Axis 252 may be the X-axis. X-axis 252 may show the amount of zoom ratio and / or EFL. Axis 254 may be the Y-axis. Y-axis 254 may represent the amount of final compensation (e.g., FC) for the performed image stabilization. Graph 250 may show the compensation ratio for electronic image stabilization at different zoom levels.

[0134] Graph 250 may include curves 256 and 258. Curve 256 may represent the amount of compensation performed solely in response to image stabilization compensation (e.g., based on lens projection functions and motion information). Image stabilization compensation curve 256 is typically presented as linear. For example, the amount of image stabilization compensation increases linearly with increasing zoom ratio. A linear increase in image stabilization compensation may not be sufficient to address larger optical zoom levels.

[0135] Curve 258 can represent the amount of compensation (e.g., total compensation) performed in response to both image stabilization compensation and additional compensation (e.g., based on calibration values). The additional compensation curve 258 is typically shown as non-linear. For example, the amount of additional compensation can increase non-linearly as the zoom ratio increases. This non-linear increase in additional compensation can accurately address distortion caused by larger optical zoom levels.

[0136] Image stabilization compensation performed by processor 102 can implement one or more lens projection functions. In one example, the lens projection function can be an isometric lens projection function (e.g., an f-theta model). The isometric lens projection function can be determined according to an equation (e.g., EQ1):

[0137] EQ1: c = f * Θ

[0138] In another example, the lens projection function can be a stereo lens projection function (e.g., a custom distortion model). The stereo lens projection function can be determined according to an equation (e.g., EQ2):

[0139] EQ2: c = 2 * f * tan(Θ / 2)

[0140] In yet another example, the lens projection function can be the pinhole lens projection function. The pinhole lens projection function can be determined according to an equation (e.g., EQ3):

[0141] EQ3: c = f * tan(Θ)

[0142] In yet another example, the lens projection function can be pre-calculated for various values ​​of Θ and stored in a lookup table in memory 150. The value of Θ can represent the angle of incidence for each pixel of the image. Image stabilization compensation performed by processor 102 can determine the result of the lens projection function to provide image stabilization compensation. The specific lens projection function implemented and / or the method for determining the result of the lens projection function can vary depending on the design criteria of the particular implementation.

[0143] Additional compensation can be added to image stabilization compensation for any amount of optical zoom 202 (e.g., from 1x zoom to the maximum zoom level). Typically, for larger optical zooms, the amount of compensation can be outside the range of the general lens projection function. For example, the larger the optical zoom 202, the greater the weighting applied by the additional compensation. In one example, a large optical zoom ratio could be greater than approximately 10x zoom. In another example, a large optical zoom ratio could be when the EFL is greater than 50mm. The amount of optical zoom ratio and / or EFL that can be considered a large optical zoom can vary depending on the design criteria of the specific implementation.

[0144] Points 260a and 260b are shown on image stabilization compensation curve 256. Points 260a-260b can be compensation amounts (e.g., CA-CB, respectively). The compensation amount CA-CB can represent the contribution of image stabilization compensation generated by the processor 102 based on the lens projection function and motion information. The compensation amount CA and point 260a can correspond to a zoom level of 28x. The compensation amount CB and point 260b can correspond to a zoom level of 31x. Compared to the compensation CA corresponding to the difference between the zoom level of 31x and the zoom level of 28x, the compensation amount CB can increase linearly.

[0145] Points 262a and 262b are shown on the additional compensation curve 258 above the corresponding points 260a-260b. The additional compensation amount RA is shown extending from point 260a to point 262a, and the additional compensation amount RB is shown extending from point 260b to point 262b. The additional compensation amount RA-RB can represent the contribution of the additional compensation generated by the additional compensation performed by the processor 102 based on the calibration value. Points 262a-262b can be the total compensation amount (e.g., CA+RA and CB+RB, respectively). The additional compensation amount RA and point 262a can correspond to a zoom level of 28x. The additional compensation amount RB and point 262b can correspond to a zoom level of 31x. Compared to the additional compensation amount RA corresponding to the difference between the zoom level of 31x and the zoom level of 28x, the additional compensation amount RB can increase non-linearly. The non-linear relationship between the increase in optical zoom and the additional compensation amount can enable the final compensation to accurately compensate for the distortion caused by the larger optical zoom ratio.

[0146] The total compensation can be expressed by an equation (e.g., EQ4):

[0147] EQ4: Total compensation = R(k1*EFL*r,k2*h,k3*r^2)

[0148] Equation EQ4, used to determine the total compensation, can be calculated through a combination of image stabilization compensation and / or additional compensation modules. The value of EFL*r can be determined in response to the optical zoom ratio and / or pixel offset. The optical zoom ratio can be determined in response to an optical zoom reading from lens 160. For example, the optical zoom ratio can be converted to an effective focal length (e.g., expressed in pixel unit values). Processor 102 can implement a zoom driver configured to read from a zoom lens motor that adjusts the zoom level of lens 160. The zoom lens motor can present the zoom level and / or changes in zoom level to the zoom driver, and the zoom driver can convert the optical zoom ratio 202 into an EFL value expressed in pixel unit values ​​in real time. Due to the optical path affected by zoom, the pixel offset can be determined based on ideal geometric distortion projection. The pixel offset amount can be determined based on the distortion caused by zoom and / or the shape of lens 160. The pixel offset can be an additional pixel offset from the ideal geometric distortion projection, which can be caused by the effect of zoom on the optical path. Pixel offset can be a result of the shape of lens 160 and / or the wide-angle effect of lens 160. The amount of pixel offset can be determined in response to performing image stabilization compensation (e.g., based on lens projection function and motion information). The value h can represent a frequency. This frequency can be an external vibration frequency that can be generated in response to analysis of IMU 106. The radius can be a radius measurement indicating the distance to the center of the image. The radius can be determined based on a point at the optical center of lens 160, which depends on the focal length f and the angle of incidence Θ. In one example, the radius can be determined according to the equation LR = f(Θ).

[0149] Values ​​k1, k2, and k3 can be calibration values. A self-designed calibration system can be used to perform the calibration operation to determine the calibration values. Calibration values ​​k1, k2, and k3 can be scalar values. In one example, calibration value k1 can be a scalar value for pixel offset and optical zoom, calibration value k2 can be a scalar value for motion information, and calibration value k3 can be a scalar value for image center distance. Details regarding pixel offset, optical zoom, and / or image center distance can be described in connection with U.S. Patent Application No. 18 / 602,416, filed March 12, 2024, the appropriate portions of which are incorporated herein by reference.

[0150] The EIS implemented by processor 102 may include contributions from two components (e.g., image stabilization compensation and additional compensation). Image stabilization compensation (e.g., c) may be a function of lens geometric distortion projection and / or vibration modes. Lens geometric distortion can be determined based on various different lens optical projection designs. For example, various lens optical projection designs are shown as EQ1-EQ3. Other lens optical projection designs (e.g., c = 2*sin(Θ / 2)) can be implemented. Other more complex lens optical projection designs can be implemented. For example, lookup table 292 in memory 150 may be implemented to describe geometric distortion compensation for lens optical projections from a specific point to the optical center of lens 160 at different angles and / or distances. Motion information may include the frequency 232a-232n and amplitude 234 of vibration 230 for roll rotation 220, the frequency 238a-238n and amplitude 240 of vibration 236 for pitch rotation 222, and / or the frequency 244a-244n and amplitude 246 of vibration 242 for yaw rotation 224.

[0151] Additional compensation (e.g., r) can be determined based on the inherent behavior of the lens and / or motion information. Total compensation for EIS can include a variable ratio of additional compensation to image stabilization compensation. This ratio can vary at different zoom values ​​and / or different distances. As the zoom value increases, the ratio of additional compensation to image stabilization compensation can increase (e.g., non-linearly). Additional compensation can include a combination of several factors. One factor can be the zoom value. Another factor can be the additional pixel offset relative to ideal geometric distortion projection due to the optical path caused by the zoom effect. Yet another factor can be the distance from a point (e.g., pixel position) to the optical center of the lens 160. A further factor can be motion information. The contribution of each of these factors can be determined by calibration values.

[0152] refer to Figure 6 A graph illustrating additional compensation for a range of optical zoom ratios is shown. Graph 280 is shown. Graph 280 may include axes 282 and 284. Axis 282 may be the X-axis. X-axis 282 may show the amount of zoom ratio and / or EFL. Axis 284 may be the Y-axis. Y-axis 284 may represent the amount of additional compensation (e.g., r) for the electronic image stabilization performed. Graph 280 can show the compensation ratio for different zoom lens suppliers even at the same zoom level. Since each lens supplier can have a different compensation ratio, calibration can be performed on each camera lens during manufacturing.

[0153] Graph 280 may include curves 290a-290c. Curves 290a-290c may represent additional compensation determined for different types of zoom lenses (e.g., zoom lens A, zoom lens B, and zoom lens C). Additional compensation curves 290a-290c may represent filler values ​​for different zoom lenses (e.g., filler for additional compensation to be added to image stabilization compensation). In one example, zoom lens A-zoom lens C may represent various large optical zoom lens products (e.g., Thorlab zoom lenses, Navitar zoom lenses, Opteka zoom lenses, Canon EF-S55 lenses, etc.). The specific lens implemented may vary depending on the design criteria of a particular implementation.

[0154] Each of the additional compensation curves 290a-290c can increase at a different rate depending on the inherent characteristics of the zoom lens. For example, additional compensation curve 290c for zoom lens C may have the highest value of additional compensation, and additional compensation curve 290a for zoom lens A may have the lowest value of additional compensation. Generally, at low values ​​of optical zoom 202 (e.g., low range of optical zoom ratios), the amount of additional compensation from additional compensation curves 290a-290c is negligible. In the example shown, a negligible value for the additional compensation curves may correspond to a zoom value of approximately 10x zoom. However, for each lens, there may not be a specific threshold zoom value where the contribution from additional compensation is negligible (e.g., some lenses may have a large amount of additional compensation at 10x zoom and lower, which may be non-negligible). Typically, a negligible contribution to additional compensation can be a comparison relative to the amount of compensation provided by image stabilization compensation (e.g., from a universal lens projection function). Generally, each zoom ratio may have a specific compensation factor. Each lens can have a unique additional compensation curve (for example, even if some of the additional compensation values ​​may overlap with the additional compensation curves for other types of zoom lenses).

[0155] In one example, zoom lens curve 290b can correspond to... Figure 5 The associated curve 258. Dashed lines 292 and 294 are shown. Dashed line 292 could represent an additional compensation amount (e.g., RA) for a zoom lens B with a zoom ratio of 28x. Similarly, dashed line 294 could represent an additional compensation amount (e.g., RB) for a zoom lens B with a zoom ratio of 31x. For example, dashed line 292 could correspond to... Figure 5 The associated RA values ​​between points 260a and 262a are shown, and the dashed line 294 can correspond to the values ​​shown. Figure 5 The associated RB value is shown between points 260b and 262b.

[0156] In general, for small zoom ratios, the additional compensation factor may be very small (e.g., negligible, close to zero, relatively small compared to image stabilization compensation values, etc.). Typically, for larger zoom ratios, the additional compensation factor may be a more prominent value in the final (or total) compensation. For example, as the optical zoom ratio increases by 202, the additional compensation can be the dominant factor in EIS correction. In one example, a large zoom ratio could be a value greater than 10x. Figure 7 In a related example, at approximately 31x zoom, additional compensation can begin to dominate the total compensation. While graph 280 shows the upper zoom ratio of 31x, additional compensation (e.g., and equation EQ4) can be applied to even larger zoom ratios. The additional compensation curves 290a-290c can increase non-linearly. In one example, the increase in additional compensation curves 290a-290c can be an exponential function (e.g., e^x). In another example, the increase in additional compensation curves 290a-290c can be a power function or a cubic function (e.g., x^2, x^3). The type of non-linear increase in additional compensation curves 290a-290c can vary depending on the design criteria of a particular implementation.

[0157] In general, additional compensation can provide an additional amount of compensation that can be related to the amount of optical zoom 202 and / or the inherent characteristics of lens 160. For example, curves 290a-290c can represent additional compensation for different lens types. Curves 290a-290c can be represented by fitted lines up to specific values ​​of optical zoom 202. In one example, a slope of 1 on the fitted line can be considered a threshold for large or low compensation amounts. For example, when the slope of the fitted line in one of curves 290a-290c is greater than 1, the compensation amount can be considered large or high. In another example, when the slope of the fitted line in one of curves 290a-290c is less than 1, the compensation amount can be considered small or low. The specific slope values ​​that can be used as thresholds for low or high compensation amounts can vary depending on the design criteria of the particular implementation.

[0158] refer to Figure 7 This diagram illustrates a high zoom ratio lens calibration for electronic image stabilization. A calibration component 300 is shown. The calibration component 300 may include a camera system 100, playback recording 302, calibration operation 304, and / or calibration value 306. The calibration component 300 may include other components (not shown). In some embodiments, the calibration component 300 may also include a vibration device 204. The number, type, and / or arrangement of the components of the calibration component 300 may vary according to the design criteria of a particular implementation.

[0159] Camera system 100 is shown as including lens 160, button 310, and / or block (or circuitry) 312. Button 310 is shown as receiving a signal (e.g., INPUT). The signal INPUT can be a physical input provided by a user (e.g., a technician / engineer performing calibration of camera system 100). In the example shown, button 310 can be a physical button. In some embodiments, button 310 can be a software button (e.g., implemented on a touchscreen interface). In some embodiments, the signal INPUT can be provided electronically. For example, calibration operation 304 can provide the signal INPUT to begin calibration. The implementation of button 310 and / or the initiation of calibration can vary depending on the design criteria of a particular implementation.

[0160] Camera system 100 can be configured to generate signals (e.g., CVAL) and / or generate signals (e.g., ASFREC). The CVAL signal may include a calibration value generated by calibration operation 304. The ASFREC signal may include a playback recording generated by circuitry 312. In some embodiments, the ASFREC signal may be presented to a device external to camera system 100. For calibration operation 304, the ASFREC signal may be used internally by camera system 100. In some embodiments, the CVAL signal may be presented to a device external to camera system 100. For calibration operation 304, the CVAL signal may be used internally by camera system 100. The amount, type, and / or format of signals generated and / or received by camera system 100 may vary according to the design criteria of a particular implementation.

[0161] Circuit 312 may include one or more of the components implemented by camera system 100 (e.g., with...). Figure 3 (Associated components shown). In one example, circuit 312 may include a circuit board. In another example, circuit 312 may be a SoC. In the example shown, circuit 312 may be an SoC implementing processor 102 and / or memory 150. SoC 312 may be configured to perform EIS by implementing total compensation (e.g., a combination of image compensation stabilization and additional compensation as shown in Equation EQ4). In the example shown, circuit 312 may include processor 102 and / or memory 150, and IMU 106 is implemented separately. In some embodiments, circuit 312 may be an SoC implementing each of processor 102, memory 150, and / or IMU 106. The number, type, and / or arrangement of components in circuit 312 may vary depending on the design criteria of a particular implementation.

[0162] SoC 312 is shown as including calibration operation 304 and / or block (or circuitry) 320. Calibration operation 304 and / or block 320 may each include computer-executable instructions. The computer-executable instructions of block 320 can implement an emulation framework when executed by processor 102. In one example, emulation framework 320 may be a proprietary emulation framework dedicated to processor 102. The combination of processor 102 and memory 150 can implement the computer-executable instructions for calibration operation 304 and / or emulation framework 320. For example, the computer-executable instructions may be stored in memory 150, read from memory 150 by processor 102, and executed by processor 102. SoC 312 may include other computer-readable data and / or computer-executable instructions.

[0163] The simulation framework 320 can be configured to generate playback recordings. For example, the simulation framework 320 can generate data transmitted in the ASFREC signal. The simulation framework 320 can run / operate on (e.g., be performed by) the camera system 100. The simulation framework 320 can be configured to record motion data that provides vibration context for video frames. The simulation framework 320 can be configured to package and / or format the playback recording data in a format that can be recognized by the calibration operation 304. In some embodiments, the simulation framework 320 implemented by the camera system 100 may depend on the calibration operation 304 for the calibration component 300.

[0164] The simulation framework 320 can be configured to synchronize video frames generated by the processor 102 with metadata. The metadata may include at least motion information captured by the IMU 106. The metadata may also include data such as the exposure shutter timing of the image sensor 180 and / or the start timing of the exposure of the image sensor 180. For example, during calibration, the vibration device 204 may generate a vibration mode VPTN, while the camera system 100 generates video frames and / or reads motion information. The simulation framework 320 can synchronize the video frames generated by the processor 102 during the vibration mode VPTN with the motion data of the vibration mode VPTN generated by the IMU 106. The synchronized video frames and motion information can be recorded as playback recordings.

[0165] Playback record 302 is shown. In the example shown, a single playback record 302 is shown for illustrative purposes. In general, the simulation framework 320 can generate more than one playback record. Playback record 302 can be a recording phase for calibration operation 304. During the recording phase, camera system 100 can be configured with a zoom lens 160 up to a specific zoom ratio (e.g., Nx zoom ratio) and then begin recording metadata and video frames. For example, a playback record can be generated for each zoom level (e.g., where Nx iterates from 1x zoom to the maximum zoom level). Playback record 302 can be stored in memory 150. For example, memory 150 can store playback record 302 for later use by calibration operation 304 performed by SoC 312. Playback record 302 can include blocks (or computer-readable data) 330 and / or blocks (or computer-readable data) 332. Block 330 can implement a video sequence. Block 332 can implement metadata. Playback record 302 can include other data (not shown). The type and / or format of the data in replay record 302 can vary depending on the design standards of a particular implementation.

[0166] Video sequence 330 may include video frames 340a-340n. Video frames 340a-340n of video sequence 330 may include multiple video frames captured during the vibration mode VPTN. For example, the number of video frames 340a-340n captured in video sequence 330 may depend on the length of the vibration mode VPTN. In one example, the number of video frames 340a-340n in video sequence 330 may be at least 600 video frames. When the vibration mode VPTN is complex, more video frames 340a-340n may be captured. Generally, the more complex the vibration in the vibration mode VPTN, the more video frames 340a-340n can be captured in video sequence 330 for playback recording 302. In one example, each of video frames 340a-340n may include a YUV frame. In another example, each of video frames 340a-340n may include a compressed video frame. For example, video frames 340a-340n may include high-quality HEVC I frames. In another example, video frames 340a-340n may include AV1 encoded frames. In yet another example, video frames 340a-340n may include x264 encoded frames. The number and / or formation of video frames 340a-340n in video sequence 330 may vary according to the design criteria of a particular implementation.

[0167] Playback recording 302 may include video sequence 330, wherein metadata 332 is synchronized with each of video frames 340a-340n. Metadata 332 may include blocks (or computer-readable data) 342a-342n. In the example shown, block 342a may include motion information, block 342b may include image sensor exposure shutter timing, block 342c may include exposure start timing, and block 342n may include other data. The type of data stored in metadata 332 may vary depending on the design criteria of a particular implementation.

[0168] Motion information 342a can be generated by IMU 106. For example, motion information 342a may include a sequence of gyroscope sample data (e.g., real-time gyroscope data). Motion information 342a may also include the gyroscope output data frequency and / or the full scale range of the raw gyroscope data. Motion information 342a can be generated during vibration mode VPTN and synchronized to video sequence 330 by simulation frame 320. Image sensor exposure shutter timing 342b and / or exposure start timing 342c may include sensor timing data for capturing video frames 340a-340n during vibration mode VPTN. Sensor timing data can be synchronized to video sequence 330 by simulation frame 320. Other data 342n may include other types of metadata (e.g., the resolution of video sequence 330, which may be the same as the original image resolution).

[0169] The playback recording 302 may be stored in memory 150 and / or used for calibration operation 304. In some embodiments, calibration operation 304 may be performed by SoC 312 (e.g., processor 102). In some embodiments, calibration operation 304 may be performed by a separate computing device. In one example, the playback recording 302 may be transferred to a separate computing device (e.g., desktop computer, laptop computer, smartphone, tablet computing device, cloud computing device, etc.) configured to perform calibration operation 304. For example, the calibration device configured to perform calibration operation may include a CPU and / or memory. The CPU of the external calibration device may be configured to receive and / or analyze data and make decisions in response to input. In one example, the CPU may implement one of a 32-bit instruction set (e.g., x86), a 64-bit instruction set (e.g., AMD64), an ARM instruction set, a RISC-V instruction set, etc. Memory may store data. The memory of an external calibration device can be implemented using various types of memory, including but not limited to cache, flash memory, memory cards, random access memory (RAM), dynamic RAM (DRAM), etc., configured to store the operating system (OS) and programs / applications (e.g., applications, dashboard interfaces, etc.). These applications can operate on various operating system platforms (e.g., Windows, iOS, Linux, Android, Windows Phone, macOS, Chromium OS / Chrome OS, Fuchsia, Blackberry, etc.). Overall, the SoC 312 can include sufficient computing resources to perform calibration operations 304 locally (e.g., without a separate calibration device).

[0170] Calibration operation 304 can be performed locally on camera system 100. Calibration operation 304 can be a playback stage. During playback, on-camera calculations can be run to calculate the checkerboard position based on captured video frames 340a-340n and a pixel offset difference table, followed by curve fitting to obtain the final result. In one example, calibration operation 304 can be configured to implement a companion application that can be configured to interact with camera system 100. This companion application allows users to view video captured by edge devices 100a-100n (e.g., video directly from edge devices 100a-100n and / or streaming via a cloud service). Calibration operation 304 can include blocks (or circuits) 350, 352, 354, and / or 356. Circuit 350 can implement a video processing pipeline. Circuit 352 can implement a pixel difference module. Circuit 354 can implement a curve fitting module. Circuit 356 can implement an evaluation tool. Calibration operation 304 may include other components (not shown). For example, one or more of components 350-356 may be implemented by a combination of hardware modules 190a-190n. The number, type, and / or arrangement of the components of calibration operation 304 may vary depending on the design criteria of a particular implementation.

[0171] Video processing pipeline 350 can be configured to analyze and / or manipulate video frames 340a-340n in video sequence 330. In one example, video processing pipeline 350 can be an iDSP of processor 102. The iDSP implementing video processing pipeline 350 can receive playback recording 302. The iDSP implementing video processing pipeline 350 can be configured to parse and play back video sequences 340a-340n and metadata 332. Video processing pipeline 350 can be configured to detect objects (e.g., calibration targets). Video processing pipeline 350 can be configured to detect intersections (e.g., golden coordinates, true coordinates, etc.) on calibration targets. For example, video processing pipeline 350 can determine the coordinates in 3D space of a grid displayed on the calibration target.

[0172] Pixel difference module 352 can be configured to generate a pixel difference matrix. Pixel difference module 352 can be configured to compare and / or determine the difference between the position of a grid determined by video processing pipeline 350 and the position of a grid in video sequence 330 that has achieved image compensation stabilization (but not additional compensation). For example, the pixel difference matrix generated by pixel difference module 352 can represent the amount of additional compensation that can be applied to EIS.

[0173] Curve fitting module 354 can be configured to generate calibration values. Curve fitting module 354 can be configured to analyze the pixel difference matrix to determine the amount of additional compensation potentially desired to achieve accurate electronic image stabilization. Based on the potentially desired amount of additional compensation, curve fitting module 354 can determine calibration values. Curve fitting module 354 can be configured to implement various types of curve fitting techniques to solve multivariate equations to determine calibration values. A combination of variables fitted to the curve based on the pixel difference matrix for additional compensation can be used as calibration values.

[0174] Evaluation tool 356 can be configured to test the accuracy of the determined calibration values. In one example, evaluation tool 356 can apply calibration values ​​to video frames 340a-340n of playback record 302. Evaluation tool 356 can determine whether the determined calibration values ​​result in the desired image-stabilized video frame output. For example, evaluation tool 356 can use the applied calibration values ​​to analyze playback record 302 to determine the accuracy of the stabilized video frames when a total calibration is applied. In some embodiments, new video frames can be captured for use by evaluation tool 356. If evaluation tool 356 determines that the calibration values ​​do not provide accurate calibration, the calibration process can be repeated. If evaluation tool 356 determines that the calibration values ​​do provide accurate calibration, calibration operation 304 can present the signal CVAL to the memory 150 of camera system 100 (e.g., calibration for a specific zoom level may have been completed, while calibration for another zoom level can be performed and / or calibration for another of camera systems 100a-100n can be initiated).

[0175] Calibration value 306 can be stored in memory 150. Calibration value 306 can be used to determine the total compensation for EIS at various zoom levels. Calibration value 306 can include blocks 380a-380n. Each of blocks 380a-380n can include a set of calibration values ​​for a specific zoom ratio. Each of the calibration value sets 380a-380n can include at least scalar values ​​k1, k2, and k3. For each zoom ratio implemented by lens 160, a set of fixed values ​​can be stored in each of the calibration value sets 380a-380n. The fixed calibration values ​​in each of the calibration value sets 380a-380n can be unique for each specific zoom ratio (e.g., they are determined individually, but some values ​​can be the same at different zoom levels). For example, calibration operation 304 can be iterated once for each zoom level (e.g., from 1x to nx zoom ratio levels) to determine a unique set of calibration values ​​380a-380n. In one example, calibration value set 380a may correspond to calibration values ​​for a 1x zoom ratio, calibration value set 380b may correspond to calibration values ​​for a 2x zoom ratio, calibration value set 380c may correspond to calibration values ​​for a 3x zoom ratio, and so on. Each of the calibration value sets 380a-380n may include a set (and only one set) of calibration values ​​k1, k2, and k3 for the corresponding zoom level. The number of calibration value sets 380a-380n may depend on the number of zoom levels achieved by lens 160. The number of calibration value sets 380a-380n may vary according to the design criteria of a particular implementation.

[0176] refer to Figure 8 A diagram illustrating a camera system for capturing a calibration target is shown. A calibration environment 400 is shown. Calibration environment 400 can be used to calibrate camera systems 100a-100n during mass production. Calibration environment 400 can implement calibration techniques that can be used to determine calibration variable 306.

[0177] The calibration environment 400 may include camera system 100i, vibration device 204, vibration 402, and / or calibration targets 410a-410n. The calibration environment 400 may include one of the camera systems 100a-100n with lenses 160 supported by vibration device 204. Vibration 402 may correspond to the vibration mode VPTN generated by vibration device 204. Camera system 100i may be a representative example of camera systems 100a-100n. For example, each of the camera systems 100a-100n may be individually calibrated in the calibration environment 400 during mass production.

[0178] In the example shown, camera system 100i is depicted mounted on vibration device 204 and aimed at calibration targets 410a-410n. Typically, during calibration, each of calibration targets 410a-410n can be positioned within the field of view of lens 160. During mass production, personnel (e.g., technicians / engineers) can move camera systems 100a-100n to specific distances from calibration targets 410a-410n so that the field of view can capture all calibration targets 410a-410n at multiple predetermined zoom levels. Camera system 100i mounted on vibration device 204 may represent one of camera systems 100a-100n configured for calibration.

[0179] Vibration device 204 is shown supporting camera system 100i and applying vibration 402 to lens 160 at different distances (e.g., FDA-FDN) and angles (e.g., FAA-FAN) from the respective calibration targets 410a-410n. In the example, the distance FDA-FDN can be at least 1.5 meters from the calibration targets 410a-410n. Vibration device 204 can be placed relative to the calibration targets 410a-410n at a fixed distance FDA-FDN and a fixed orientation (e.g., a fixed angle FAA-FAN). For calibration techniques, the position of camera system 100i can remain unchanged. Not moving the position of camera system 100i can save time for calibration techniques. For example, simulation frame 320 can enable playback recording 302 to contain video sequence 330, which includes video frames 340a-340n captured at various predefined zoom levels (e.g., iterating through each zoom level from 1x to nx) without readjusting the position of camera system 100i.

[0180] The calibration target 410b shown includes calibration patterns 412aa-412nn. The calibration patterns 412aa-412nn shown on calibration target 410b can be representative examples of calibration patterns 412aa-412nn that can be used for each of calibration targets 410a-410n (e.g., calibration targets 410a-410n can each have the same pattern). In the example shown, calibration patterns 412aa-412nn can include alternating light and dark (e.g., white and black) rectangles (e.g., a checkerboard / chessboard pattern). For example, square 412aa could be a black square located in the upper left corner of calibration target 410b, square 412ba could be a white square located in the top row and adjacent to 412aa to the right, square 412ab could be a white square located in the second row from the top and directly below square 412aa, and so on. The number of squares in calibration patterns 412aa-412nn can provide data points for the calibration results. More rows and columns in calibration patterns 412aa-412nn can provide more grids (e.g., data points). The more data points, the more accurate the calibration value 306 will be. The grid size of calibration patterns 412aa-412nn can be no less than a predefined number of pixels. In one example, regardless of the physical dimensions of calibration targets 410a-410n, the grid size of calibration patterns 412aa-412nn can be no less than 50x50 pixels. For example, the 50x50 grid size limit can be defined based on the standard of the minimum detectable bounding box in pixels used for CNN computer vision operations. Overall, the user can set calibration targets 410a-410n in calibration environment 400 and adjust the distance to FDA-FDN and physical size accordingly.

[0181] The calibration patterns 412aa-412nn can be implemented with various patterns. For example, some libraries (e.g., OpenCV) allow the use of dot patterns instead of checkerboard patterns 412aa-412nn for calibration targets 410a-410n. The calibration patterns 412aa-412nn can be consistent for each of the calibration targets 410a-410n. Generally, the size, dimensions, and / or distance from the camera system 100i of the checkerboard image 412aa-412nn can vary depending on the design criteria of the specific implementation. Generally, the video frames 340a-340n used for the full calibration technique can be YUV and / or RGB images.

[0182] To perform calibration techniques during mass production, personnel can mount camera system 100i to vibration device 204. Each customer (e.g., each camera provider) can request and / or request calibration of zoom lens 160 at different zoom levels (e.g., multiple predetermined zoom levels). In some embodiments, the customer (e.g., lens manufacturer) can provide camera / lens projection type (e.g., lens projection function). Calibration techniques may include performing a calibration phase for EIS correction at each of the predetermined zoom levels. In one example, the customer may provide one of camera systems 100a-100n with a predetermined zoom level ranging from 1x to 31x. In another example, the customer may provide one of camera systems 100a-100n with a predetermined zoom level ranging from 1x to 50x. Each calibration phase may provide video frames 340a-340n corresponding to different zoom levels. For each calibration phase, the optical zoom 202 of lens 160 can be set to different zoom levels, and video frames 340a-340n of the checkerboard pattern 412aa-412nn of calibration targets 410a-410 can be repeatedly captured to determine a set of calibration values ​​380a-380n for a specific zoom level. The captured video frames 340a-340n can be synchronized with the metadata 332 of the vibration mode VPTN (e.g., including at least motion information 342a) via simulation frame 320.

[0183] Calibration techniques can be a time-consuming process. The simulation framework 320 can be implemented to save significant effort for technicians / engineers performing calibration techniques. For example, a user can arrange calibration targets 410a-410n for predetermined zoom levels and then provide an initialization input signal INPUT (e.g., pressing button 310) to enable the simulation framework 320 to initiate the calibration technique. During the calibration technique, the simulation framework 320 can automatically record data for playback of recording 302 (e.g., recording phases) for each predetermined zoom level (e.g., 1x to 30x).

[0184] During the calibration technique, after the set of calibration values ​​has been determined (e.g., after the playback phase), evaluation tool 356 can analyze the EIS performed using calibration values ​​306. If calibration values ​​306 determine that the stabilized video frames are accurate, then calibration values ​​306 can be stored in memory 150. After the set of calibration values ​​380a-380n has been determined for each zoom level, the next camera system 100a-100n can be mounted to vibration device 204 for calibration. If the evaluation result determined by evaluation tool 356 is determined to be insufficient (e.g., the accuracy of the stabilization result is inadequate), it may be necessary to recapture the calibration images and recalculate the updated calibration values ​​306. Calibration operation 304 may require multiple iterations to capture calibration images of calibration targets 410a-410n to generate calibration values ​​that provide accurate compensation for camera system 100i.

[0185] In some embodiments, for calibration operation 304, predefined known values ​​for the positions of the checkerboard patterns 412aa-412nn can be stored in memory 150. In the example, the total number of intersections (e.g., rows and columns) of each of the squares in calibration patterns 412aa-412nn and the distance between the intersections of the squares in calibration patterns 412aa-412nn of calibration targets 410a-410n can be known in advance (and stored). In some embodiments, calibration operation 304 can perform computer vision operations based on playback recording 302 to detect the position of the intersection of each of the squares in calibration patterns 412aa-412nn of calibration targets 410a-410. The predefined known values ​​(or detected values) can be target values ​​for the calibration result (e.g., a gold standard or a perfect value).

[0186] refer to Figure 9 A diagram illustrating a predefined arrangement of calibration targets is shown. An example predefined arrangement 450 is shown. The example predefined arrangement 450 can be shown as a portion of video frame 452. Video frame 452 can be a representative example of one of video frames 340a-340n in playback recording 302. Video frame 452 can represent the field of view captured by lens 160.

[0187] The predefined arrangement 450 may include nine calibration targets 410a-410i. Calibration targets 410a-410i may be captured at different locations within the image view (e.g., field of view) of video frame 452. For example, each of calibration targets 410a-410i may be located at a different distance FDA-FDN and at a different angle FAA-FAN relative to lens 160 to achieve the predefined arrangement 450.

[0188] Each of the calibration targets 310a-310i in the predefined arrangement 450 may include calibration patterns 412aa-412nn. In the example shown, calibration patterns 412aa-412nn may include a 7x5 grid of alternating light / dark squares. Although a 7x5 pattern is shown for illustrative purposes, calibration patterns 412aa-412nn may include more than 5 rows and 7 columns. Generally, the more rows and columns the calibration patterns 412aa-412nn have, the more likely the calibration operation 304 is to produce accurate results.

[0189] The predefined arrangement 450 may include nine calibration targets 310a-310i. These nine targets 310a-310i can be arranged in a 3x3 pattern within video frame 452. In the example shown, calibration target 410a may be the top left target, calibration target 410b may be the top center target, calibration target 410c may be the top right target, calibration target 410d may be the center left target, calibration target 410e may be the center target, calibration target 410f may be the center right target, calibration target 410g may be the bottom left target, calibration target 410h may be the bottom center target, and calibration target 410i may be the bottom right target. The distance between each of the calibration targets 410a-410i can be greater than 1.5m. Larger distances can be used for the predefined arrangement 450. Generally, the larger the optical zoom 202, the larger the distance used for the predefined arrangement 450 to perform accurate calibration. The calibration targets 310a-310i can be arranged at a distance such that the calibration patterns 412aa-412nn can have a unit size of at least 50×50 pixels.

[0190] A white space 454 is shown for video frame 452. The white space 454 is not necessarily white or empty. The white space 454 can represent a portion of video frame 452 that may be outside the space occupied by calibration targets 410a-410i. For example, if a predefined arrangement 450 of calibration targets 410a-410i is circumscribed within a rectangle, then the white space 454 may be outside the circumscribed rectangle. Calibration targets 410a-410i are shown as representing a majority of the available space in video frame 452 (e.g., calibration targets 410a-410i occupy a majority of the field of view of lens 160). Calibration targets 410a-410i may occupy more of video frame 452 than the white space 454. In the example, a portion of the space occupied by calibration targets 410a-410n of video frame 452 (e.g., the field of view and / or VIN domain) may be greater than 75% (e.g., the white space 454 outside the circumscribed rectangle of calibration targets 410a-410i may be less than 25%).

[0191] Each of the calibration targets 410a-410i may have a corresponding orientation 460a-460i. The corresponding orientation 460a-460i may include the orientation of the plane of the calibration targets 410a-410i. A predefined arrangement 450 may include the calibration targets 410a-410i, which are arranged such that the corresponding orientations 460a-460i may be different from each other. In one example, for accurate calibration during the calibration technique, the difference in the corresponding orientations 460a-460i may include an angular difference of more than 10 degrees between the planes of the calibration targets 410a-410i.

[0192] Double-ended arrows (e.g., VDH) and VDV are shown. The VDH arrow can represent the distance between an edge of one of the calibration targets 410f and an edge of the video frame 452 (e.g., the edge of the horizontal field of view). The VDV arrow can represent the distance between an edge of one of the calibration targets 410i and an edge of the video frame 452 (e.g., the edge of the vertical field of view). Distances VDH and VDV can represent the size of the white space 454 surrounding the calibration targets 410a-410i. Although only two distances, VDH and VDV, are shown for illustrative purposes, various distances can exist from each of the calibration targets 410a-410i to the left, right, top, and bottom edges of the video frame 452. In one example, distances VDH and VDV can represent the distance from the bounding rectangle containing the calibration targets 410a-410i to the edge of the video frame 452.

[0193] During the calibration process, vibration device 204 can apply a vibration mode VPTN to camera system 100i. The vibration mode VPTN can cause camera system 100i to shake. This shaking during calibration can cause the predefined arrangement 450 of calibration targets 410a-410i to move around within video frame 452 (e.g., vertical, horizontal, and / or forward / backward translation). The greater the vibration (e.g., roll amplitude 234, pitch amplitude 240, and / or yaw amplitude 246), the more the calibration targets 410a-410i can be offset within the field of view of video frame 452. To ensure accurate calibration, the circumscribed rectangle of calibration targets 410a-410i may need to always be within the field of view. For example, the predefined arrangement 450 can be implemented such that the distances VDH and VDV are large enough to provide a sufficient amount of white space 454 to ensure that none of the calibration targets 410a-410i (or portions of calibration targets 410a-410i) moves outside the field of view of the video frame 452. Overall, the circumscribed rectangle of the predefined arrangement 450 of the calibration targets 410a-410i may need to remain within the VIN domain during vibration modes.

[0194] The distance to the FDA-FDN can be determined based on the optical zoom level in a predefined zoom level used for the calibration technique. The vibration device 204 and camera system 100i can be positioned to ensure that, during the vibration pattern VPTN for each of the optical zoom levels, the checkerboard pattern 412aa-412nn of each of the calibration targets 410a-410i occupies almost the entire field of view of the lens 160, wherein the checkerboard pattern 412aa-412nn has a pixel unit size of at least 50x50. For mass production using the calibration technique, the vibration device 204 and camera system 100i can be manually adjusted once (e.g., initial setup) at the start of the calibration technique for each of the camera systems 100a-100n. Since the capture of playback recording 302 and the calibration operation 304 performed locally on camera systems 100a-100n can be automated processes (e.g., initialized by pressing button 310), the calibration operation 304 can be performed without additional manual interaction after camera system 100i has been installed and calibration targets 410a-410i have been placed in the predefined arrangement 450. For example, the FOV of lens 160 can change accordingly when the zoom ratio changes. As long as the predefined arrangement 450 of calibration targets 410a-410i meets the criteria (e.g., distance, angle, FOV, white space, pixel unit size, etc.), the capture of playback recording 302 and / or the calibration operation 304 can be performed without any personnel changing the settings for calibration targets 410a-410i. If a single arrangement of calibration targets 410a-410i does not meet the criteria, personnel can change the settings after one or more changes in the zoom ratio to meet the calibration criteria.

[0195] refer to Figure 10 A graph illustrating the total compensation achieved by the capture device is shown. An example total compensation curve 500 is shown. The example total compensation curve 500 can be represented as a three-dimensional graph 502. The three-dimensional graph 502 may include axes 504, 506, and / or axis 508. Axis 504 may be the X-axis. Axis 506 may be the Y-axis. Axis 508 may be the Z-axis. The example total compensation curve 500 can be used to determine a calibration value 306 based on the desired total compensation.

[0196] The calibration technique may include applying a vibration mode VPTN to one of the camera systems 100a-100n, while one of the camera systems 100a-100n captures video data of a predefined arrangement of calibration targets 410a-410n for each of the predefined zoom levels. The specific mode applied to the vibration mode VPTN may not significantly affect the result of calibration operation 304. The specific mode applied to the vibration mode VPTN may not need to be known in advance (e.g., playback recording 302 can capture the vibration mode VPTN in real time). Overall, if the vibration mode VPTN is complex, more video frames 340a-3340n can be captured, which can be useful for calculating a more accurate result for calibration value 306. Overall, the technician performing the calibration during manufacturing can find a balance between accuracy and calibration time. A simulation framework 320 implemented by SoC 312 can generate playback recording 302, and playback recording 302 can be used to perform calibration operation 304.

[0197] In response to playback recording 302, calibration operation 304 can be performed by SoC 312. Video processing pipeline 350 can be configured to analyze and / or perform computer vision operations on each of video frames 340a-340n in video sequence 330. In one example, video sequence 330 may include at least 600 video frames 340a-340n (e.g., YUV frames). If the vibration mode VPTN is complex, more than 600 video frames 340a-340n can improve the accuracy used to determine calibration values. In another example, video sequence 340a-340n may include compressed video frames (e.g., high-quality HEVC I-frames) to save bandwidth in the DRAM of memory 150 and / or reduce I / O throughput to external storage devices.

[0198] To perform calibration operation 304, video processing module 350 can be configured to determine the position of calibration patterns 412aa-412nn (e.g., a checkerboard grid) for each of calibration targets 410a-410n in each of video frames 340a-340a. In the example, if calibration patterns 412aa-412nn include a 7x5 grid (as shown in the example), Figure 9 (As shown in the associated diagram), the video processing module 350 can generate a 7x5 matrix for each of the calibration targets 410a-410n in one of the video frames 340a-340n. For example, for each of the calibration targets 410a-410n in each of the video frames 340a-340n (e.g., n = 1, 2, 3, ..., frameID), there can be a grid position matrix. In the example, one of the grid position matrices can be shown in a table (e.g., Table 1):

[0199]

[0200]

[0201] The values ​​in Table 1 can represent the positions of the squares of the calibration pattern 412aa-412nn for the calibration target 310a (e.g., the upper left target in the predefined arrangement 450) of the video frame 340j in the video sequence 330 of the playback record 302. In general, the calibration targets 410a-410n can include more than one 7x5 arrangement of the calibration patterns 412aa-412nn, and Table 1 can include more position values ​​than are shown.

[0202] For each of the video frames 340a-340n in the replay record 302, there may exist a larger combination matrix representing all positions of the squares of the calibration patterns 412aa-412nn for all calibration targets 410a-410n. Figure 9 The example of a predefined 9x9 arrangement 450 shown in association may contain three rows and three columns of 7x5 grid calibration targets 410a-410n, resulting in a larger matrix of (7*3)x(5*3) = 21x15 (e.g., 315 checkerboard grid positions for a single YUV frame). In total, there can be more than 315 checkerboard grid positions (e.g., each of the calibration targets 410a-410n may include a grid larger than 7x5). The video processing module 350 can generate the grid positions for all calibration targets 410a-410n for all video frames 340a-340n in the playback record 302. The grid positions can be determined using computer vision operations. The grid positions can be considered as estimated ground truth values ​​and / or golden coordinate values.

[0203] The pixel difference module 352 can be configured to compare the grid positions generated by the video processing module 350 based on the analysis of video frames 340a-340n with data in the playback record 302. For example, the playback record 302 may include motion information 342a generated by the IMU 106. Based on the playback record 302, the calibration operation 304 can be configured to determine the grid positions for calibration targets 410a-410n in response to an EIS performed using a lens projection function and motion information 342a (e.g., performing image stabilization using contributions from image stabilization compensation, but without additional compensation). In response to a comparison of the EIS performed without additional compensation with the grid position coordinates determined by the video processing module 350 (e.g., partially shown in Table 1), the pixel difference module 352 can generate a pixel difference matrix. Overall, the lens manufacturer can provide information about the camera / lens projection type. After calibration, lens distortion information can be calculated through subsequent curve fitting.

[0204] The pixel difference matrix may include a table of values. The values ​​in the pixel difference matrix may include numerical values ​​indicating the amount by which the result of image stabilization differs from the actual grid positions (e.g., golden coordinate values) of the calibration pattern 412aa-412nn. In one example, the pixel difference matrix may include values ​​expressed in terms of the number of pixels (e.g., pixel unit values). In another example, the pixel difference matrix may include values ​​expressed in millimeters. In yet another example, the pixel difference matrix may include values ​​expressed in centimeters. The pixel difference matrix / table may be calculated by the pixel difference module 352 between the golden coordinates determined based on the checkerboard grid positions for each frame and the image stabilization operation performed without additional compensation. An example pixel difference matrix is ​​shown in the table (e.g., Table 2):

[0205]

[0206] Example pixel differences are shown in the pixel difference matrix in Table 2. Generally, for each of video frames 340a-340n, the same number of pixel differences as in the large matrix of chessboard positions can exist. In the example predefined arrangement 450 (e.g., a 3x3 arrangement with calibration targets 410a-410n including a 7x5 chessboard pattern), 315 pixel differences (e.g., 21x15) can exist in the pixel difference matrix. The number of pixel differences and / or the size of the pixel difference matrix can vary depending on the design criteria of a particular implementation.

[0207] In some embodiments, for a checkerboard calibration pattern 412aa-412nn in an image captured by camera system 100i during calibration, calibration operation 304 may calculate a stabilization difference between the grid positions from the captured data and a reference (e.g., a predefined known value) for the real-world positions of calibration targets 410a-410n. The comparisons provided in the pixel difference matrix used by calibration operation 304 can be used to quantify the inherent characteristics of lens 160. In response to the quantization of the inherent characteristics of lens 160, calibration operation 304 can generate accurate calibration values.

[0208] The curve fitting module 354 can use the pixel difference matrix to determine the calibration value for the total compensation equation. The total compensation can be determined according to equation EQ4. In equation EQ4, variables k1, k2, and k3 can be calibration values ​​306 of one of the sets of calibration values ​​380a-380n for a specific zoom ratio level. The parameter (k1*EFL*r) can be configured to satisfy the pitch / yaw / roll model. The parameter k2*h can be configured to provide the actual conversion from the raw data of motion information 342a (e.g., raw data from IMU 106) to the angle. The parameter k3*r^2 can be the distance from the point to the center of the image. For example, r^2 may be equal to x^2+y^2, which can be simplified to k3*r^2. The value r can be determined according to the length (footage) of the checkerboard grid pattern 412aa-412nn recorded in the playback record 302.

[0209] Curve 510 is shown in 3D plot 502. Curve 510 can plot values ​​in the pixel difference matrix. Plotted points 512a-512n are shown on curve 510. Plotted points can represent values ​​in the pixel difference matrix. Curve fitting module 354 can be configured to evaluate curve points 512a-512n of curve 510 to determine a calibration variable 306 for a specific one in a set of calibration values ​​380a-380n (e.g., for one of the zoom ratio levels). For example, calibration value 306 can be a value that allows observations from the pixel difference matrix to fit equation EQ4. In the example, two values ​​can be determined by fitting curve 510. One variable can be a point (e.g., x, y) of r, which serves as an input value to equation EQ4. The other variable can be the vibration frequency (e.g., h).

[0210] Reference Figure 11 The diagram illustrates a curve fitting to determine calibration values ​​used for additional compensation. An example total compensation curve fit 530 is shown. The example total compensation curve fit 530 may include a three-dimensional curve plot 502. The three-dimensional curve plot 502 may include, as shown in the diagram... Figure 10 The X-axis 504, Y-axis 506, and Z-axis 508 are shown in association. Curve 510 is shown as having curve points 512a-512n.

[0211] The curve fitting module 354 can be configured to determine calibration variables 306 (e.g., k1, k2, and k3 for each of the calibration value sets 380a-380n). The curve fitting calculation performed by the curve fitting module 354 can be a nonlinear fit. In one example, the curve fitting implemented by the curve fitting module 354 can be a polynomial curve fitting utilizing Taylor's theorem.

[0212] A calibration value 306 can be determined at each zoom level in the predefined zoom levels. Examples of calibration values ​​306 determined for predefined zoom levels (e.g., calibration value sets 380a-380n) can be shown in a table (e.g., Table 3):

[0213]

[0214] In general, for any given zoom ratio, there can exist a set of fixed values ​​(e.g., k1, k2, k3) that can be stored as one of the calibration value sets 380a-380n. For each zoom ratio, the calibration values ​​(e.g., k1, k2, k3) can have different values. Graph 530 can represent a curve fit performed to determine one of the calibration value sets 380a-380n at a specific example zoom ratio level (e.g., calibration value 306 for a 20x zoom ratio). Through compensation calculated by equation EQ4, the total compensation (e.g., image stabilization with additional compensation) can produce results very close to the golden position coordinates. For example, if the camera system 100i implements zoom ratio levels from 1x to 30x, the calibration operation 304 can be iterated thirty times (e.g., performing thirty different pixel difference matrices and performing thirty different curve fits) to generate a calibration value set 380a-380n including thirty sets of values ​​for k1, k2, and k3.

[0215] Lines 532a-532n are shown on curve plot 502. Lines 532a-532n can represent curve fitting calculations. Curve fitting calculations 532a-532n can align the results for curve points 512a-512n with curve 510 to determine calibration value 306. Curve fitting calculations 532a-532n can determine calibration values ​​that allow observations from the pixel difference matrix to fit equation EQ4.

[0216] Calibration operation 304 can be configured to evaluate the calibration value 306 determined by curve fitting module 354. Evaluation tool 356 can be configured to test electronic image stabilization by performing total compensation using a set of calibration values ​​380a-380n determined for a specific zoom ratio. For example, using playback recording 302, evaluation tool 356 can apply image stabilization compensation and additional compensation. Evaluation tool 356 can test electronic image stabilization for each zoom level to ensure accuracy.

[0217] Evaluation tool 356 can evaluate the results (e.g., the generated accurate calibration value 306). Evaluation tool 356 can check whether the stabilized video frame provides an accurate result for lens 160. Evaluation tool 356 can output a fitting error result. The fitting error result can be compared with a predetermined accuracy threshold. In the example, the predetermined threshold can be a value within 1 / 16 of a pixel. The predetermined threshold can be determined based on the accuracy of the twisted hardware block implemented by processor 102. For example, 1 / 16 of a subpixel can be the limit of the accuracy of the twisted hardware block, which can set the minimum error for calibration operation 304. After curve fitting, evaluation tool 356 can ensure that the increment between the EIS-compensated pixel position and the golden position calculated from the checkerboard grid position is accurate within the minimum error.

[0218] The fitting error result quantifies the level of error between the actual distortion from the captured image and the distortion calculated according to the fitting function. If the fitting error is close to zero pixels, the calculated fitting function can accurately stabilize the movement based on the intrinsic quality of lens 160 at a predefined zoom level. If the error level is higher than a predetermined threshold, the calibration technique can be repeated. For example, the calibration technique (e.g., using different vibration modes and generating different playback recordings) can be repeated until the fitting error determined by evaluation tool 356 is within a predetermined threshold (e.g., 1 / 16 of a pixel).

[0219] In some embodiments, the evaluation tool 356 can test the accuracy of the results of the calibration operation 304 without capturing new video frames. For example, the video sequence 330 replaying the recording 302 may include sufficient data for the evaluation tool 356. A portion (e.g., the majority) of video frames 340a-340n may be used to determine the calibration value 306, and another portion (e.g., a small portion) of video frames 340a-340n may be reserved for verification using the evaluation tool 356. In one example, 80% of video frames 340a-340n may be used to determine the calibration value 306, and 20% of video frames 340a-340n may be used for verification using the evaluation tool 356.

[0220] Once the evaluation tool 356 determines that the calibration value 306 is within a predetermined accuracy threshold, a calibration operation 304 can be completed for one of the zoom ratio levels. The generated calibration value 306 can be stored by the SoC 312 of the camera system 100i (e.g., in memory 150). For mass production, after determining and storing the calibration value 306 for the camera system 100i (e.g., for all zoom ratio levels), a technician / engineer can perform the calibration operation 304 for the next of the camera systems 100a-100n. For example, the camera system 100i can be removed from the vibration device 204, and the next camera system 100j can be set on the vibration device 204, and the calibration targets 410a-410n can be arranged in a predetermined arrangement 450, wherein the distance and / or angle are suitable for each of the predetermined zoom levels specific to the camera system 100j.

[0221] refer to Figure 12 The diagram illustrates method (or process) 580. Method 580 can perform large zoom ratio lens calibration for electronic image stabilization. Method 580 generally includes steps (or states) 582, 584, 586, 588, 590, 592, 594, 596, a decision step (or state) 598, step (or state) 600, and step (or state) 602.

[0222] Step 582 can initiate method 580. In step 584, the capture device 104 can receive pixel data of the environment and the IMU 106 can receive motion information. For example, the lens 160 can receive light input LIN, and the image sensor 180 can convert the light into raw pixel data (e.g., the signal VIDEO). In parallel, the gyroscope 186 can measure the motion MTN of the camera system 100, and the IMU 106 can convert the measurement into motion information (e.g., the signal M_INFO). Next, in step 586, the processor 102 can process the pixel data arranged into video frames. In some embodiments, in addition to EIS, the processor 102 can also perform other operations on the pixel data arranged into video frames (e.g., perform computer vision operations, calculate depth data, determine white balance, etc.). In step 588, processor 102 and / or memory 150 may generate playback record 302, wherein video sequence 330 (e.g., video frames 340a-340n) captures calibration targets 410a-410n, and metadata 332 includes at least motion information 342a corresponding to vibration mode VPTN. For example, one of camera systems 100a-100n may be connected to vibration device 204, which may apply vibration mode VPTN while SoC 312 of simulation framework 320 generates playback record 302. For example, capturing video frames and metadata and generating playback record 302 may be a recording part of calibration techniques. Next, method 580 may move to step 590.

[0223] In step 590, the processor 102 and / or memory 150 can determine the coordinates of calibration targets 410a-410n based on video frames 340a-340n in the playback record 302. For example, the SoC 312 can perform calibration operation 304. Calibration operation 304 can be a playback portion of the calibration technique. The SoC 312 can perform computer vision operations to determine the position of each grid of the calibration pattern 412aa-412nn. Next, in step 592, the processor 102 can perform image stabilization compensation on the playback record 302. Image stabilization compensation can be based on the lens optical projection function and motion information without additional compensation. In step 594, processor 102 may generate a pixel difference matrix (e.g., a pixel difference table) in response to a comparison of the coordinates of the calibration patterns 412aa-412nn of the calibration targets 410a-410n determined using computer vision (e.g., determined in step 590) with the coordinates of the calibration patterns 412aa-412nn of the calibration targets 410a-410n determined based on image stabilization compensation (e.g., determined in step 592). Next, in step 596, processor 102 may generate calibration values ​​306 for additional compensation in response to performing curve fitting on the pixel difference matrix. For example, the curve fitting may be configured to determine a solution to equation EQ4 based on the pixel difference matrix. Next, method 580 may move to decision step 598.

[0224] In decision step 598, processor 102 may determine whether there are additional zoom levels for lens 160 to be calibrated. For example, calibration operation 304 may be programmed with multiple zoom ratio levels for a specific lens 160 and / or camera model. These multiple zoom ratio levels may be provided by the camera and / or lens manufacturer. If additional zoom ratio levels exist for calibration, method 580 may return to step 584. For example, calibration operation 304 may be performed to determine a set of calibration values ​​for each of the calibration value sets 380a-380n. If no additional zoom ratio levels exist for calibration, method 580 may proceed to step 600. In step 600, processor 102 may implement EIS using total compensation (e.g., a combination of image stabilization compensation and additional compensation using one of the calibration value sets 380a-380n corresponding to the zoom ratio level of the captured video frame). Next, method 580 may proceed to step 602. Step 602 may end method 580.

[0225] refer to Figure 13The diagram illustrates method (or process) 620. Method 620 can perform a calibration operation in response to a playback recording. Method 620 generally includes steps (or states) 622, 624, 626, 628, 630, 632, 634, 636, 638, a decision step (or state) 640, 642, and 644.

[0226] Step 622 can initiate method 620. In step 622, simulation framework 320 can generate playback record 302. Next, in step 626, video processing pipeline 350 (e.g., iDSP) can parse and replay video frames 340a-340n and metadata 332 from playback record 302. Next, method 620 can move to steps 628 and 630. In one example, steps 628-630 can be performed in parallel and / or substantially in parallel. In another example, steps 628-630 can be performed sequentially. If performed sequentially, the order in which steps 628-630 are performed can vary depending on the design criteria of the specific implementation.

[0227] In step 628, the video processing pipeline 350 may perform a computer vision operation to detect the golden coordinates of calibration targets 410a-410n. For example, the golden coordinates may be the estimated true real-world position of each grid 412aa-412nn of each of the calibration targets 410a-410n in one of the video frames 340a-340n. Overall, if applied to video frames 340a-340n, a calibration value 306 can be calculated so that the EIS results achieve the golden coordinates. Next, method 620 may move to step 632. In step 630, the video processing pipeline 350 may use image stabilization to determine the grid positions without additional compensation. For example, the video processing pipeline 350 may apply an image stabilization operation to video frames 340a-340n and then perform a computer vision operation to determine the position of each of the grids 412aa-412nn. Next, method 620 may move to step 632.

[0228] In step 632, the pixel difference module 532 can compare the grid positions determined using image stabilization compensation with the golden coordinates. Next, in step 634, the pixel difference module 532 can generate a pixel difference matrix (e.g., Table 2). The pixel difference matrix can be determined based on the comparison between the grid positions determined using image stabilization compensation and the golden coordinates. In step 636, the curve fitting module 354 can perform curve fitting based on equation EQ4 to determine a specific set of calibration values ​​from the calibration value set 380a-380n corresponding to the current zoom ratio. For example, this set of calibration values ​​may include calibration values ​​k1, k2, and k3 that enable equation EQ4 to provide a solution for the values ​​in the pixel difference matrix. Next, in step 638, the evaluation tool 356 can test the values ​​of k1, k2, and k3. For example, the evaluation tool 356 can apply electronic image stabilization to a subset of video frames 340a-340n in the playback record 302 using total compensation based on a specific set of calibration values ​​from the calibration value set 380a-380n for the current zoom level. Next, method 620 can be moved to decision step 640.

[0229] In decision step 640, evaluation tool 356 can determine whether the accuracy of EIS provides results within a sub-pixel error threshold when using the determined calibration values ​​k1, k2, and k3. For example, the sub-pixel error threshold could be 1 / 16 of a pixel. If the accuracy of EIS is not within the sub-pixel error threshold, method 620 can return to step 626 to repeat calibration operation 304 and recalculate the calibration values. If the accuracy of EIS is within the sub-pixel error threshold, method 620 can move to step 642. In step 642, processor 102 can store the calibration value determined for the current zoom level as one of a set of calibration values ​​380a-380n as calibration value 306. For example, calibration value 306 can be stored in memory 150. Next, method 620 can move to step 644. Step 644 can end method 620.

[0230] refer to Figure 14 The diagram illustrates method (or process) 680. Method 680 can capture a recording of a calibration target during vibration mode for playback of the recording. Method 680 generally includes steps (or states) 682, 684, 686, 688, a decision step (or state) 690, 692, 694, and 696.

[0231] Step 682 can initiate method 680. In step 684, vibration device 204 can apply the vibration mode VPTN to one of the camera systems 100a-100n. Next, in step 686, IMU 106 can measure the vibration mode VPTN as motion input MTN. In step 688, capture device 104 can capture video frames for playback recording 302. Overall, steps 686 and 688 can be captured in parallel (e.g., motion input MTN and pixel data can be captured simultaneously). Next, method 680 can move to decision step 690.

[0232] In decision step 690, simulation framework 320 can determine whether there is a sufficient number of video frames for the complexity of the vibration mode VPTN. Generally, at least 600 of the video frames 340a-340n can be captured for playback recording 302. However, depending on the complexity of the vibration mode VPTN, additional video frames may be captured. If the number of captured video frames is insufficient, method 680 can return to step 688. If the number of captured video frames is sufficient, method 680 can move to step 692. In step 692, simulation framework 320 can synchronize metadata 332 with the video sequence 330. For example, motion information, image sensor exposure shutter timing, exposure start timing, resolution information, etc., can be synchronized with the captured video frames 340a-340n. Next, in step 694, playback recording 302 can be stored in memory 150. Next, method 680 can move to step 696. Step 696 can end method 680.

[0233] In some embodiments, calibration operation 304 may be performed after generating playback record 302 for one of the zoom levels of lens 160. For example, lens 160 may be set to a 1x zoom level, playback record 302 may be generated for the 1x zoom level, and then before changing the zoom level of lens 160 to a 2x zoom level, calibration operation 304 may determine a set of calibration values ​​380a, etc., corresponding to the 1x zoom level. In some embodiments, before performing any of the calibration operations 304, the calibration technique may first capture one of the playback records 302 for all zoom levels. For example, playback record 302 may be generated for the 1x zoom level, then for the 2x zoom level, then for the 3x zoom level, etc., until playback record 302 has been generated for each of the zoom levels. Then, calibration operation 304 may subsequently use the playback record 302 generated for each zoom level to determine all sets of calibration values ​​380a-380n. Whether replay records 302 are generated one at a time or sequentially can vary depending on the design standards of the specific implementation.

[0234] refer to Figure 15 The diagram illustrates method (or process) 720. Method 720 can generate calibration values ​​for a camera during camera manufacturing. Method 720 generally includes steps (or states) 722, 724, 726, 728, 730, 732, a decision step (or state) 734, 736, 738, and step (or state) 740.

[0235] Step 722 can initiate method 720. In step 724, the user (e.g., a technician / engineer / calibrator) can arrange calibration targets 410a-410n in a predefined pattern 450 for the next field of view in camera systems 100a-100n. Ideally, for cameras implementing autofocus, calibration targets 410a-410n can be arranged such that the criteria for calibration operation 304 (e.g., distance, corresponding angle, calibration targets 410a-410n remaining within the field of view, grid size of calibration patterns 412aa-412nn not less than the minimum pixel unit threshold, etc.) can be satisfied for all zoom levels of lens 160. Next, in step 726, vibration device 204 can apply a vibration mode VPTN to one of the camera systems 100a-100n. In step 728, simulation frame 320 can generate a playback recording 302. For example, the user can apply the signal INPUT (e.g., pressing button 310) to initiate the calibration technique. In some embodiments, the entire calibration process can be performed automatically by pressing button 310 (e.g., if a change in the arrangement of calibration targets 410a-410n does not provide any benefit). Next, in step 730, SoC 312 can perform calibration operation 304. In step 732, calibration operation 304 can generate one of the calibration value sets 380a-380n for the current zoom level. Next, method 720 can move to decision step 734.

[0236] In decision step 734, calibration operation 304 can determine whether there are additional zoom levels for lens 160 to be calibrated. For example, if zoom levels from 1x to 30x exist, there can be thirty iterations in steps 726-732 to generate each of the calibration value sets 380a-380n. If more zoom levels are available for calibration, method 720 can move to step 736. In step 736, lens 160 can be set to the next optical zoom level. For example, the zoom lens motor can adjust the zoom level. If the arrangement of calibration targets 410a-410n does not meet the standards of the calibration technique, method 720 can move to step 724 and the calibration targets 410a-410n can be adjusted. If no adjustment to the calibration targets 410a-410n is required, method 720 can return to step 726 (e.g., start the next zoom level iteration to determine a calibration value for a specific one of camera systems 100a-100n). In decision step 734, if no further zoom levels are available for calibration for one of the camera systems 100a-100n, method 720 can move to decision step 738. For example, once all calibration value sets 380a-380n have been determined for one of the camera systems 100a-100n, calibration for one camera can be completed.

[0237] In decision step 738, the user can determine whether more camera systems 100a-100n are available for calibration. The number of camera systems 100a-100n to be calibrated may depend on the number of cameras being manufactured. If more camera systems 100a-100n are available for calibration, method 720 can return to step 724. For example, calibration targets 410a-410n can be recalibrated for the next camera among camera systems 100a-100n to begin calibration for the next camera. If no more cameras 100a-100n are available for calibration, method 720 can proceed to step 740. Step 740 can end method 720.

[0238] Depend on Figure 1-14The functions executed by the diagram can be implemented using one or more of a conventional general-purpose processor, digital computer, microprocessor, microcontroller, RISC (Reduced Instruction Set Computer) processor, CISC (Complex Instruction Set Computer) processor, SIMD (Single Instruction Multiple Data) processor, signal processor, central processing unit (CPU), arithmetic logic unit (ALU), video digital signal processor (VDSP), and / or similar computing machines programmed according to the teachings of this specification, as will be apparent to those skilled in the art. A skilled programmer can readily prepare appropriate software, firmware, code, routines, instructions, opcodes, microcode, and / or program modules based on the teachings of this disclosure, again as will be apparent to those skilled in the art. The software is typically executed from one or more media by one or more processors in a machine-implemented manner.

[0239] The present invention can also be implemented by preparing an ASIC (Application-Specific Integrated Circuit), a platform ASIC, an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), a CPLD (Complex Programmable Logic Device), a sea-of-gate, an RFIC (Radio Frequency Integrated Circuit), an ASSP (Application-Specific Standard Product), one or more monolithic integrated circuits, one or more chips or dies arranged as flip-chip modules and / or multi-chip modules, or by interconnecting conventional component circuits through a suitable network, as described herein, modifications of which will be apparent to those skilled in the art.

[0240] Therefore, the present invention may also include a computer product, which may be one or more storage media and / or one or more transmission media, including instructions that can be used to program a machine to perform one or more processes or methods according to the present invention. The execution of the instructions contained in the computer product by the machine, and the operation of the surrounding circuitry, can convert input data into one or more files on the storage medium and / or one or more output signals representing physical objects or substances, such as audio and / or visual depictions. The execution of the instructions contained in the computer product by the machine can be performed on data stored on the storage medium and / or user input and / or in combination with values ​​generated by a random number generator implemented by the computer product. The storage medium may include, but is not limited to, any type of disk, including floppy disks, hard disks, magnetic disks, optical disks, CD-ROMs, DVDs, and magneto-optical disks, as well as circuits such as ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), UVPROM (Ultraviolet Erasable Programmable ROM), flash memory, magnetic cards, optical cards, and / or any type of medium suitable for storing electronic instructions.

[0241] The elements of this invention can form part or all of one or more devices, units, components, systems, machines, and / or apparatuses. Devices may include, but are not limited to, servers, workstations, storage array controllers, storage systems, personal computers, laptop computers, notebook computers, handheld computers, cloud servers, personal digital assistants, portable electronic devices, battery-powered devices, set-top boxes, encoders, decoders, transcoders, compressors, decompressors, preprocessors, post-processors, transmitters, receivers, transceivers, cryptographic circuits, cellular phones, digital cameras, positioning and / or navigation systems, medical devices, head-up displays, wireless devices, audio recording, audio storage and / or audio playback devices, video recording, video storage and / or video playback devices, gaming platforms, peripheral devices, and / or multi-chip modules. Those skilled in the art will understand that the elements of this invention can be implemented in other types of devices to meet specific application standards.

[0242] When used herein in conjunction with "is" and verbs, the terms "may" and "usually" are intended to convey the intention that the description is exemplary and is considered broad enough to encompass both the specific examples presented in this disclosure and the alternative examples that may be derived based on this disclosure. As used herein, the terms "may" and "usually" should not be construed as necessarily implying the desirability or possibility of omitting the corresponding element.

[0243] When the notations “a”-“n” are used herein for various components, modules, and / or circuits, a single component, module, and / or circuit or multiple such components, modules, and / or circuits are disclosed, wherein the notation “n” is applied to represent any particular integer. Each different component, module, and / or circuit having an instance (or event) notated as “a”-“n” may indicate that the different component, module, and / or circuit may have a matching number of instances or a different number of instances. An instance designated as “a” may represent the first of multiple instances, and an instance “n” may refer to the last of multiple instances without implying a specific number of instances.

[0244] Although the invention has been specifically shown and described with reference to embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention.

Claims

1. An apparatus comprising: An interface configured to receive: (i) pixel data of the environment; and (ii) information about the movement of the device; as well as The processor is configured to: (i) process pixel data arranged as video frames; (ii) Measuring motion information; (iii) Generating a playback recording in response to (a) the video frame of the calibration target captured during a vibration mode applied to the device and (b) the motion information of the vibration mode; (iv) Implementing image stabilization compensation in response to (a) the lens projection function and (b) the motion information; (v) Perform additional compensation in response to the calibration value; and (vi) performing a calibration operation to determine the calibration value, wherein (i) the playback recording includes the video frames generated at multiple predefined zoom levels of the lens, and (ii) the calibration operation includes: (I) Determine the coordinates of the calibration target for each video frame in the playback record. (II) In response to a comparison of the coordinates of the video frame determined using the image stabilization compensation with the coordinates of the calibration target, a pixel difference matrix is ​​determined, and (III) The calibration value is generated in response to curve fitting performed on the pixel difference matrix.

2. The apparatus according to claim 1, wherein, (i) the calibration operation is further configured to implement an evaluation tool, and (ii) the evaluation tool is configured to test the stabilized video frames generated by the processor in response to the application of the image stabilization compensation and the additional compensation using the calibration value.

3. The apparatus according to claim 2, wherein, The evaluation tool is configured to ensure that the accuracy of the stabilized video frames is within 1 / 16 of the pixel.

4. The apparatus according to claim 1, wherein, The calibration target comprises a set of nine chessboard patterns arranged in a predefined pattern.

5. The apparatus according to claim 4, wherein, The predefined pattern is (i) at a distance of at least 1.5 m from the lens, and (ii) within the field of view of the lens during the vibration mode.

6. The apparatus according to claim 4, wherein, (i) the predefined pattern includes each of the nine chessboard patterns having planes arranged at different angles, and (ii) the different angles are at least 10 degrees apart between each of the planes.

7. The apparatus according to claim 4, wherein, (i) The pixel difference matrix includes data points corresponding to each grid coordinate for each calibration target in the calibration targets, and (ii) Each grid coordinate is not less than 50×50 pixel units.

8. The apparatus according to claim 1, wherein, (i) While generating the video frame of the calibration target, the vibration pattern is applied to the device, and (ii) the inertial measurement unit of the device measures the movement information of the vibration pattern.

9. The apparatus according to claim 8, wherein, The vibration pattern is generated by a vibration device used on the device.

10. The apparatus according to claim 1, wherein, (i) the processor is configured to implement a simulation framework, and (ii) the simulation framework is configured to synchronize the video frames with the motion information used for the playback recording.

11. The apparatus according to claim 10, wherein, (i) the playback record includes metadata, and (ii) the metadata includes the motion information, image sensor exposure shutter timing, and exposure start timing.

12. The apparatus according to claim 10, wherein, The video frames recorded during playback include at least 600 YUV sequences of the pixel data.

13. The apparatus according to claim 10, wherein, The video frames recorded during playback include high-quality compressed frames of the pixel data.

14. The apparatus according to claim 10, wherein, The simulation framework is configured to automatically generate the playback record for each of the plurality of predefined zoom levels in response to a trigger input.

15. The apparatus according to claim 1, wherein, (i) the processor is configured to implement a calibration phase for each of the plurality of predefined zoom levels, and (ii) the calibration phase includes: (a) determining the coordinates of the calibration target for each video frame in the playback record corresponding to one of the plurality of predefined zoom levels, (b) determining the pixel difference matrix for each of the predefined zoom levels, and (c) generating a set of calibration values ​​for each of the predefined zoom levels.

16. The apparatus according to claim 1, wherein, (i) the vibration mode includes a combination of vibrations in the pitch direction, vibrations in the roll direction, and vibrations in the yaw direction, and (ii) each of the vibrations in the pitch direction, the vibrations in the roll direction, and the vibrations in the yaw direction includes a corresponding vibration frequency and vibration amplitude.

17. The apparatus according to claim 1, wherein, The movement information includes one or more of the following: maximum amplitude, actual amplitude, actual angle value, vibration frequency, and vibration duration.

18. The apparatus according to claim 1, wherein, The total compensation generated by the processor for the stabilized video frame is a combination of the image stabilization compensation and the additional compensation.

19. The apparatus according to claim 18, wherein, (i) The total compensation is determined according to the following equation: Final_comp = R[k1*EFL*r,k2*h,k3*r^2], and (ii) The calibration value includes k1, k2 and k3.

20. The apparatus according to claim 1, wherein, The curve fitting performed on the pixel difference matrix includes polynomial curve fitting using Taylor's theorem.

Citation Information

Patent Citations

  • Memory hierarchy to transfer vector data for operators of a directed acyclic graph

    US10437600B1

  • Using camera data to manage a vehicle parked outside in cold climates

    US11001231B1

  • Generating training data for speed bump detection

    US11586843B1

  • Generating detection parameters for a rental property monitoring solution using computer vision and audio analytics from a rental agreement

    US11645706B1

  • Object-aware temperature anomalies monitoring and early warning by combining visual and thermal sensing sensing

    US20220044023A1