Fixed pattern calibration for multi-view stitching
Through the fixed pattern calibration method of the interface and processor device, the inherent parameters and distortion parameters of the camera are utilized to solve the problem of multiple calibrations in multi-view stitching, and achieve efficient seamless stitching at different distances.
Patent Information
- Application Number
- CN202110040734.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-13
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-03-06
AI Technical Summary
The existing technology requires multiple calibrations for camera systems at different distances in multi-view stitching, which makes the process cumbersome and inefficient and cannot achieve fixed pattern calibration.
The device adopts an interface and a processor, and realizes fixed pattern calibration by utilizing the intrinsic parameters and distortion parameters of the camera in combination with a calibration plate through the geometric calibration and posture calibration process. The device is suitable for stitching calibration of short, medium and long distances.
It simplifies the multi-view stitching process, achieves seamless stitching at any distance, reduces parallax problems, and improves stitching efficiency and accuracy.
Smart Images

Figure CN114765667B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to multi-view stitching, and more particularly to methods and / or apparatus for implementing fixed pattern calibration for multi-view stitching. Background Art
[0002] Due to parallax, it is not possible to perfectly stitch images for all distances. If two images are well aligned, you can use similar triangles to get the relationship between the two images based on parallax:
[0003] Parallax = focal length * baseline / distance Equation 1 From the above equation, we can see that parallax is related to the camera system. Parallax is more obvious at short distances than at long distances.
[0004] In conventional calibration methods, multi-view stitching calibration is based on a calibration target at a specific distance from the camera lens. After the calibration process, the images from the camera will be stitched at the target's location. At the target's location, the disparity will be zero, but at other locations (e.g., close to or far from the target's distance), the disparity should be the actual disparity minus the disparity at the target's location.
[0005] Different camera systems (e.g., different lenses / sensors / baselines) have different stitching distance requirements. If a customer wishes to use the same camera system to achieve stitching results at different distances, they will need to run the calibration process multiple times, placing the calibration plate at the target distance for each calibration run. Therefore, the conventional process requires multiple calibrations for each stitching position.
[0006] It is desirable to implement fixed pattern calibration for multi-view stitching. Summary of the Invention
[0007] The present invention covers aspects relating to an apparatus comprising an interface and a processor. The interface can be configured to receive video signals from two or more cameras arranged to obtain a predetermined field of view, wherein the respective fields of view of each pair of the two or more cameras overlap. The processor can be configured to perform a fixed pattern calibration for facilitating multi-view stitching. The fixed pattern calibration generally comprises: (a) performing a geometric calibration process to obtain intrinsic parameters and distortion parameters of each lens of the two or more cameras; and (b) applying a pose calibration process to the video signals using (i) the intrinsic parameters and distortion parameters of each lens of the two or more cameras and (ii) a calibration plate to obtain configuration parameters for the respective fields of view of the two or more cameras.
[0008] In some embodiments of the apparatus aspects described above, the pose calibration process includes detecting circle centers or corners, respectively, on a calibration board using at least one of a circle detector and a checkerboard detector; and determining whether detected points in one view match in another view.
[0009] In some embodiments of the apparatus aspects described above, a pose calibration process includes performing extrinsic calibration for each lens of the two or more cameras using intrinsic parameters and distortion parameters of each lens of the two or more cameras, wherein the extrinsic parameters of each lens of the two or more cameras include a corresponding rotation matrix and a corresponding translation vector. In some embodiments, the pose calibration process further includes changing a z value of the corresponding translation vector of each lens of the two or more cameras to a medium-range value or a long-range value while maintaining the corresponding rotation matrix of each lens of the two or more cameras unchanged.
[0010] In some embodiments of the apparatus aspects described above, the pose calibration process includes projecting keypoints from world coordinates to image coordinates.
[0011] In some embodiments of the apparatus aspects described above, the pose calibration process includes computing a corresponding homography matrix between each adjacent pair of two or more cameras. In some embodiments, the corresponding homography matrix between each adjacent pair of two or more cameras is based on a central view.
[0012] In some embodiments of the apparatus aspects described above, the pose calibration process includes: applying intrinsic parameters and distortion parameters of each lens of the two or more cameras to corresponding images; and warping the corresponding images using corresponding homography matrices.
[0013] In some embodiments of the apparatus aspects described above, the pose calibration process includes applying a projection model to the views of the two or more cameras. In some embodiments of applying the projection model, for horizontal stitching, the projection model includes at least one of a perspective model, a cylindrical model, and an equirectangular model; and for vertical stitching, the projection model includes at least one of a transverse cylindrical model and a Mercator model.
[0014] The present invention also covers aspects relating to a method for fixed pattern calibration for multi-view stitching using multiple cameras, comprising: (a) arranging two or more cameras to obtain a predetermined field of view, wherein the corresponding fields of view of each adjacent pair of two or more cameras overlap; (b) performing a geometric calibration process to obtain intrinsic parameters and distortion parameters of each lens of the two or more cameras; and (c) applying a pose calibration process to video signals from the two or more cameras using (i) the intrinsic parameters and the distortion parameters of each lens of the two or more cameras and (ii) a calibration plate to obtain configuration parameters for the corresponding fields of view of the two or more cameras.
[0015] In some embodiments of the method aspects described above, the pose calibration process includes: using at least one of a circle detector and a checkerboard detector to detect circle centers or corners on a calibration board, respectively; and determining whether the detected points in one view match in another view.
[0016] In some embodiments of the method aspects described above, a pose calibration process includes performing extrinsic calibration for each lens of the two or more cameras using intrinsic parameters and distortion parameters of each lens of the two or more cameras, wherein the extrinsic parameters of each lens of the two or more cameras include a corresponding rotation matrix and a corresponding translation vector. In some embodiments, the pose calibration process further includes changing a z value of the corresponding translation vector of each lens of the two or more cameras to a medium-range value or a long-range value while maintaining the corresponding rotation matrix of each lens of the two or more cameras unchanged.
[0017] In some embodiments of the method aspects described above, the pose calibration process includes projecting keypoints from world coordinates to image coordinates.
[0018] In some embodiments of the method aspects described above, the pose calibration process includes computing a corresponding homography matrix between each adjacent pair of two or more cameras. In some embodiments, the corresponding homography matrix between each adjacent pair of two or more cameras is based on a central view.
[0019] In some embodiments of the method aspects described above, the pose calibration process includes: (a) applying intrinsic parameters and distortion parameters of each lens of two or more cameras to corresponding images; and warping the corresponding images using corresponding homography matrices.
[0020] In some embodiments of the method aspects described above, the pose calibration process includes applying a projection model to the views of the two or more cameras. In some embodiments of applying the projection model, for horizontal stitching, the projection model includes at least one of a perspective model, a cylindrical model, and an equirectangular model; and for vertical stitching, the projection model includes at least one of a transverse cylindrical model and a Mercator model. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Embodiments of the invention will become apparent from the following detailed description, appended claims and accompanying drawings.
[0022] Figure 1 is a diagram illustrating an example context for implementing an example embodiment of the present invention.
[0023] Figure 2 is a diagram illustrating a calibration layout according to an example embodiment of the present invention.
[0024] Figure 3 is a diagram illustrating an object imaged by two adjacent cameras according to an embodiment of the present invention.
[0025] Figure 4 is a diagram illustrating frames from two adjacent cameras imaging a common object, according to an embodiment of the present invention.
[0026] Figure 5 is a diagram illustrating a two-dimensional (2D) array of cameras according to an embodiment of the present invention.
[0027] Figure 6 is a diagram illustrating a camera system in which example embodiments of the present invention may be implemented.
[0028] Figure 7 is a diagram illustrating a calibration process according to an example embodiment of the present invention.
[0029] Figure 8 is a diagram illustrating a posture calibration process according to an example embodiment of the present invention.
[0030] Figure 9 is a diagram of a camera system illustrating an example implementation of a computer vision system in which calibration techniques according to example embodiments of the present invention may be implemented.
[0031] Figure 10 is a diagram illustrating a context in which calibration techniques according to example embodiments of the present invention may be implemented.
[0032] Figure 11 It shows Figure 10 Figure 2 shows a general implementation of a hardware engine. DETAILED DESCRIPTION
[0033] Embodiments of the present invention include providing a fixed pattern calibration for multi-view stitching that can (i) perform stitching calibration using a target (e.g., a calibration plate, etc.) at short distances; (ii) calculate parameters for stitching at medium and long distances; (iii) utilize intrinsic and distortion information of the lens and extrinsic calibration information for the target to obtain good stitching calibration information for any position; (iv) be implemented as a pipelined process; (v) apply different projection models for horizontal stitching and vertical stitching; (vi) be applied to a two-dimensional (2D) camera array and / or (vii) be implemented as one or more integrated circuits.
[0034] In various embodiments, a new stitching calibration method is provided that uses a short-distance target (e.g., a calibration plate, etc.) for stitching calibration and calculates parameters for medium-distance and long-distance stitching, especially for long-distance stitching. In various embodiments, the new stitching calibration method generally has a first step including lens calibration and a second step including posture calibration. In the first step, a geometric calibration is generally performed for each lens to obtain intrinsic parameters and distortion parameters of each lens. The lens calibration step is generally important for the distance measurement accuracy in the second posture calibration step. In the second step, calculations are generally performed using the intrinsic parameters and distortion parameters of each lens and non-intrinsic calibration information for the target to obtain stitching parameter information for any position.
[0035] refer to Figure 1 , a diagram illustrating an example context of the present invention is shown. A house 50 and a vehicle 52 are shown. A camera system 100a-100n is shown. Each of the cameras 100a-100n can be configured to operate independently of each other. Each of the cameras 100a-100n can be configured to generate a video signal that transmits an image for a corresponding field of view (FOV). In one example, multiple cameras 100a-100n can be configured to have overlapping corresponding fields of view. The corresponding video frames from all cameras 100a-100n or a subset thereof can be stitched together to provide a common field of view that is larger than any one of the corresponding fields of view. In an example, the corresponding video frames can be stored locally (e.g., on a microSD card, to a local network attached storage (NAS) device, etc.).
[0036] Each of cameras 100a-100n can be configured to detect different or the same events / objects that can be considered interesting. For example, camera system 100b can capture an area near the entrance of house 50. For the entrance of house 50, the object / event of interest can be detecting a person. Camera system 100b can be configured to analyze video frames to detect a person, and can slow down when a person is detected (e.g., select a video frame to be encoded at a higher frame rate). In another example, camera system 100d can capture an area near vehicle 52. For vehicle 52, the object / event of interest can be detecting other vehicles and pedestrians when on the road, and detecting suspicious activity or people when parked at house 50. Camera system 100d can be configured to analyze video frames to detect vehicles (or road signs) and people, and can slow down when a vehicle or person is detected.
[0037] In general, as used herein, a "frame" generally refers to a single captured image of one or more objects in the respective fields of view of a camera. Motion video generally includes a temporal sequence of frames captured at a video rate of typically ten to thirty frames per second. However, other frame rates may be implemented to meet the design criteria of a particular implementation. In general, as used herein, the term "parallax" refers to the apparent change in the position of an object due to a change in the viewing position that provides a new line of sight. Parallax generally occurs because two cameras have different centers of projection (COPs). When images from cameras with different COPs are combined without modification, parallax generally results between the objects of the combined image, and the objects appear to have "ghost images" or double images in the composite image. The term "common field of view" is generally used herein to describe an image or portion of an image of a scene viewed (or captured) by more than one camera.
[0038] In various embodiments, multiple video cameras can be mounted on a rigid substrate so that the respective field of view of each camera overlaps the respective field of view of each adjacent (neighboring) camera. Multiple video cameras can be calibrated using fixed pattern calibration techniques according to example embodiments of the present invention. The resulting images can be aligned using digital warping, objects can be matched using parallax reduction techniques, and the images can be stitched together to form a large composite image. The result is a seamless, high-resolution video image that spans the common field of view imaged by the multiple video cameras.
[0039] refer to Figure 2, a diagram illustrating a calibration layout according to an example embodiment of the present invention is shown. In various embodiments, the camera (or imaging) system 100 may include a plurality of video cameras arranged in a spaced-apart radially oriented array to collectively capture a panoramic field of view. In an example, the camera system 100 may include a plurality of lenses 60a-60n. Each of the plurality of lenses 60a-60n may be associated with (attached to) a respective image capture device. Each of the plurality of lenses 60a-60n may have a respective field of view centered about the optical axis of each lens (e.g., shown by a dashed center line). The lenses 60a-60n are typically arranged so that the respective fields of view of each adjacent pair of lenses 60a-60n overlap.
[0040] Each of the plurality of lenses 60a-60n is associated with a corresponding image sensor configured to capture images of the corresponding field of view through the lens 60a-60n. Each image sensor associated with the plurality of lenses 60a-60n is typically configured to simultaneously transmit a stream of digital or analog output to a processor 110. The processor 110 processes the multiple signals to seamlessly stitch the multiple images from the adjacent cameras into a single stream of digital or analog video output of the wide-angle scene. When generating the single stream of digital or analog video output of the wide-angle scene, the processor 110 typically removes any distortion created by the image capture process and removes redundant pixels recorded in overlapping fields of view.
[0041] Because it is impractical to build a camera array with each camera having a common center of projection, the camera array can have a small but noticeable baseline separation. Baseline separation is a problem when combining images of objects at different distances from the baseline, because a single warping function will only work perfectly for one specific distance. Images of objects not at that distance may be warped to different positions and may appear to have double images ("ghost images") or be truncated when the images are combined.
[0042] In various embodiments, the camera system 100 can be calibrated using a fixed pattern calibration scheme according to an embodiment of the present invention so that objects at any distance can be combined without visible parallax. The camera calibration scheme according to an embodiment of the present invention can greatly simplify the stereo matching problem. In an example, the calibration plate 102a can be placed at a position that appears in the captured images of the left camera and the center camera. Similarly, the calibration plate 102b can be placed at a position that appears in the captured images of the center camera and the right camera. In various embodiments, it is only necessary to place the calibration plates 102a and 102b at a short distance (e.g., 3 meters (m)) from the lenses 60a-60n. The calibration parameters for farther positions (e.g., 10m, 50m, etc.) can be calculated using a multi-view stitching calibration technique according to an embodiment of the present invention.
[0043] Typically, calibration of all camera views can be based on a center view (or baseline). In various embodiments, the minimum number of cameras for multi-view stitching is two. In examples where the total number of cameras is two, any camera can be selected as the center view (or baseline). In examples where the total number of cameras is an even number and greater than two, the camera closest to the center can be selected as the baseline. In examples where the total number of cameras is an odd number, the center camera is typically selected as the baseline. In general, calibration techniques according to embodiments of the present invention can also be applied to two-dimensional (2D) camera arrays.
[0044] refer to Figure 3 , shows a diagram illustrating an object imaged by two adjacent cameras according to an embodiment of the present invention. In various embodiments, adjacent frames may be stitched together using a method or combination of methods for combining separate images into a panoramic image or a combined image. In an example, a spatial transformation (warping) of a quadrilateral region may be used that merges at least two images into one larger image without losing the commonality of the multiple images. First, a plurality of image registration points may be determined. For example, a fixed point may be imaged at a known position in each sub-image, either manually or automatically. In either case, the calibration process involves pointing the array of cameras at a known structured scene and finding corresponding points. For example, Figure 3 The view of two cameras trained on a scene including a rectangular target 102 is shown. In an example, the rectangular target 102 can be implemented as a checkerboard calibration plate or a circular calibration plate. The rectangular target 102 is usually placed in the viewing Figure 1 and visual Figure 2 In the overlapping area between Figure 1 Heshi Figure 2 Each corresponds to an approximately corresponding field of view captured from one of the two cameras. The corners of the rectangular target 102 may constitute image registration points (e.g., points common to the views of each of the two or more cameras of the camera array). In embodiments where the rectangular target 102 comprises a checkerboard calibration plate or a circular calibration plate, a checkerboard detector or a circle detector may be used to obtain the image registration points.
[0045] refer to Figure 4 , showing the captured rectangular target 102 corresponding to Figure 3 Vision Figure 1 Heshi Figure 2Figure 1 is a diagram of an actual frame of two cameras. The frames of the two cameras are shown as adjacent frames - frame 1 and frame 2. The target can be captured as a quadrilateral region 104 with corners A, B, C and D in frame 1 and a quadrilateral region 106 with corners A', B', C' and D' in frame 2. Because the camera angles are slightly different, the quadrilateral regions 104 and 106 are inconsistent in angular construction and are generally captured at different positions relative to other objects in the image. In the example, the adjacent frames can be matched by warping each of the quadrilateral regions 104 and 106 into a common coordinate system. Note that the sides of the quadrilateral region 104 are shown as straight, but in reality may be affected by some barrel or pincushion distortion, which can also be approximately corrected via the warping operation. In the example, a radial (rather than a piecewise linear) transformation can be used to correct for barrel / pincushion distortion. The piecewise linear transformation can fix the approximation of the curve.
[0046] In another example, only one of the images may be warped to match the coordinate system of the other image. For example, the warping of quadrilateral region 104 may be performed via a perspective transformation. Thus, quadrilateral 104 in frame 1 may be transformed into quadrilateral 106 in the coordinate system of frame 2.
[0047] Because it's often impractical to build a camera array with every camera having a common COP, camera arrays typically have a small but noticeable baseline separation. This baseline separation becomes a problem when combining images of objects at different distances from the baseline, as a single warping function will only work perfectly for one specific distance. Images of objects not at that distance may be warped to different positions and may appear double ("ghost") or truncated when the images are combined.
[0048] In various embodiments, the camera array can be calibrated so that images of an object or smooth background at a specific distance can be combined without visible parallax. The minimum parallax can be found by determining how much to shift one image to match the other. Because the images are warped into corresponding squares, all that needs to be done is find a specific shift that matches the corresponding square.
[0049] An imaging system according to an embodiment of the present invention typically includes a plurality of video cameras arranged in a spaced-apart radially oriented array to collectively capture a panoramic or spherical panorama (panospheric) field of view. The imaging system may further include a processor circuit configured to simultaneously receive each stream of digital or analog output from the plurality of cameras. The processor circuit may be configured to process the collection of signals to remove any distortion created by the image capture process, seamlessly merge multiple images from adjacent cameras, by removing redundant pixels recorded in overlapping fields of view, so as to generate a single stream of digital or analog video output of a wide-angle scene. The processor may present the wide-angle scene to display it on a display device such as a monitor, a virtual reality helmet, or a projection display for viewing the image.
[0050] refer to Figure 5 , other configurations of the camera lenses 60a-60n are also possible. In an example, according to another example embodiment of the present invention, the camera lenses 60a-60n can be mounted in a planar array 108. In an example, the camera lenses 60a-60n can be aligned in various directions. In an example, the camera lenses 60a-60n can be aligned in the same direction (with overlapping views as described above). However, other more diverse alignments can be achieved (e.g., generally still having overlapping and adjacent areas) to meet the design criteria of a particular implementation. For example, the camera array can be implemented to have an arrangement including, but not limited to, radial, linear, and / or planar. In some embodiments, the position and angle of each of the cameras in the camera array can be fixed relative to another camera. Therefore, the entire camera array can be moved without any recalibration (re-registration) or recalculation of the matrix equations used to transform the individual images into a single panoramic scene.
[0051] refer to Figure 6, a diagram illustrating a camera system 100 in which an example embodiment of the present invention may be implemented is shown. In an example, the camera system 100 may be implemented as a camera system-on-a-chip connected to multiple lens and sensor assemblies. In an example embodiment, the camera system 100 may include lenses 60a-60n, motion sensors 70a-70n, capture devices 80a-80n, a processor / SoC 110, a block (or circuit) 112, a block (or circuit) 114, and / or a block (or circuit) 116. Circuit 112 may be implemented as a memory. Block 114 may be a communication module. Block 116 may be implemented as a battery. In some embodiments, the camera system 100 may include lenses 60a-60n, motion sensors 70a-70n, capture devices 80a-80n, a processor / SoC 110, a memory 112, a communication module 114, and a battery 116. In another example, camera system 100 may include lenses 60a-60n, motion sensors 70a-70n, and capture devices 80a-80n, and processor / SoC 110, memory 112, communication module 114, and battery 116 may be components of a single device. The implementation of camera system 100 may vary depending on the design criteria of a particular implementation.
[0052] Lenses 60a-60n are shown attached to respective capture devices 80a-80n. In an example, capture devices 80a-80n are shown as including blocks (or circuits) 82a-82n, blocks (or circuits) 84a-84n, and blocks (or circuits) 86a-86n, respectively. Circuits 82a-82n may be sensors (e.g., image sensors). Circuits 84a-84n may be processors and / or logic. Circuits 86a-86n may be memory circuits (e.g., frame buffers).
[0053] The capture devices 80a-80n can be configured to capture video image data (e.g., light collected and focused by the lenses 60a-60n). The capture devices 80a-80n can capture data received through the lenses 60a-60n to generate a video bitstream (e.g., a sequence of video frames). The lenses 60a-60n can be oriented, tilted, panned, zoomed, and / or rotated to capture the environment surrounding the camera system 100 (e.g., to capture data from a corresponding field of view).
[0054] The capture devices 80a-80n can convert the received light into a digital data stream. In some embodiments, the capture devices 80a-80n can perform analog-to-digital conversion. For example, the capture devices 80a-80n can perform photoelectric conversion on the light received by the lenses 60a-60n. The image sensors 80-80n can convert the digital data stream into a video data stream (or bitstream), a video file, and / or a plurality of video frames. In an example, each of the capture devices 80a-80n can present the video data as a digital video signal (e.g., signals VIDEO_A-VIDEO_N). The digital video signal can include video frames (e.g., sequential digital images and / or audio).
[0055] The video data captured by the capture devices 80a-80n can be represented as signals / bitstreams / data VIDEO_A-VIDEO_N (e.g., digital video signals). The capture devices 80a-80n can present the signals VIDEO_A-VIDEO_N to the processor / SoC 110. The signals VIDEO_A-VIDEO_N can represent video frames / video data. The signals VIDEO_A-VIDEO_N can be video streams captured by the capture devices 80a-80n.
[0056] The image sensors 82a-82n can receive light from the corresponding lenses 60a-60n and convert the light into digital data (e.g., a bit stream). For example, the image sensors 82a-82n can perform photoelectric conversion on the light from the lenses 60a-60n. In some embodiments, the image sensors 82a-82n can have additional margin that is not used as part of the image output. In some embodiments, the image sensors 82a-82n may not have additional margin. In some embodiments, some of the image sensors 82a-82n may have additional margin, and some of the image sensors 82a-82n may not have additional margin. In some embodiments, the image sensors 82a-82n can be configured to generate a monochrome (B / W) video signal. In some embodiments, the image sensors 82a-82n can be configured to generate a color (e.g., RGB, YUV, RGB-1R, YCbCr, etc.) video signal. In some embodiments, the image sensors 82a-82n can be configured to generate a video signal in response to visible light and / or infrared (IR) light.
[0057] The processor / logic 84a-84n can convert the bitstream into human-viewable content (e.g., video data, e.g., video frames, that can be understood by an average person regardless of image quality). For example, the processor 84a-84n can receive pure (e.g., raw) data from the camera sensor 82a-82n and generate (e.g., encode) video data (e.g., a bitstream) based on the raw data. The capture device 80a-80n can have a memory 86a-86n to store the raw data and / or the processed bitstream. For example, the capture device 80a-80n can implement a frame memory and / or buffer 86a-86n to store one or more of the video frames (e.g., video signals) (e.g., provide temporary storage and / or caching therefor). In some embodiments, the processor / logic 84a-84n can perform analysis and / or correction on the video frames stored in the memory / buffer 86a-86n of the capture device 80a-80n.
[0058] Motion sensors 70a-70n can be configured to detect motion (e.g., in a corresponding field of view corresponding to the viewing angle of lenses 60a-60n). Detection of motion can be used as a threshold for activating capture devices 80a-80n. Motion sensors 70a-70n can be implemented as internal components of camera system 100 and / or components external to camera system 100. In an example, sensors 70a-70n can be implemented as passive infrared (PIR) sensors. In another example, sensors 70a-70n can be implemented as smart motion sensors. In an example, a smart motion sensor can include a low-resolution image sensor configured to detect motion and / or a person. Motion sensors 70a-70n can each generate a corresponding signal (e.g., SENS_A-SENS_N) in response to detecting motion in one of the corresponding areas (e.g., FOV). Signals SENS_A-SENS_N can be presented to processor / SoC 110. In an example, motion sensor 70a may generate signal SENS_A (assert signal SENS_A) when motion is detected in the corresponding FOV of lens 60a; and motion sensor 70a may generate signal SENS_N (assert signal SENS_N) when motion is detected in the corresponding FOV of lens 60n.
[0059] The processor / SoC 110 may be configured to execute computer-readable code and / or process information. The processor / SoC 110 may be configured to receive input and / or present output to the memory 112. The processor / SoC 110 may be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of the processor / SoC 110 may vary according to the design criteria of a particular implementation. The processor / SoC 110 may be configured for low-power (e.g., battery) operation.
[0060] Processor / SoC 110 may receive signals VIDEO_A-VIDEO_N and signals SENS_A-SENS_N. Processor / SoC 110 may generate signal META based on signals VIDEO_A-VIDEO_N, signals SENS_A-SENS_N, and / or other inputs. In some embodiments, signal META may be generated based on analysis of signals VIDEO_A-VIDEO_N and / or objects detected in signals VIDEO_A-VIDEO_N. In various embodiments, processor / SoC 110 may be configured to perform one or more of feature extraction, object detection, object tracking, and object identification. For example, processor / SoC 110 may determine motion information by analyzing a frame from signals VIDEO_A-VIDEO_N and comparing the frame to a previous frame. This comparison may be used to perform digital motion estimation.
[0061] In various embodiments, the processor / SoC 110 may perform a video stitching operation. The video stitching operation may be configured to facilitate seamless tracking as an object moves through the respective fields of view associated with the capture devices 80a-80n. The processor / SoC 110 may generate a plurality of signals VIDOUT_A-VIDOUT_N and STITCHED_VIDEO. The signals VIDOUT_A-VIDOUT_N may be portions (components) of a multi-sensor video signal. In some embodiments, the processor / SoC 110 may be configured to generate a single video output signal (e.g., STITCHED_VIDEO). Video output signal(s) (e.g., STITCHED_VIDEO or VIDOUT_A-VIDOUT_N) may be generated that include video data from one or more of the signals VIDEO_A-VIDEO_N. The video output signal(s) (e.g., STITCHED_VIDEO or VIDOUT_A-VIDOUT_N) may be presented to the memory 112 and / or the communication module 114.
[0062] The memory 112 can store data. The memory 112 can be implemented as a cache memory, flash memory, a memory card, a DRAM memory, etc. The type and / or size of the memory 112 can vary according to the design standards of a particular implementation. The data stored in the memory 112 can correspond to video files, motion information (e.g., readings from sensors 70a-70n, video stitching parameters, image stabilization parameters, user input, etc.) and / or metadata information.
[0063] The lenses 60a-60n (e.g., camera lenses) can be oriented to provide a view of the environment surrounding the camera 100. The lenses 60a-60n can be designed to capture environmental data (e.g., light). The lenses 60a-60n can be wide-angle lenses and / or fisheye lenses (e.g., lenses capable of capturing a wide field of view). The lenses 60a-60n can be configured to capture and / or focus light for the capture devices 80a-80n. Typically, image sensors 82a-82n are located behind the lenses 60a-60n. Based on the light captured from the lenses 60a-60n, the capture devices 80a-80n can generate a bitstream and / or video data.
[0064] The communication module 114 can be configured to implement one or more communication protocols. For example, the communication module 114 can be configured to implement Wi-Fi, Bluetooth, Ethernet, etc. In embodiments where the camera 100 is implemented as a wireless camera, the protocol implemented by the communication module 114 can be a wireless communication protocol. The type of communication protocol implemented by the communication module 114 can vary depending on the design criteria of a particular implementation.
[0065] The communication module 114 can be configured to generate a broadcast signal as an output from the camera 100. The broadcast signal can transmit the video data VIDOUT to an external device. For example, the broadcast signal can be sent to a cloud storage service (e.g., a storage service that can be expanded on demand). In some embodiments, the communication module 114 may not transmit data until the processor / SoC 110 has performed video analysis to determine that an object is in the field of view of the camera 100.
[0066] In some embodiments, the communication module 114 may be configured to generate a manual control signal. The manual control signal may be generated in response to a signal from a user received by the communication module 114. The manual control signal may be configured to activate the processor / SoC 110. The processor / SoC 110 may be activated in response to the manual control signal regardless of the power state of the camera 100.
[0067] The camera system 100 may include a battery 116 configured to provide power to the various components of the camera 100. A multi-step method for activating and / or deactivating the capture devices 80a-80n and / or any other power consumption features of the camera system 100 based on the output of the motion sensors 70a-70n may be implemented to reduce the power consumption of the camera 100 and extend the operating life of the battery 116. The motion sensors 70a-70n may have a low power drain on the battery 116 (e.g., less than 10W). In an example, the motion sensors 70a-70n may be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 110. The video analytics performed by the processor / SoC 110 may have a significant power drain on the battery 116 (e.g., greater than the motion sensors 70a-70n). In an example, the processor / SoC 110 may be in a low-power state (or powered down) until the motion sensors 70a-70b detect certain motion.
[0068] The camera system 100 can be configured to operate using various power states. For example, in a powered-down state (e.g., a sleep state, a low-power state), the motion sensors 70a-70n and the processor / SoC 110 can be powered on, while other components of the camera 100 (e.g., image capture devices 80a-80n, memory 112, communication module 114, etc.) can be powered off. In another example, the camera 100 can operate in an intermediate state. In the intermediate state, one of the image capture devices 80a-80n can be powered on, and the memory 112 and / or communication module 114 can be powered off. In yet another example, the camera system 100 can operate in a powered-on (or high-power) state. In the powered-on state, the motion sensors 70a-70n, the processor / SoC 110, the capture devices 80a-80n, the memory 112, and / or the communication module 114 can be powered on. In the powered-down state, the camera system 100 can consume some power (e.g., a relatively small and / or minimal amount of power) from the battery 116. In the powered-on state, camera system 100 may consume more power from battery 116. The number of power states and / or the number of components of camera system 100 that are powered on when camera system 100 operates in each of the power states may vary according to the design criteria of a particular implementation.
[0069] refer to Figure 7, a diagram illustrating a calibration process 200 according to an example embodiment of the present invention is shown. In various embodiments, process 200 generally implements new calibration techniques. The new calibration techniques can automatically generate stitching calibration distance parameters, and the calibration process may be more efficient. As technology advances, stitching calibration can be constructed to utilize intrinsic and distortion information of the lens, as well as extrinsic calibration information for multiple targets, to obtain good stitching calibration information in any position.
[0070] In an example embodiment, process (or method) 200 may include step (or state) 202, step (or state) 204, and step (or state) 206. Step 202 generally implements the first stage of the multi-view stitching process pipeline. Step 204 generally implements the second stage of the multi-view stitching process pipeline. Step 206 generally presents system configuration parameters for multi-view stitching calculated in the multi-view stitching process pipeline.
[0071] Step 202 typically performs a lens calibration process. In an example, the lens calibration process may include geometric calibration. Geometric calibration should be performed for each lens in the camera system 100. Geometric calibration is typically performed for each channel (e.g., channel 0 to channel N) to obtain intrinsic parameters and distortion parameters for each of the lenses in the system 100. The lens calibration step is important for the accuracy of distance measurements in the second step 204 of the multi-view stitching process pipeline. Generally, it is better to cover the entire field of view (FOV) using multiple images from the calibration target pattern.
[0072] In one example, lens calibration can be performed using techniques found in Zhengyou Zhang, "A Flexible New Technique for Camera Calibration" (IEEE Trans. Pattern Analysis and Machine Intelligence, December 2000, Vol. 22: pp. 1330-1334) and Zhengyou Zhang, "Flexible Camera Calibration By Viewing a Plane From Unknown Orientations" (Computer Vision, 1999, The Proceedings of the Seventh IEEE International Conference, September 1999, IEEE Press), both of which are incorporated herein by reference. Generally, because each lens may have distortion, distortion parameters (e.g., k1, k2, k3, p1, p2, etc.) can be defined for each lens. Intrinsic parameters (e.g., fx, fy, cx, cy, etc.) can also be defined. In one example, the distortion parameters and intrinsic parameters can be fitted by taking many snapshots of a calibration pattern (e.g., a checkerboard, a circular plate, etc.). Examples of calibration patterns can be found in “A Flexible New Technique for Camera Calibration” by Zhengyou Zhang (IEEE Trans. Pattern Analysis and Machine Intelligence, December 2000, Vol. 22: pp. 1330-1334) and “Flexible Camera Calibration By Viewing a Plane From Unknown Orientations” by Zhengyou Zhang (Computer Vision, 1999, The Proceedings of the Seventh IEEE International Conference, September 1999, IEEE Press), which are incorporated herein by reference.
[0073] In step 204, the process 200 typically performs pose calibration using inputs from channels 0 to N and the intrinsic and distortion parameters obtained for each of the lenses in step 202. In step 204, the process 200 typically performs stitching calibration using a calibration plate (or target) at a short distance (e.g., 3 meters), and then calculates system configuration parameters for multi-view stitching for medium-range and long-range stitching.
[0074] In step 206, process 200 generally presents system configuration parameters for multi-view stitching calculated in the multi-view stitching process pipeline. The system configuration parameters generally facilitate multi-view stitching at any position of the lens in the corresponding FOV.
[0075] refer to Figure 8 , a diagram illustrating an example implementation of a pose calibration process 204 according to an example embodiment of the present invention is shown. In various embodiments, the pose calibration process 204 receives input from channels 0 to N and receives the intrinsic parameters and distortion parameters of each of the lenses obtained in the lens calibration step 202. In an example embodiment, the pose calibration process (or method) 204 includes steps (or states) 210, 212, 214, 216, 218, 220, 222, 224, and 226. The pose calibration process 204 generally begins in step 210.
[0076] In step 210, a circle detector or a checkerboard detector can be used to detect the center of a circle or a corner on the calibration plate. A flat and rigid calibration plate is typically placed between the two views at a short distance from the lens (e.g., 3 meters). In the example where three cameras are used, calibration plate A can be placed at the position shown in the respective FOVs of the left camera and the center camera, and calibration plate B can be placed at the position shown in the respective FOVs of the right camera and the center camera. The detected points in one view should match in the other view. In step 212, an extrinsic calibration is performed for each lens using the intrinsic parameters (e.g., fx, fy, cx, cy, etc.) and distortion parameters (e.g., k1, k2, k3, p1, p2, etc.) from the lens calibration step 202. In the example, the extrinsic parameters typically include a rotation matrix (e.g., R 3x3 ) and the translation vector (e.g., T 3x1 ). In step 214, the z value of the translation vector is typically changed to the desired specific medium or long distance, and the rotation matrix remains the same.
[0077] In step 216, the keypoints are typically projected from world coordinates to image coordinates. In an example, the following equation 2 may be used:
[0078]
[0079] Where fx represents the focal length in the horizontal axis, fy represents the focal length in the vertical axis, x0 represents the center coordinate on the horizontal axis, y0 represents the center coordinate on the vertical axis, and R 3x3 represents the rotation matrix, T 3x1 represents the translation vector, u and v represent the coordinates in the image, X W 、Y W 、Z W represents the real world coordinates, and Z c Represents the depth of a point in camera coordinates.
[0080] In step 218, a homography matrix between two adjacent views can be calculated. Generally, the matrices for all views can be based on the center view (or baseline). In examples where the total number of cameras is two, any camera can be selected as the baseline. In examples where the total number of cameras is an even number and greater than two, the camera closest to the center can be selected as the baseline. In examples where the total number of cameras is an odd number, the center camera is generally selected as the baseline.
[0081] In step 220, the intrinsic parameters and distortion parameters can be applied to the corresponding images. The corresponding images can be warped using a homography matrix. In step 222, a cylindrical projection model can be applied to all views. However, other projection models can be applied to meet the design criteria of a specific implementation. In the example of horizontal stitching, models such as perspective, cylindrical, equirectangular, etc. can be applied. In another example of vertical stitching, models such as transverse cylindrical, Mercator, etc. can be applied.
[0082] In step 224, invalid regions caused by distortion or projection can be removed, each view can be cropped to an inscribed quadrilateral while maintaining the aspect ratio of the quadrilateral, and scaled to near its original size. In step 226, the overlap area of two adjacent views can be calculated for medium and / or long distances. Configuration parameters are typically obtained for the offset / width / height of each view.
[0083] refer to Figure 9, a diagram of a camera system 900 is shown, which illustrates an example implementation of a computer vision system in which a fixed pattern calibration scheme for multi-view stitching according to an example embodiment of the present invention can be implemented. In one example, the electronics of the camera system 900 can be implemented as one or more integrated circuits. In an example, the camera system 900 can be built around a processor / camera chip (or circuit) 902. In an example, the processor / camera chip 902 can be implemented as an application specific integrated circuit (ASIC) or a system on a chip (SOC). The processor / camera circuit 902 is typically combined with hardware and / or software / firmware that can be configured to implement the above combined Figures 1 to 8 Describe the circuit and process.
[0084] In an example, the processor / camera circuit 902 may be connected to a lens and sensor assembly 904. In some embodiments, the lens and sensor assembly 904 may be a component of the processor / camera circuit 902 (e.g., a SoC component). In some embodiments, the lens and sensor assembly 904 may be a separate component from the processor / camera circuit 902 (e.g., the lens and sensor assembly may be an interchangeable component compatible with the processor / camera circuit 902). In some embodiments, the lens and sensor assembly 904 may be part of a separate camera connected to the processor / camera circuit 902 (e.g., via a video cable, a High-Definition Media Interface (HDMI) cable, a Universal Serial Bus (USB) cable, an Ethernet cable, or a wireless link).
[0085] The lens and sensor assembly 904 may include a block (or circuit) 906 and / or a block (or circuit) 908. Circuit 906 may be associated with the lens assembly. Circuit 908 may be implemented as one or more image sensors. In one example, circuit 908 may be implemented as a single sensor. In another example, circuit 908 may be implemented as a stereo pair of sensors. The lens and sensor assembly 904 may include other components (not shown). The number, type, and / or function of the components of the lens and sensor assembly 904 may vary depending on the design criteria of a particular implementation.
[0086] Lens assembly 906 can capture and / or focus light input received from the environment surrounding camera system 900. Lens assembly 906 can capture and / or focus light for image sensor(s) 908. Lens assembly 906 can implement one or more optical lenses. Lens assembly 906 can provide zoom features and / or focusing features. Lens assembly 906 can be implemented with additional circuitry (e.g., a motor) to adjust the direction, zoom, and / or aperture of lens assembly 906. Lens assembly 906 can be oriented, tilted, translated, zoomed, and / or rotated to provide a desired view of the environment surrounding camera system 900.
[0087] Image sensor(s) 908 may receive light from lens assembly 906. Image sensor(s) 908 may be configured to convert the received focused light into digital data (e.g., a bitstream). In some embodiments, image sensor(s) 908 may perform analog-to-digital conversion. For example, image sensor(s) 908 may perform photoelectric conversion on the focused light received from lens assembly 906. Image sensor(s) 908 may present the converted image data as a bitstream in a color filter array (CFA) format. Processor / camera circuitry 902 may convert the bitstream into video data, a video file, and / or a video frame (e.g., human-readable content).
[0088] The processor / camera circuit 902 may also be connected to: (i) optional audio input / output circuitry, which may include an audio codec 910, a microphone 912, and a speaker 914; (ii) memory 916, which may include dynamic random access memory (DRAM); (iii) non-volatile memory (e.g., NAND flash memory) 918, removable media (e.g., SD, SDXC, etc.) 920, one or more serial (e.g., RS-485, RS-232, etc.) devices 922, one or more universal serial bus (USB) devices (e.g., a USB host) 924, and a wireless communication device 926.
[0089] In various embodiments, the processor / camera circuit 902 may include a plurality of blocks (or circuits) 930a-930n, a plurality of blocks (or circuits) 932a-932n, a block (or circuit) 934, a block (or circuit) 936, a block (or circuit) 938, a block (or circuit) 940, a block (or circuit) 942, a block (or circuit) 944, a block (or circuit) 946, a block (or circuit) 948, a block (or circuit) 950, a block (or circuit) 952, and / or a block (or circuit) 954. The plurality of circuits 930a-930n may be processor circuits. In various embodiments, the circuits 930a-930n may include one or more embedded processors (e.g., ARM, etc.). The circuits 932a-932n may implement a plurality of computer vision-related processor circuits. In an example, one or more of the circuits 932a-932n may implement various computer vision-related applications. The circuit 934 may be a digital signal processing (DSP) module. In some embodiments, circuit 934 may implement separate image DSP modules and video DSP modules.
[0090] Circuit 936 may be a storage interface. Circuit 936 may interface the processor / camera circuit 902 with the DRAM 916, non-volatile memory 918, and removable media 920. One or more of the DRAM 916, non-volatile memory 918, and / or removable media 920 may store computer-readable instructions. The computer-readable instructions may be read and executed by processors 930a-930n. In response to the computer-readable instructions, processors 930a-930n may be operable to act as controllers for processors 932a-932n. For example, the resources of processors 932a-932n may be configured to efficiently perform various specific operations in hardware, and processors 930a-930n may be configured to make decisions regarding how to handle input / output to / from the various resources of processor 932.
[0091] Circuit 938 may implement a local memory system. In some embodiments, local memory system 938 may include, but is not limited to, a cache memory (e.g., L2CACHE), a direct memory access (DMA) engine, a graphics direct memory access (GDMA) engine, and fast random access memory. In an example, DAG memory 968 may be implemented in local memory system 938. Circuit 940 may implement sensor inputs (or interfaces). Circuit 942 may implement one or more control interfaces, including, but not limited to, an inter-device communication (IDC) interface, an inter-integrated circuit (I2C) interface, a serial peripheral interface (SPI), and a pulse width modulation (PWM) interface. Circuit 944 may implement an audio interface (e.g., an I2S interface, etc.). Circuit 946 may implement clock circuits, including, but not limited to, a real-time clock (RTC), a watchdog timer (WDT), and / or one or more programmable timers. Circuit 948 may implement an input / output (I / O) interface. Circuit 950 may be a video output module. Circuit 952 may be a communications module. Circuit 954 may be a security module. Circuits 930 - 954 may be connected to each other using one or more buses, interfaces, traces, protocols, or the like.
[0092] Circuit 918 may be implemented as a non-volatile memory (e.g., NAND flash memory, NOR flash memory, etc.). Circuit 920 may include one or more removable media cards (e.g., Secure Digital (SD), Secure Digital Extended Capacity (SDXC), etc.). Circuit 922 may include one or more serial interfaces (e.g., RS-485, RS-232, etc.). Circuit 924 may be an interface for connecting to a Universal Serial Bus (USB) host or acting as a USB host. Circuit 926 may be a wireless interface for communicating with a user device (e.g., a smartphone, a computer, a tablet computing device, a cloud resource, etc.). In various embodiments, circuits 904-926 may be implemented as components external to processor / camera circuit 902. In some embodiments, circuits 904-926 may be on-board components of processor / camera circuit 902.
[0093] The control interface 942 may be configured to generate signals (e.g., IDC / I2C, STEPPER, IRIS, AF / ZOOM / TILT / PAN, etc.) for controlling the lens and sensor assembly 904. The signal IRIS may be configured to adjust the aperture of the lens assembly 906. The interface 942 may enable the processor / camera circuitry 902 to control the lens and sensor assembly 904.
[0094] The storage interface 936 can be configured to manage one or more types of storage and / or data access. In one example, the storage interface 936 can implement a direct memory access (DMA) engine and / or graphics direct memory access (GDMA). In another example, the storage interface 936 can implement a secure digital (SD) card interface (e.g., to connect to the removable media 920). In various embodiments, programming code (e.g., executable instructions for controlling the various processors and encoders of the processor / camera circuit 902) can be stored in one or more of the memories (e.g., DRAM 916, NAND 918, etc.). When executed by one or more of the processors 930, the programming code typically causes one or more components in the processor / camera circuit 902 to configure video synchronization operations and initiate video frame processing operations. The resulting compressed video signal can be presented to the storage interface 936, the video output 950, and / or the communication interface 952. The storage interface 936 can transfer program code and / or data between external media (e.g., DRAM 916, NAND 918, removable media 920, etc.) and the local (internal) memory system 938.
[0095] The sensor input 940 can be configured to send data to / receive data from the image sensor 908. In one example, the sensor input 940 can include an image sensor input interface. The sensor input 940 can be configured to send a captured image (e.g., picture elements, pixels, data) from the image sensor 908 to the DSP module 934, one or more of the processors 930, and / or one or more of the processors 932. The data received by the sensor input 940 can be used by the DSP 934 to determine luma (Y) and chroma (U and V) values from the image sensor 908. The sensor input 940 can provide an interface to the lens and sensor assembly 904. The sensor input interface 940 can enable the processor / camera circuit 902 to capture image data from the lens and sensor assembly 904.
[0096] The audio interface 944 may be configured to transmit / receive audio data. In one example, the audio interface 944 may implement an audio inter-IC sound (I2S) interface. The audio interface 944 may be configured to transmit / receive data in a format implemented by the audio codec 910.
[0097] The DSP module 934 can be configured to process digital signals. The DSP module 934 can include an image digital signal processor (IDSP), a video digital signal processor DSP (VDSP), and / or an audio digital signal processor (ADSP). The DSP module 934 can be configured to receive information from the sensor input 940 (e.g., pixel data values captured by the image sensor 908). The DSP module 934 can be configured to determine pixel values (e.g., RGB, YUV, luminance, chrominance, etc.) based on the information received from the sensor input 940. The DSP module 934 can be further configured to support or provide a sensor RGB to YUV raw image pipeline to improve image quality, perform bad pixel detection and correction, demosaicing, white balance, color and tone correction, gamma correction, hue adjustment, saturation, brightness and contrast adjustment, and chroma and luminance noise filtering.
[0098] The I / O interface 948 may be configured to send / receive data. The data sent / received by the I / O interface 948 may be miscellaneous information and / or control data. In one example, the I / O interface 948 may implement one or more of a general-purpose input / output (GPIO) interface, an analog-to-digital converter (ADC) module, a digital-to-analog converter (DAC) module, an infrared (IR) remote interface, a pulse width modulation (PWM) module, a universal asynchronous receiver transmitter (UART), an infrared (IR) remote interface, and / or one or more synchronous data communication interfaces (IDC SPI / SSI).
[0099] The video output module 950 can be configured to transmit video data. For example, the processor / camera circuit 902 can be connected to an external device (e.g., a television, a monitor, a laptop computer, a tablet computing device, etc.). The video output module 950 can implement a High-Definition Multimedia Interface (HDMI), a PAL / NTSC interface, an LCD / TV / parallel interface, and / or a DisplayPort interface.
[0100] The communication module 952 may be configured to send / receive data. The data sent / received by the communication module 952 may be transmitted / received according to a specific protocol (e.g., In one example, the communication module 952 may implement a secure digital input output (SDIO) interface. The communication module 952 may include a communication interface that receives data transmitted via one or more wireless protocols (e.g., The processor / camera circuit 902 may also include support for wireless communications using Z-Wave, LoRa, Institute of Electrical and Electronics Engineers (IEEE) 802.11a / b / g / n / ac (WiFi), IEEE 802.15, IEEE 802.15.1, IEEE 802.15.2, IEEE 802.15.3, IEEE 802.15.4, IEEE 802.15.5, and / or IEEE 802.20, GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, 5G, LTE-M, NB-IoT, SMS, etc. The communication module 952 may also include support for communication using one or more universal serial bus protocols (e.g., USB 1.0, 2.0, 3.0, etc.). The processor / camera circuit 902 may also be configured to be powered via a USB connection. However, other communication and / or power interfaces may be implemented accordingly to meet the design criteria of a particular application.
[0101] The security module 954 may include a set of advanced security features to achieve advanced on-device physical security, including OTP, secure boot, As well as I / O visualization and DRAM scrambling. In an example, the security module 958 can include a true random number generator. In an example, the security module 954 can be used for DRAM communication encryption on the processor / camera circuit 902.
[0102] The processor / camera circuit 902 can be configured (e.g., programmed) to control one or more lens assemblies 906 and one or more image sensors 908. The processor / camera circuit 902 can receive raw image data from the image sensor(s) 908. The processor / camera circuit 902 can simultaneously (in parallel) encode the raw image data into multiple encoded video streams. The multiple video streams can have various resolutions (e.g., VGA, WVGA, QVGA, SD, HD, Ultra HD, 4K, etc.). The processor / camera circuit 902 can receive encoded and / or unencoded (e.g., raw) audio data at the audio interface 944. The processor / camera circuit 902 can also receive encoded audio data from the communication interface 952 (e.g., USB and / or SDIO). The processor / camera circuit 902 can provide the encoded video data to the wireless interface 926 (e.g., using a USB host interface). The wireless interface 926 can include support for video signals transmitted via one or more wireless and / or cellular protocols (e.g., The processor / camera circuit 902 may also include support for wireless communications using Z-Wave, LoRa, Wi-Fi (IEEE 802.11a / b / g / n / ac, IEEE 802.15, IEEE 802.15.1, IEEE 802.15.2, IEEE 802.15.3, IEEE 802.15.4, IEEE 802.15.5, IEEE 802.20, GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, 5G, SMS, LTE-M, NB-IoT, etc. The processor / camera circuit 902 may also include support for communication using one or more universal serial bus protocols (e.g., USB 1.0, 2.0, 3.0, etc.).
[0103] refer to Figure 10 , a diagram illustrating a processing circuit 902 in the context of a fixed pattern calibration scheme for multi-view stitching according to an example embodiment of the present invention is shown. In various embodiments, the processing circuit 902 can be implemented as part of a computer vision system. In various embodiments, the processing circuit 902 can be implemented as part of a camera, a computer, a server (e.g., a cloud server), a smart phone (e.g., a cellular phone), a personal digital assistant, etc. In an example, the processing circuit 902 can be configured for applications including, but not limited to, autonomous and semi-autonomous vehicles (e.g., cars, trucks, motorcycles, agricultural machinery, drones, aircraft, etc.), manufacturing and / or security and surveillance systems. In contrast to a general-purpose computer, the processing circuit 902 typically includes hardware circuitry that is optimized to provide high-performance image processing and computer vision pipelines in a minimal area and with minimal power consumption. In an example, various operations for performing image processing, feature detection / extraction, and / or object detection / classification for computer (or machine) vision can be implemented using hardware modules designed to reduce computational complexity and use resources efficiently.
[0104] In an example embodiment, processing circuit 902 may include block (or circuit) 930i, block (or circuit) 932i, block (or circuit) 916, and / or memory bus 917. Circuit 930i may implement a first processor. Circuit 932i may implement a second processor. In an example, circuit 932i may implement a computer vision processor. In an example, processor 932i may be an intelligent vision processor. Circuit 916 may implement external memory (e.g., memory external to circuits 930i and 932i). In an example, circuit 916 may be implemented as a dynamic random access memory (DRAM) circuit. Processing circuit 902 may include other components (not shown). The number, type, and / or arrangement of components of processing circuit 902 may vary depending on the design criteria of a particular implementation.
[0105] Circuit 930i may implement a processor circuit. In some embodiments, processor circuit 930i may be implemented using a general-purpose processor circuit. Processor 930i may be operable to interact with circuit 932i and circuit 916 to perform various processing tasks. In an example, processor 930i may be configured as a controller for circuit 932i. Processor 930i may be configured to execute computer-readable instructions. In one example, the computer-readable instructions may be stored by circuit 916. In some embodiments, the computer-readable instructions may include controller operations. Processor 930i may be configured to communicate with circuit 932i and / or access results generated by components of circuit 932i. In an example, processor 930i may be configured to utilize circuit 932i to perform operations associated with one or more neural network models.
[0106] In an example, the processor 930i can be configured to program the circuit 932i using a fixed pattern calibration (FPC) 200 for a multi-view stitching scheme. In various embodiments, the FPC technology 200 can be configured for operation in an edge device. In an example, the processing circuit 902 can be coupled to a sensor (e.g., a video camera, etc.) configured to generate a data input. The processing circuit 902 can be configured to generate one or more outputs in response to the data input from the sensor. The data input can be processed by the FPC technology 200. The operations performed by the processor 930i can vary according to the design criteria of the specific implementation.
[0107] In various embodiments, circuitry 916 may implement dynamic random access memory (DRAM) circuitry. Circuitry 916 is generally operable to store multi-dimensional arrays of input data elements and various forms of output data elements. Circuitry 916 may exchange input data elements and output data elements with processor 930 i and processor 932 i.
[0108] The processor 932i may implement a computer vision processor circuit. In an example, the processor 932i may be configured to implement various functions for computer vision. The processor 932i is generally operable to perform specific processing tasks assigned by the processor 930i. In various embodiments, all or part of the processor 932i may be implemented solely in hardware. The processor 932i may directly execute a data stream generated by software (e.g., a directed acyclic graph, etc.) for a fixed pattern calibration scheme for multi-view stitching and a specified processing (e.g., computer vision) task. In some embodiments, the processor 932i may be a representative example of several computer vision processors implemented by the processing circuit 902 and configured to operate together.
[0109] In an example embodiment, processor 932i generally includes block (or circuit) 960, one or more blocks (or circuits) 962a-962n, block (or circuit) 960, path 966, and block (or circuit) 968. Block 960 can implement a scheduler circuit. Blocks 962a-962n can implement hardware resources (or engines). Block 964 can implement a shared memory circuit. Block 968 can implement a directed acyclic graph (DAG) memory. In an example embodiment, one or more of circuits 962a-962n can include blocks (or circuits) 970a-970n. In the example shown, circuits 970a, 970b, and 970n are implemented.
[0110] In an example embodiment, circuit 970a may implement a convolution operation, circuit 970b may be configured to provide an n-dimensional (nD) dot product operation, and circuit 970n may be configured to perform a transcendental operation. Circuits 970a-970n may be utilized to provide a fixed pattern calibration scheme for multi-view stitching according to an example embodiment of the present invention. Convolution, nD dot product, and transcendental operations may be used to perform computer (or machine) vision tasks (e.g., as part of an object detection process, etc.). In yet another example, one or more of circuits 962c-962n may include blocks (or circuits) 970c-970n (not shown) to provide multi-dimensional convolution calculations.
[0111] In an example, circuit 932i may be configured to receive a directed acyclic graph (DAG) from processor 930i. The DAG received from processor 930i may be stored in DAG memory 968. Circuit 932i may be configured to execute the DAG for a fixed pattern calibration scheme for multi-view stitching using circuits 960, 962a-962n, and 964.
[0112] A plurality of signals (e.g., OP_A to OP_N) may be exchanged between circuit 960 and corresponding circuits 962a-962n. Each signal OP_A to OP_N may convey information about an execution operation and / or a generation operation. A plurality of signals (e.g., MEM_A to MEM_N) may be exchanged between corresponding circuits 962a-962n and circuit 964. Signals MEM_A to MEM_N may carry data. Signals (e.g., DRAM) may be exchanged between circuit 916 and circuit 964. Signal DRAM may transmit data between circuits 916 and 960 (e.g., on memory bus 966).
[0113] Circuit 960 may implement a scheduler circuit. Scheduler circuit 960 is generally operable to schedule tasks between circuits 962a-962n to perform various computer vision-related tasks defined by processor 930i. Scheduler circuit 960 may assign individual tasks to circuits 962a-962n. Scheduler circuit 960 may assign individual tasks in response to parsing a directed acyclic graph (DAG) provided by processor 930i. Scheduler circuit 960 may time-multiplex tasks to circuits 962a-962n based on the availability of circuits 962a-962n to perform the work.
[0114] Each circuit 962a-962n can implement a processing resource (or hardware engine). The hardware engines 962a-962n are generally operable to perform specific processing tasks. The hardware engines 962a-962n can be implemented to include dedicated hardware circuits that are optimized to have high performance and low power consumption when performing specific processing tasks. In some configurations, the hardware engines 962a-962n can operate in parallel and independently of each other. In other configurations, the hardware engines 962a-962n can operate together with each other to perform assigned tasks.
[0115] The hardware engines 962a-962n may be homogeneous processing resources (e.g., all circuits 962a-962n may have the same capabilities) or heterogeneous processing resources (e.g., two or more circuits 962a-962n may have different capabilities). The hardware engines 962a-962n are generally configured to execute operators, which may include, but are not limited to, resampling operators, warping operators, component operators that manipulate lists of components (e.g., components may be regions of a vector that share common properties and may be grouped together with bounding boxes), matrix inverse operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, min-max-reduce operators, pooling operators, non-min-non-max suppression operators, aggregation operators, scattering operators, statistical operators, classifier operators, integral image operators, upsampling operators, and downsampling operators that are powers of two, among others.
[0116] In various embodiments, the hardware engines 962a-962n can be implemented as hardware circuits only. In some embodiments, the hardware engines 962a-962n can be implemented as general-purpose engines that can be configured to operate as special-purpose machines (or engines) through circuit customization and / or software / firmware. In some embodiments, the hardware engines 962a-962n can be alternatively implemented as one or more instances or threads of program code executed on the processor 930i and / or one or more processors 932i (including but not limited to vector processors, central processing units (CPUs), digital signal processors (DSPs), or graphics processing units (GPUs)). In some embodiments, the scheduler 960 can select one or more of the hardware engines 962a-962n for a specific process and / or thread. The scheduler 960 can be configured to assign the hardware engines 962a-962n to specific tasks in response to parsing the directed acyclic graph stored in the DAG memory 968.
[0117] Circuit 964 can implement shared memory circuitry. Shared memory 964 can be configured to store data in response to input requests and / or present data in response to output requests (e.g., requests from processor 930i, DRAM 916, scheduler circuit 960, and / or hardware engines 962a-962n). In an example, shared memory circuit 964 can implement on-chip memory for computer vision processor 932i. Shared memory 964 is generally operable to store all or part of a multidimensional array (or vector) of input data elements and output data elements generated and / or utilized by hardware engines 962a-962n. Input data elements can be transferred from DRAM circuit 916 to shared memory 964 via memory bus 917. Output data elements can be sent from shared memory 964 to DRAM circuit 916 via memory bus 917.
[0118] Path 966 may implement a transfer path within processor 932 i. Transfer path 966 is generally operable to move data from scheduler circuit 960 to shared memory 964. Transfer path 966 may also be operable to move data from shared memory 964 to scheduler circuit 960.
[0119] Processor 930i is shown in communication with computer vision processor 932i. Processor 930i can be configured as a controller for computer vision processor 932i. In some embodiments, processor 930i can be configured to transmit instructions to scheduler 960. For example, processor 930i can provide one or more directed acyclic graphs to scheduler 960 via DAG memory 968. Scheduler 960 can initialize and / or configure hardware engines 962a-962n in response to parsing the directed acyclic graphs. In some embodiments, processor 930i can receive status information from scheduler 960. For example, scheduler 960 can provide status information and / or readiness of outputs from hardware engines 962a-962n to processor 930i to enable processor 930i to determine one or more next instructions to be executed and / or decisions to be made. In some embodiments, processor 930i can be configured to communicate with shared memory 964 (e.g., directly or through scheduler 960, which receives data from shared memory 964 via path 966). Processor 930i can be configured to retrieve information from shared memory 964 to make decisions. The instructions executed by processor 930i in response to information from computer vision processor 932i can vary according to the design criteria of a particular implementation.
[0120] Circuit 970a can implement a convolution circuit. Convolution circuit 970a can communicate with memory 964 to receive input data and present output data. Convolution circuit 970a is generally operable to fetch multiple data vectors from shared memory circuit 964. Each data vector can include multiple data values. Convolution circuit 970a can also be operable to fetch a kernel from shared memory 964. The kernel typically includes multiple kernel values. Convolution circuit 970a can also be operable to fetch a block from shared memory 964 into an internal (or local) buffer. The block typically includes multiple input tiles. Each input tile can include multiple input values of multiple dimensions. Convolution circuit 970a can also be operable to calculate multiple intermediate values in parallel by multiplying each input tile in the internal buffer by a corresponding one of the kernel values, and to calculate an output tile including multiple output values based on the intermediate values. In various embodiments, convolution circuit 970a can be implemented solely in hardware. An example of a convolution computation scheme that can be used to implement circuit 970a can be found in U.S. Patent No. 10,210,768, which is incorporated herein by reference in its entirety. Circuit 970b can implement an nD dot product process. Circuit 970n can implement a transcendental operation process. In various embodiments, a fixed pattern calibration scheme for multi-view stitching according to embodiments of the present invention can be performed according to the implementation description provided herein.
[0121] refer to Figure 11 , shows the description Figure 10 FIG. 1 is a diagram of an example implementation of a general hardware engine 962x. Hardware engine 962x may represent hardware engines 962a-962n. Hardware engine 962x generally includes a block (or circuit) 980, a block (or circuit) 982, a block (or circuit) 984, and a plurality of blocks (or circuits) 986a-986n. Circuit 980 may be implemented as a memory (or buffer) pair 980a and 980b. Circuit 982 may implement a controller circuit. In an example, circuit 982 may include one or more finite state machines (FSMs) configured to control various operators implemented by hardware engine 962x. Circuit 984 may implement a processing pipeline for hardware engine 962x. Circuits 986a-986n may implement a first-in, first-out (FIFO) memory. Circuits 986a-986n may be configured as input buffers for processing pipeline 984. Shared memory 964 may be configured (eg, by signals from circuitry 982 ) as a plurality of shared input buffers 988 a - 988 n and one or more output buffers 990 .
[0122] Signals (e.g., ADDR / CONFIG) may be generated by scheduler circuit 960 and received by hardware engine 962x. Signal ADDR / CONFIG may carry address information and configuration data. Signals (e.g., BUSY_LEVEL) may be generated by circuit 982 and transmitted to scheduler circuit 960. Signal BUSY_LEVEL may convey the busy level of hardware engine 962x. Signals (e.g., STATUS / TARGETS) may be generated by circuit 982 and transmitted to scheduler circuit 960. Signal STATUS / TARGETS may provide status information about hardware engine 962x and target information for operands.
[0123] In an example embodiment, buffers 980a and 980b can be configured as dual-bank configuration buffers. The dual-bank buffers can be operable to store configuration information for a currently running operation in one buffer (e.g., buffer 980b) and move configuration information for the next operation to another buffer (e.g., buffer 980a). Scheduler 960 typically loads operator configuration information (including status words in the case where an operator has been partially processed in a previous operator block) into the dual-bank buffers. Once circuit 982 completes configuration information for the currently running operation and has received configuration information for the next operation, buffers 980a and 980b can be swapped.
[0124] Circuitry 982 typically implements control circuitry for hardware engine 962x. Circuitry 982 determines when to switch from a currently running operator to a new operator. Controller 982 is typically operable to control the movement of information into, out of, and within hardware engine 982x. Typically, the operation of hardware engine 962x is pipelined. During an operator switch, the front end of pipeline 984 may already be processing data for the new operator while the tail end of pipeline 984 is still completing processing associated with the old operator.
[0125] The circuitry 984 may implement pipeline circuitry. The pipeline circuitry 984 is generally operable to process operands received from the shared memory 964 using functions designed into the hardware engines 962x. The circuitry 984 may transfer data generated by the executed functions to one or more shared buffers 990.
[0126] Buffers 986a-986n may implement FIFO buffers. FIFO buffers 986a-986n may be operable to store operands received from shared buffers 988a-988n for processing in pipeline 984. In general, the number of FIFO buffers and the number of shared buffers implemented may vary to meet the design criteria of a particular application.
[0127] As will be appreciated by those skilled in the relevant art(s), one or more of conventional general-purpose processors, digital computers, microprocessors, microcontrollers, RISC (Reduced Instruction Set Computer) processors, CISC (Complex Instruction Set Computer) processors, SIMD (Single Instruction Multiple Data) processors, signal processors, central processing units (CPUs), arithmetic logic units (ALUs), video digital signal processors (VDSPs), distributed computer resources, and / or similar computing machines programmed according to the teachings of this specification may be used to design, model, simulate, and / or emulate the system. Figure 1-11 The functions performed by the diagrams and the structures shown. As will be apparent to those skilled in the relevant art(s), skilled programmers can readily prepare appropriate software, firmware, coding, routines, instructions, opcodes, microcodes, and / or program modules based on the teachings of this disclosure. The software is typically embodied in one or more media (e.g., non-transitory storage media) and can be executed sequentially or in parallel by one or more processors.
[0128] Embodiments of the present invention may also be implemented in one or more of an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), a PLD (programmable logic device), a CPLD (complex programmable logic device), an ocean of gates, an ASSP (application specific standard product), and an integrated circuit. The circuit may be implemented based on one or more hardware description languages. Embodiments of the present invention may be utilized in conjunction with flash memory, nonvolatile memory, random access memory, read-only memory, magnetic disks, floppy disks, optical disks such as DVDs and DVD RAMs, magneto-optical disks, and / or distributed storage systems.
[0129] When the terms "may" and "typically" are used herein in conjunction with the verb "is," it is intended to convey the intention that the description is exemplary and is considered broad enough to encompass both the specific examples set forth in this disclosure and alternative examples that can be derived based on this disclosure. The terms "may" and "typically" as used herein should not be interpreted as necessarily implying the desirability or possibility of omitting the corresponding element.
[0130] While the invention has been particularly shown and described with reference to embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the invention.
Claims
1. A device comprising: an interface configured to receive video signals from two or more cameras arranged to obtain a predetermined field of view, wherein the respective fields of view of each pair of the two or more cameras overlap; and A processor configured to perform fixed pattern calibration for facilitating multi-view stitching, wherein the fixed pattern calibration comprises: (a) performing a geometric calibration process to obtain intrinsic parameters and distortion parameters of each lens of the two or more cameras; and (b) applying a pose calibration process to the video signals using (i) the intrinsic parameters, corresponding rotation matrices, corresponding translation vectors, and the distortion parameters of each lens of the two or more cameras and (ii) a calibration plate to obtain configuration parameters for the corresponding fields of view of the two or more cameras; The posture calibration process further includes: changing the z value of the corresponding translation vector of each lens of the two or more cameras to a medium distance value or a long distance value, while maintaining the corresponding rotation matrix of each lens of the two or more cameras unchanged.
2. The device according to claim 1, wherein The posture calibration process includes: using at least one of a circle detector and a checkerboard detector to respectively detect a center of a circle or a corner on the calibration plate; and Determines whether detected points in one view have matches in the other view.
3. The device according to claim 1, wherein The posture calibration process also includes: using the intrinsic parameters and the distortion parameters of each lens of the two or more cameras to perform non-intrinsic calibration for each lens of the two or more cameras, wherein the non-intrinsic parameters of each lens of the two or more cameras include the corresponding rotation matrix and the corresponding translation vector.
4. The device according to claim 1, wherein The pose calibration process includes projecting keypoints from world coordinates to image coordinates.
5. The device according to claim 1, wherein The pose calibration process includes computing a corresponding homography matrix between each adjacent pair of the two or more cameras.
6. The device according to claim 5, wherein The corresponding homography matrix between each adjacent pair of the two or more cameras is based on a central view.
7. The device according to claim 5, wherein The posture calibration process includes: applying the intrinsic parameters and the distortion parameters of each lens of the two or more cameras to corresponding images; and The corresponding images are warped using the corresponding homography matrices.
8. The device according to claim 1, wherein The pose calibration process includes applying a projective model to the views of the two or more cameras.
9. The apparatus according to claim 8, wherein: For horizontal stitching, the projection model includes at least one of a perspective model, a cylindrical model, and an equirectangular model; and For vertical stitching, the projection model includes at least one of a transverse cylindrical model and a Mercator model.
10. A method for fixed pattern calibration for multi-view stitching using multiple cameras, comprising: arranging two or more cameras to obtain a predetermined field of view, wherein respective fields of view of each adjacent pair of the two or more cameras overlap; performing a geometric calibration process to obtain intrinsic parameters and distortion parameters of each lens of the two or more cameras; and applying a pose calibration process to the video signals from the two or more cameras using (i) the intrinsic parameters of each lens of the two or more cameras, the corresponding rotation matrix, the corresponding translation vector, and the distortion parameters and (ii) a calibration plate to obtain configuration parameters for the corresponding fields of view of the two or more cameras; The posture calibration process further includes: changing the z value of the corresponding translation vector of each lens of the two or more cameras to a medium distance value or a long distance value, while maintaining the corresponding rotation matrix of each lens of the two or more cameras unchanged.
11. The method according to claim 10, wherein: The posture calibration process includes: using at least one of a circle detector and a checkerboard detector to respectively detect a center of a circle or a corner on the calibration plate; and Determines whether detected points in one view have matches in the other view.
12. The method according to claim 10, wherein: The posture calibration process also includes: using the intrinsic parameters and the distortion parameters of each lens of the two or more cameras to perform non-intrinsic calibration for each lens of the two or more cameras, wherein the non-intrinsic parameters of each lens of the two or more cameras include the corresponding rotation matrix and the corresponding translation vector.
13. The method according to claim 10, wherein: The pose calibration process includes projecting keypoints from world coordinates to image coordinates.
14. The method according to claim 10, wherein: The pose calibration process includes computing a corresponding homography matrix between each adjacent pair of the two or more cameras.
15. The method according to claim 14, wherein The corresponding homography matrix between each adjacent pair of the two or more cameras is based on a central view.
16. The method according to claim 14, wherein The posture calibration process includes: applying the intrinsic parameters and the distortion parameters of each lens of the two or more cameras to corresponding images; and The corresponding images are warped using the corresponding homography matrices.
17. The method according to claim 10, wherein: The pose calibration process includes applying a projective model to the views of the two or more cameras.
18. The method according to claim 17, wherein: For horizontal stitching, the projection model includes at least one of a perspective model, a cylindrical model, and an equirectangular model; and For vertical stitching, the projection model includes at least one of a transverse cylindrical model and a Mercator model.
19. The method of claim 17, wherein the posture calibration process comprises: Each view is clipped to the inscribed quadrilateral while maintaining the quadrilateral's aspect ratio and scaled to approximate the view's original size.
20. The method of claim 10, wherein the configuration parameters are obtained for one or more of an offset, a width, and a height of respective fields of view of the two or more cameras.
Citation Information
Patent Citations
Systems and methods for customizing a learning experience of a user
US10210768B2
Method and device for calibrating dual fisheye lens panoramic camera, and storage medium and terminal thereof
US20190236805A1