Installation and calibration of structured light projectors in monocular camera stereo systems

By generating parallax angle maps and applying directional block matching algorithms, the problem of limited installation position of structured light projectors in monocular camera stereo system is solved, the system flexibility and installation accuracy are improved, and production costs are reduced.

CN115222782BActive Publication Date: 2025-08-08AMBARELLA INT LP
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202110410126.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-04-16
Publication Date
2025-08-08
Estimated Expiration
2041-04-16

AI Technical Summary

Technical Problem

The installation position of the structured light projector in the monocular camera stereo system is limited, resulting in strong dependence on parallax measurement, limiting the flexibility and installation accuracy of the system and increasing production costs.

Method used

The reference image and target image are generated through the interface and the processor, the parallax operation is performed, the parallax angle map is constructed, the pattern shift matching is used for parallax angle map, and the directional block matching algorithm is used to realize the installation calibration of the structured light projector.

Benefits of technology

The dependence of structured light projector installation on position is reduced, parallax angle map is generated, directional block matching is realized, system flexibility and installation accuracy are improved, and production costs are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115222782B_ABST
    Figure CN115222782B_ABST
Patent Text Reader

Abstract

An apparatus includes an interface and a processor. The interface may be configured to receive pixel data. The processor may be configured to: (i) generate a reference image and a target image based on the pixel data, (ii) perform a parallax operation on the reference image and the target image, and (iii) construct a parallax angle map in response to the parallax operation. The parallax operation may include: (a) selecting a plurality of grid pixels, (b) measuring a parallax angle for each grid pixel, (c) calculating a plurality of coefficients by solving a surface formula for the parallax angle map for the grid pixels, and (d) generating values in the parallax angle map for the pixel data using the coefficients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates generally to stereo vision, and more particularly to a method and / or apparatus for achieving installation calibration of a structured light projector in a mono camera stereo system. Background Art

[0002] A structured light projector is a point light source that spreads spots of light onto an object. When the object moves toward or away from the camera, the speckle pattern captured by the camera sensor will shift horizontally. If a reference image of the speckle pattern on the object is captured at a certain distance (R), and the object moves from distance R to a second distance (X), the movement distance (M=XR) can be calculated by measuring the shift of the speckle pattern captured by the camera sensor. The measurement of the shift of the speckle pattern between the two captured images is the principle of operation of a monocular camera stereo system. The shift offset of the speckle pattern is often called disparity.

[0003] Compared to stereo solutions utilizing two cameras, monocular stereo systems typically perform at a lower cost, smaller size, and lower power. Despite these advantages, monocular stereo systems limit the placement of the structured light projector relative to the camera.

[0004] It is expected to achieve the installation calibration of the structured light projector in a monocular camera stereo system. Summary of the Invention

[0005] The present invention covers aspects relating to an apparatus comprising an interface and a processor. The interface may be configured to receive pixel data. The processor may be configured to: (i) generate a reference image and a target image based on the pixel data, (ii) perform a parallax operation on the reference image and the target image, and (iii) construct a parallax angle map in response to the parallax operation. The parallax operation may include: (a) selecting a plurality of grid pixels, (b) measuring a parallax angle for each grid pixel, (c) calculating a plurality of coefficients by solving a surface formula for the parallax angle map for the grid pixels, and (d) generating values in the parallax angle map for the pixel data using the coefficients.

[0006] In some embodiments of the apparatus aspects described above, the processor may be further configured to construct the disparity map by performing pattern shift matching using the disparity angle map. In some embodiments in which the processor performs pattern shift matching using the disparity angle map, the processor may be further configured to perform pattern shift matching by applying a directional block matching process using the disparity angle map.

[0007] In some embodiments of the apparatus aspects described above, the processor may further be configured to construct a disparity angle map by executing an offline calibration procedure comprising capturing a first image comprising a speckled pattern projected onto a wall at a first distance, and capturing a second image comprising a speckled pattern projected onto the wall at a second distance.

[0008] In some embodiments of the apparatus aspects described above, the processor may further be configured to construct a disparity angle map by executing an online calibration procedure comprising capturing a first image comprising a speckle pattern projected onto an object at a first distance, and capturing a second image comprising a speckle pattern projected onto an object at a second distance.

[0009] In some embodiments of the apparatus aspects described above, the processor may be further configured to calculate the plurality of coefficients by applying a regression algorithm to a parametric surface comprising the grid pixels to solve a surface formula for the disparity angle map of the grid pixels. In some embodiments, the regression algorithm comprises a least squares regression algorithm. In some embodiments, the parametric surface comprises a cubic parametric surface.

[0010] In some embodiments of the apparatus aspects described above, the apparatus may further include a camera configured to generate pixel data and a structured light projector configured to project a speckle pattern. In some embodiments in which the apparatus further includes a camera and a structured light projector, the processor may further be configured to: generate a reference image including the speckle pattern projected onto an object at a first distance from the camera, generate a target image including the speckle pattern projected onto the object at a second distance from the camera, and measure a disparity angle for each grid pixel by determining a pattern shift between the speckle pattern in the reference image and the speckle pattern in the target image.

[0011] The present invention also covers aspects relating to a method for mounting and calibrating a structured light projector in a monocular camera stereo system, the method comprising: (i) receiving pixel data at an interface, (ii) generating a reference image and a target image based on the pixel data using a processor, (iii) performing a parallax operation on the reference image and the target image, and (iv) constructing a parallax angle map in response to the parallax operation, wherein the parallax operation comprises: (a) selecting a plurality of grid pixels, (b) measuring a parallax angle for each grid pixel, (c) calculating a plurality of coefficients by solving a surface formula for the parallax angle map for the grid pixels, and (d) generating values in the parallax angle map for the pixel data using the coefficients.

[0012] In some embodiments of the method aspects described above, the method further comprises constructing the disparity map using the processor by performing pattern shift matching using the disparity angle map. In some embodiments, the method further comprises performing pattern shift matching by applying a directional block matching process using the disparity angle map.

[0013] In some embodiments of the method aspects described above, the method also includes constructing a disparity angle map by performing an offline calibration procedure, wherein the offline calibration procedure includes: capturing a first image comprising a speckle pattern projected onto a wall at a first distance, and capturing a second image comprising the speckle pattern projected onto the wall at a second distance.

[0014] In some embodiments of the method aspects described above, the method also includes constructing a disparity angle map by performing an online calibration procedure, wherein the online calibration procedure includes: capturing a first image comprising a speckle pattern projected onto an object at a first distance, and capturing a second image comprising the speckle pattern projected onto an object at a second distance.

[0015] In some embodiments of the method aspects described above, the method further comprises solving a surface formula for the disparity angle map of the grid pixels by applying a regression algorithm to a parametric surface comprising the grid pixels, thereby calculating the plurality of coefficients. In some embodiments, the regression algorithm comprises a least squares regression algorithm. In some embodiments, the parametric surface comprises a cubic parametric surface.

[0016] In some embodiments of the method aspects described above, the method further comprises generating pixel data using a camera and projecting a speckle pattern onto the object using a structured light projector. In some embodiments, the method further comprises: (i) generating a reference image comprising the speckle pattern projected onto the object at a first distance from the camera, (ii) generating a target image comprising the speckle pattern projected onto the object at a second distance from the camera, and (iii) measuring, using a processor, a disparity angle for each grid pixel to determine a pattern shift between the speckle pattern in the reference image and the speckle pattern in the target image. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Embodiments of the present invention will become apparent from the following detailed description, appended claims and accompanying drawings.

[0018] Figure 1 is a diagram illustrating a monocular camera stereo system according to an example embodiment of the present invention.

[0019] Figure 2 is a diagram illustrating elements of a monocular camera stereo system according to an example embodiment of the present invention.

[0020] Figure 3is a diagram showing example geometry of a monocular camera stereo system.

[0021] Figure 4 It is shown by Figure 3 Figure 3. A structured light projector projects a structured light (or speckle) pattern onto a wall.

[0022] Figure 5 is a diagram illustrating disparity determination for a monocular camera stereo system using a well-mounted structured light projector.

[0023] Figure 6 is a diagram showing how the disparity determination of a monocular camera stereo system changes when the structured light projector is mounted off the x-axis.

[0024] Figure 7 is a diagram showing how the disparity determination of a monocular camera stereo system changes when the structured light projector is moved along the z-axis.

[0025] Figure 8 is a diagram illustrating block matching in a local range.

[0026] Figure 9 is a flowchart illustrating a calibration process according to another example embodiment of the present invention.

[0027] Figure 10 is a flow chart illustrating a process for measuring disparity angles for grid pixels.

[0028] Figure 11 is a flow chart illustrating a process for generating a disparity angle map for all pixels in a target image.

[0029] Figure 12 is a diagram illustrating results between performing a typical block matching routine and performing a directional block matching routine according to an example embodiment of the present invention.

[0030] Figure 13 is a diagram illustrating an apparatus according to another example embodiment of the present invention. DETAILED DESCRIPTION

[0031] Embodiments of the present invention include providing a mounting calibration for a structured light projector in a monocular camera stereo system that can (i) overcome mounting limitations, (ii) reduce the dependency of disparity measurements on the mounting of the structured light projector, (iii) generate a disparity angle map, (iv) implement an oriented block matching (OBM) algorithm, and / or (v) be implemented as one or more integrated circuits.

[0032] refer to Figure 1, a block diagram of an apparatus illustrating an example implementation of a monocular camera stereo system according to an example embodiment of the present invention is shown. In an example, system 100 can implement a 3D sensing platform that includes a monocular camera stereo system using an infrared (IR) or RGB-IR image sensor with a structured light projector. In an example, system 100 can include block (or circuit) 102, block (or circuit) 104, block (or circuit) 106, and / or block (or circuit) 108. Circuit 102 can be implemented as a control circuit (e.g., a dedicated circuit, an embedded controller, a processor, a microprocessor, etc.). Circuit 104 can implement an infrared structured light projector. Circuit 106 can implement a security / surveillance camera (or module). Circuit 108 can implement an image signal processing (ISP) circuit (or processor or front end). In an example, circuit 108 may be capable of performing multi-channel ISP.

[0033] In an example, circuit 102 may include block (or circuit) 110. Block 110 may implement structured light (SL) control circuitry (or functionality). In another example, circuits 102 and 110 may be implemented as separate circuit cores that may be instantiated on a single integrated circuit substrate (or die) or in a multi-chip module (MCM). In an example, circuits 102 and 108 (and circuit 110 when separate from circuit 102) may be implemented in a single integrated circuit or system on chip (SOC) 112.

[0034] In various embodiments, circuitry 102 may be configured to implement a disparity angle map generation (DAMG) technique 114 according to an example embodiment of the present invention. In an example, the disparity angle map generation (DAMG) technique 114 may be implemented using hardware, software, or a combination of hardware and software. In embodiments where the disparity angle map generation (DAMG) technique 114 is implemented in software or as a combination of hardware and software, computer-executable instructions implementing the disparity angle map generation (DAMG) technique 114 may be stored on circuitry 102 or in a memory associated with circuitry 102. In various embodiments, circuitry 102 may also be configured to implement an oriented block matching (OMB) technique 116 according to an example embodiment of the present invention. In an example, the oriented block matching (OMB) technique 116 may be implemented using hardware, software, or a combination of hardware and software. In embodiments where the oriented block matching (OMB) technique 116 is implemented in software or as a combination of hardware and software, computer-executable instructions implementing the oriented block matching (OMB) technique 116 may be stored on circuitry 102 or in a memory associated with circuitry 102.

[0035] In various embodiments, the circuit 102 can be connected to the IR structured light projector 104, the camera 106, and the ISP circuit 108. The camera 106 can also be connected to the ISP circuit 108. In an example, the circuit 102 generally provides a central control mechanism to synchronize the timing of the IR projector 104 and the camera 106. In an example, the circuit 102 can be configured to control the structured light source 120 of the IR projector 104. In various embodiments, the circuit 102 can also be configured to perform a structured light projector installation calibration routine (algorithm) using images captured by the camera 106. In one example, the installation calibration can be performed offline using images of a wall, one image at a close (e.g., 20 cm) distance and the other image at a far (e.g., 50 cm) distance. In another example, the installation calibration can be performed in real time (e.g., online after deployment) using images of objects, one object at a close distance and the second object at a far distance. In another example, the installation calibration can be performed periodically to maintain and / or update the performance of the system 100.

[0036] In an example, the circuit 102 may also be configured to control an infrared (IR) or RGB-IR sensor 130 of the camera 106. In an example, the circuit 102 may also be configured to control the ISP circuit 108 to synchronize with the output of the camera 106. In various embodiments, the circuit 102 may be configured to generate one or more video output signals (e.g., VIDOUT) and a signal (e.g., DISPARITY MAP). In an example, the signal DISPARITY MAP may be used to determine depth information.

[0037] In various embodiments, the circuit 102 may be configured to perform a calibration process according to an embodiment of the present invention. In an example, the circuit 102 may be configured to generate a disparity angle map that includes disparity angle information for each pixel of an image captured by the camera 106. In various embodiments, the circuit 102 may be configured to execute an oriented block matching (OBM) algorithm that may cooperate with the calibration process. In an example, the OBM algorithm may use the disparity angle map generated during the calibration process to construct a disparity map by pattern shift matching. In some embodiments, the circuit 102 may also be configured to output the disparity angle map generated during the calibration.

[0038] In various embodiments, the video output signal VIDOUT generated by the processor 102 can encode various video streams for different purposes. In one example, IR channel data without structured light pattern contamination and without additional IR illumination can be used for face detection, face recognition and / or real-time video feed in day mode. In another example, IR channel data without structured light pattern contamination and with additional IR illumination can be used for face detection, face recognition and / or real-time video feed in night mode. In yet another example, IR channel data with structured light pattern and without additional IR illumination can be used for depth analysis and presence determination in both day mode and night mode. In an embodiment utilizing an RGB-IR sensor, RGB channel data without structured light pattern contamination can be used for face detection and face recognition and / or real-time video feed in day mode. In an example, RGB channel data with structured light pattern contamination can be discarded.

[0039] In some embodiments, circuit 106 can be configured to present a signal (e.g., ES). Signal ES can indicate when sensor 130 begins exposure (or provide information to facilitate calculation using a predefined formula for when sensor 130 begins exposure). In one example, a flash pin of rolling shutter sensor 130 can be configured to generate signal ES. In another example, other sensor signals from circuit 106 can be used to calculate when to begin exposure (e.g., using a predefined formula, etc.). Circuit 110 can utilize signal ES from circuit 106 to control circuit 104. In another example, signal ES can be configured to generate an interrupt in response to sensor 130 beginning exposure. The interrupt can cause circuit 110 to begin a predefined on-period of time for structured light source 120 of IR projector 104. In an example, circuit 110 can be configured to program a timer with a predefined on-period. In response to receiving signal ES, circuit 110 can start a timer to turn on the structured light source for a predefined time period.

[0040] In an example, circuit 102 may have an input that may receive signal ES, a first input / output that may communicate with a first input / output of circuit 108 via a signal (e.g., ISP_SYNC), a second input / output that may communicate an infrared image channel (e.g., IR DATA) with a second input / output of circuit 108, an optional third input / output that may communicate a color image channel (e.g., RGB DATA) with a third input / output of circuit 108, a first output that may present a signal (e.g., SL_TRIG), a second output that may present one or more video output signals VIDOUT, and a third output that may present a DISPARITY MAP signal. In an example, circuit 104 may have an input that may receive signal SL_TRIG. Circuit 104 may be configured to generate a structured light pattern based on signal SL_TRIG.

[0041] In an example, circuit 106 can have an output that can present signal ES (or another signal communicating information that can be used to calculate the start of exposure), and an input / output that can transmit a signal (e.g., RAWVIDEO) to a fourth input / output of circuit 108. In an example, signal RAW VIDEO can transmit a single channel of video pixel information (e.g., IR) to circuit 108. In another example, signal RAW VIDEO can transmit four channels of video pixel information (e.g., R, G, B, and IR) to circuit 108. In an example, circuits 106 and 108 can also exchange control and / or status signals via the connection carrying signal RAW VIDEO.

[0042] In an example, circuit 108 can be configured to divide the four-channel RGB-IR video signal RAW VIDEO received from circuit 106 into separate IR image data channels and RGB image data channels. In embodiments utilizing an RGB-IR sensor, circuit 108 can be configured to generate a first image channel having a signal IR DATA and a second image channel having a signal RGB DATA in response to signal RAW VIDEO. The first image channel IR DATA typically includes monochrome image data. The second image channel RGB DATA typically includes color image data. In an example, the color image data can include RGB color space data or YUV color space data. When a structured light pattern is projected by circuit 104, the first image channel IR DATA typically includes both the IR image data and the structured light pattern projected by circuit 104. When a structured light pattern is not projected by circuit 104, the first image channel IR DATA typically includes the IR image data without the structured light pattern. When a structured light pattern is projected by circuit 104, the second image channel RGB DATA typically also includes the structured light pattern projected by circuit 104 and is therefore typically ignored.

[0043] In an example, the structured light pattern data carried by the first image channel IR DATA can be analyzed by the circuit 102 to obtain 3D (e.g., depth) information for the field of view of the camera 106. The circuit 102 can also be configured to perform presence determination based on the structured light pattern data carried by the first image channel IR DATA. In an example, the circuit 102 can analyze the RGB (or YUV) data and the IR data to distinguish (e.g., detect, identify, etc.) one or more features or objects in the field of view of the camera 106. In an example, the circuit 110 can be configured to generate the signal SL_TRIG. The circuit 110 can implement a structured light control timing protocol. In an example, the circuit 110 can be implemented in hardware, software (or firmware, microcode, etc.), or a combination of hardware and software.

[0044] In an example, circuit 120 can be implemented as a structured light source. In an example, circuit 120 can be implemented as an array of vertical cavity surface emitting lasers (VCSELs) and a lens. However, other types of structured light sources can be implemented to meet the design criteria of a specific application. In an example, the array of VCSELs is generally configured to generate a laser pattern. The lens is generally configured to decompose the laser pattern into a dense array of dot (or spot) patterns. In an example, circuit 120 can implement a near-infrared (NIR) light source. In various embodiments, the light source of circuit 120 can be configured to emit light with a wavelength of approximately 940 nanometers (nm), which is invisible to the human eye. However, other wavelengths can be utilized. In an example, a wavelength in the range of approximately 800nm-1000nm can be utilized. In an example, circuit 120 can be configured to emit a structured light pattern in response to signal SL_TRIG. In an example, the time period and / or intensity of the light emitted by circuit 120 can be controlled (e.g., programmed) by circuit 102. In an example, circuit 102 may configure circuit 120 before asserting signal SL_TRIG.

[0045] In an example, circuit 130 can be implemented using a global shutter image sensor or a rolling shutter image sensor. When circuit 130 is implemented using a global shutter image sensor, all pixels of the sensor can begin exposure at the same time. When circuit 130 is implemented using a rolling shutter image sensor, as each row (or line) of the sensor begins exposure, all pixels in that row (or line) can begin exposure at the same time. In an example, circuit 130 can be implemented using an infrared (IR) sensor or an RGB-IR image sensor. In an example, the RGB-IR image sensor can be implemented as an RGB-IR complementary metal oxide semiconductor (CMOS) image sensor.

[0046] In one example, circuit 130 can be configured to assert signal ES in response to starting exposure. In another example, circuit 130 can be configured to assert another signal that can be used to calculate the start of exposure of the sensor using a predefined formula. In an example, circuit 130 can be configured to generate signal RAW VIDEO. In an example, circuit 130 can apply a mask to a monochrome sensor. In an example, the mask can include multiple cells that contain a red pixel, a green pixel, a blue pixel, and an IR pixel. The IR pixel can contain red, green, and blue filter materials that can effectively absorb all light in the visible spectrum while allowing longer infrared wavelengths to pass with minimal loss. Due to hardware limitations, red, green, and blue pixels can also receive (respond to) longer infrared wavelengths. Therefore, the infrared structured light pattern (if present) will contaminate the RGB channels. Due to structured light contamination, the RGB frame can generally be ignored when the infrared structured light pattern is present.

[0047] refer to Figure 2 , a diagram illustrating elements of a monocular camera stereo system according to an example embodiment of the present invention is shown. In an example, the monocular camera stereo system 100 may include a housing 140 and a processor 102. An infrared (IR) structured light projector 104 (including a first lens and a structured light source 120) and a camera 106 (including an IR image sensor or an RGB-IR image sensor of a second lens and circuit 130) may be mounted in the housing 104. In one example, the processor 102 may also be mounted within the housing 140. In another example, the processor 102 may be located externally or remotely from the housing 140. The IR structured light projector 104 may be configured to project a structured light pattern (SLP) onto an object in the field of view of the camera 106 when turned on. The IR image sensor of the circuit 130 may be used to acquire IR image data (with and without the structured light pattern) for objects in the field of view of the camera 106. The RGB-IR image sensor of circuitry 130 (if implemented) can be used to acquire both IR image data (with and without structured light patterns) and RGB image data (without structured light patterns) for objects in the field of view of camera 106. Monocular camera stereo system 100 generally provides advantages over conventional dual-camera 3D sensing systems. By utilizing the RGB-IR image sensor to acquire both RGB image data and IR image data with and without structured light patterns (SLP), monocular camera stereo system 100 can reduce system cost and system complexity relative to conventional systems (e.g., one sensor and one lens versus two sensors and two lenses).

[0048] In an example, processor 102 may utilize RGB-IR data from an RGB-IR sensor of circuit 130, the RGB-IR data having been separated (divided into) a first image data channel and a second image data channel, the first image data channel including IR image data having a structured light pattern present, the second image data channel including RGB image data and / or IR image data without a structured light pattern present. In an example, the first image data channel and the second image data channel may be processed by processor 102 for 3D (e.g., depth) perception, presence determination, 3D facial recognition, object detection, facial detection, object recognition, and facial recognition. In an example, the first image data channel including IR image data having a structured light pattern present may be used to perform a structured light projector installation calibration process according to an example embodiment of the present invention. The first image data channel including IR image data having a structured light pattern present may also be used to perform depth analysis, presence determination, and / or 3D facial recognition. In an example, processor 102 may be configured to generate a disparity map output based on a disparity angle map generated during projector installation calibration. A second image data channel comprising IR image data without the present structured light pattern and RGB image data without the present structured light pattern can be used to generate an encoded (or compressed) video signal, bitstream, or multiple bitstreams, and to perform object detection, facial detection, object recognition, and / or facial recognition.

[0049] refer to Figure 3 , shows a schematic diagram illustrating an example arrangement of elements of a monocular camera stereo system. In various embodiments, a plane 150 of the camera sensor 130 is vertically separated from a plane 152 of the lens 154 of the camera 106. In an ideal mounting position, the center of the structured light projector 104 is constrained to be aligned horizontally (e.g., on the X-axis) and vertically (e.g., on the Z-axis) with the optical center of the lens 154 of the camera 106. Due to this limitation, the use of a monocular camera stereo system may be limited for the following reasons. In production, to ensure ideal alignment, high mounting accuracy of the structured light projector 104 is required, which may increase production costs and reduce efficiency. In design, high mounting accuracy standards for the structured light projector 104 may limit component distribution and cause size issues. In various embodiments, a solution is generally provided that overcomes mounting limitations when mounting the structured light projector 104 in a monocular camera stereo system and allows for flexibility (e.g., lower accuracy) when mounting the structured light projector 104 in a monocular camera stereo system.

[0050] refer to Figure 4 , shows the description by Figure 3FIGURE 1 illustrates a structured light (or speckle) pattern projected onto an object by a structured light projector 104. In an example, the structured light projector 104 generates a point light source that spreads light, thereby forming a dot or speckle pattern on an object (e.g., a wall, etc.). When the object 160 moves toward or away from the camera 106, the speckle pattern on the object 160 generally shifts horizontally. During a calibration process according to an embodiment of the present invention, a reference image of the speckle pattern on the object 160 may be captured at a first (e.g., close) distance (R), and a target image of the speckle pattern on the object 160 may be captured at a second (e.g., far) distance (X). In an example, the distance of the reference image may be 20 centimeters (cm), and the distance of the target image may be 50 cm. Other close and / or far distances may be used accordingly. The movement distance (M) may be calculated by measuring the displacement of the speckle pattern captured by the camera sensor 130 (e.g., M=XR). The angle of the displacement of the speckle pattern between the two captured images may also be determined. The angle of the displacement offset of the speckle pattern may be referred to as the parallax angle.

[0051] refer to Figure 5 , a graph 170 illustrating the parallax of a monocular camera stereo system 100 with a well (accurately) mounted structured light projector 104 is shown. In this example, the rectangular plane 150 of the camera sensor 130 can be parallel to the XY plane 152 containing the lens 154 of the camera 106 and the structured light projector 104, and the pixel rows of the sensor 130 (e.g., represented by the dashed lines) can be aligned with the X-axis direction. Point A can be used to represent the optical center of the lens 154, and point B can be used to represent the mounting point of the structured light projector 104. Points H and G can be used to represent one of the spots on the target wall 172 and the reference wall 174, respectively, at different distances from the camera 106. Spot G can be captured by the camera sensor 130 at point C, and spot H can be captured by the camera sensor 130 at point D. The parallax value is generally defined as the length of the line CD (e.g., when the object moves from H to G, the spot pattern shifts from D to C).

[0052] Plane ABG is generally defined by three points A, B, and G. Plane ABG intersects the plane 150 of the image sensor 130 at the parallax line CD and intersects the XY plane 152 at the X axis. Because plane ABG intersects the plane 150 of the image sensor 130 at the parallax line CD and intersects the XY plane 152 at the X axis, the parallax line CD is always parallel to the X axis. In existing parallax algorithms such as block matching (BM) and semi-global block matching (SGBM), the parallax line CD must remain aligned with the X axis without any angle with the X axis. Therefore, existing parallax algorithms limit the installation of the structured light projector 104 relative to the camera 106.

[0053] refer to Figure 6 , shows a graph 180 illustrating the parallax with the monocular camera stereo system 100 when the mounting of the structured light projector 104 is displaced in the XY plane 152. Given Figure 5 Under the same conditions in FIG, when the mounting point of the structured light projector 104 is shifted from point B on the X-axis to point (B') in the XY plane 152, an angle Theta appears between line AB' and the X-axis. Plane AB'G intersects plane 150 of image sensor 130 at parallax line CD and intersects XY plane 152 at line AB'. Since parallax line CD is always parallel to line AB', parallax line CD also makes an angle Theta with the pixel rows of image sensor 130. Figure 6 In the example shown in FIG. 1 where the structured light projector 104 is displaced in the XY plane 152 , all pixels in the image sensor 130 generally have the same parallax angle Theta.

[0054] refer to Figure 7 , shows a diagram 190 illustrating the parallax with the monocular camera stereo system 100 when the mounting of the structured light projector 104 is displaced along the Z axis. Figure 6 Under the same conditions as in (except that the structured light projector 104 is shifted up / down along the Z axis from point B to point B'), plane AB'G intersects plane 150 of image sensor 130 at parallax line CD and intersects XY plane 152 at line AF. An angle Theta appears between line AF and the X axis. Because parallax line CD is always parallel to line AF, parallax line CD also makes an angle Theta with the pixel rows of image sensor 130. Figure 7 In the example where the structured light projector 104 is shifted along the Z-axis, different pixels may have different parallax angles.

[0055] In summary, the mounting point of the structured light projector 104 can be Figure 6 , and can also be shifted as Figure 7 Shift as in . Figure 5 Compared to the example in , the parallax angle is usually as Figure 6 and Figure 7 When the displacement of the mounting point of the structured light projector 104 is Figure 6 and Figure 7 When the object 160 moves from point H to point G, the speckle pattern is generally shifted from point D to point C at an oblique angle. For existing stereo algorithms such as block matching (BM) and semi-global block matching (SGBM), the disparity angle should be strictly zero. When the disparity angle is not zero, existing stereo algorithms do not work properly because they can only perform pattern matching by horizontal shifting in one row.

[0056] In an algorithm according to an exemplary embodiment of the present invention, a disparity angle (e.g., a disparity angle map) may be assigned to each pixel. When a disparity angle is assigned to each pixel, a new block matching scheme (e.g., a directional BM or SGBM algorithm) may be defined that guides pattern matching by shifting at the assigned disparity angle for each pixel. The calibration technique according to an exemplary embodiment of the present invention is generally configured to help establish a disparity angle map for each of the pixels in the image captured by the image sensor 130. The new block matching scheme can utilize the disparity angle map to generate a disparity map that is not dependent on ideal installation constraints.

[0057] Reference Figure 8 , a diagram 200 illustrating a block matching operation utilizing a local search range, which may be part of a calibration process (or procedure), according to an example embodiment of the present invention. In various embodiments, multiple calibration steps may be performed to establish a disparity angle map for the pixels of the captured image. In examples where the structured light projector 104 is flexibly (or less than ideally) mounted on a monocular camera stereo apparatus, the structured light projector 104 may be misaligned with the optical center of the camera lens 154 (e.g., as Figure 6 and 7 , with shift offsets on the X / Y / Z axes).

[0058] In a first step of the calibration process, the camera may be placed in front of a white wall, the structured light projector may be powered on, and two speckle pattern pictures may be captured, a first (e.g., reference) picture 202 captured at a close distance (e.g., 20 cm) and a second (e.g., target) picture captured at a far distance (e.g., 50 cm). In the next step of the calibration process, a grid of pixels (e.g., at the center of the reference picture 202) may be selected. Figure 8 The calibration process can be continued by measuring the disparity angle for each grid pixel using the following sub-steps.

[0059] In a first sub-step, a local search range 206 may be estimated for the grid pixels in the target image 204 (e.g., by Figure 8 ). In an example, a rectangular search range may be suggested. In a next sub-step, the pattern block may be cropped around the grid pixels in the reference picture 202. In an example, the pattern block may be cropped around the grid pixel at (x0, y0) in the reference picture 202. In a next sub-step, the pattern block may be shifted over the pixels in the local search range 206 in the target picture 204, and a block matching operation may be performed to match the pattern block with the pixels in the local search range 206 to find the pixel with the best pattern match (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) (e.g., the pixel with the best pattern match) Figure 8 , y1) in the target image 204). In the next sub-step, a line connecting the point at coordinates (x0, y0) and the point at coordinates (x1, y1) may be determined. The line connecting the two points is a disparity line (e.g., a line drawn by Figure 8 ). In an example, the tangent of the disparity angle may be determined as (y1-y0) / (x1-x0). In an example, the disparity angle map may be generated as a two-dimensional (2D) vector by filling in the tangent value calculated for each grid pixel. This process is typically repeated for each grid pixel in the reference picture 202.

[0060] At this point, there may still be some problems to be solved: (i) only the disparity angles of the grid pixels are obtained, not the complete map of the pixels; (ii) even if the matching criterion is a best match, the disparity angles of some of the grid pixels may not have been obtained due to very low matching confidence; (iii) due to some mismatches, incorrect measurements of the disparity angles may have been obtained. In the next step of the calibration process, the incorrect disparity angles can be corrected and a complete (dense) disparity angle map can be generated. In the first sub-step, the matrix of the disparity angle map can be regarded as a parametric surface. In the example, the surface formula of the parametric surface can be expressed using the following equation 1:

[0061] z=ax^3+bx^2y+cxy^2+dy^3+ex^2+fxy+gy^2+hx+my+n Equation 1

[0062] where z is the tangent of the parallax angle and (x, y) are the pixel coordinates in the sensor plane 150. When the ten coefficients a, b, c, d, e, f, g, h, m and n are determined, the surface is typically determined (fixed) for the map.

[0063] Based on the measurements of the grid pixels, a regression method can be used to easily obtain the best fit surface. In an example, a least squares regression (LSR) method can be used. In an example with N measurement pairs (xi, yi)->zi (where (xi, yi) are the axes of the grid pixels and zi is the tangent value of the grid pixels), the coefficients a, b, c, d, ..., m and n can be calculated by solving the following formula:

[0064]

[0065] When the coefficients have been calculated, the surface formula is fixed. The solved formula can be used to predict the updated disparity angle map for all grid pixels. The updated disparity angle map for all grid pixels typically includes corrections for any erroneous disparity lines and any missing disparity lines. In the second sub-step, the disparity angle for non-grid pixels can be predicted using the following formula:

[0066]

[0067] When the disparity angles for non-grid pixels have been predicted, a complete map of disparity angles has been obtained.

[0068] After the calibration process has been performed, the map of disparity angles is complete and can be used. In an example, a typical block matching (BM) algorithm can be modified according to embodiments of the present invention to obtain an oriented block matching (OBM) process that can be coordinated with the calibration process according to embodiments of the present invention. The OBM process can be similar to the typical BM algorithm, except that the typical BM process constructs disparity by pattern shift matching in horizontal lines, while the OBM process uses the disparity angles in the map of disparity angles generated during the calibration process to establish disparity by pattern shift matching.

[0069] In an example, the calibration process may utilize the cubic parameter surface described above by Equation 2:

[0070] z=ax^3+bx^2y+cxy^2+dy^3+ex^2+fxy+gy^2+hx+my+n Equation 2

[0071] However, the calibration process is not limited to the use of cubic parameter surfaces. Parameter surfaces of other orders can be used to meet the design criteria of a specific implementation. Additionally, the regression of the parameter surface is not limited to the LSR regression method. Other regression algorithms can be implemented accordingly to meet the design criteria of a specific implementation. In various embodiments, the calibration process described above can be used to construct a disparity angle map and applied to a modified block matching algorithm (e.g., the directional block matching (OBM) algorithm described above). However, the disparity angle map according to an embodiment of the present invention can be applied to modify other stereo algorithms (e.g., SGBM, etc.). In the example, an offline calibration process using the capture of two pictures of a wall has been described, where one picture is captured at a close distance and the other picture is captured at a long distance. However, the calibration process according to an embodiment of the present invention can also be used to perform online calibration by capturing two pictures of an object, where one picture is captured when the object is at a close distance and the other picture is captured when the object is at a long distance.

[0072] refer to Figure 9, a flow chart of process 300 is shown. Process (or method) 300 may implement an installation calibration process according to another example embodiment of the present invention. Process 300 generally includes step (or state) 302, step (or state) 304, step (or state) 306, step (or state) 308, step (or state) 310, step (or state) 312, step (or state) 314, decision step (or state) 316, step (or state) 318, and step (or state) 320. Process 300 generally begins in step 302 and moves to step 304. In step 304, camera 106 may be placed in front of a white wall. In step 306, structured light projector 104 may be powered on. In step 308, a first (e.g., reference) picture 202 of the speckle pattern projected onto the wall by structured light projector 104 is captured at a close distance (e.g., 20 cm). In step 310, a second (e.g., target) picture 204 of a speckle pattern projected onto a wall is captured at a long distance (e.g., 50 cm). In step 312, a grid pixel in the reference picture 202 can be selected. In step 314, the disparity angle for the selected grid pixel can be measured. In step 316, the process 300 checks whether the disparity angle has been measured for all grid pixels. If the disparity angle has not been measured for all grid pixels, the process 300 returns to step 314 to process the next grid pixel. When the disparity angle has been measured for all grid pixels, the process 300 moves to step 318.

[0073] At this point, there may still be some issues to address: only the disparity angles for the grid pixels are obtained, not a complete map of the pixels; even if the matching criterion is a best match, the disparity angles for some of the grid pixels may not be obtained due to very low match confidence; and due to some mismatches, erroneous measurements of the disparity angles may be obtained. In step 318, process 300 corrects the erroneous disparity angles and generates a complete (dense) disparity angle map for the non-grid pixels. When the complete (dense) disparity angle map has been generated, process 300 moves to step 320 and terminates.

[0074] refer to Figure 10, a flow chart of process 400 is shown. Process (or method) 400 generally illustrates a process for measuring disparity angles for grid pixels according to another example embodiment of the present invention. Process 400 generally includes step (or state) 402, step (or state) 404, step (or state) 406, step (or state) 408, step (or state) 410, step (or state) 412, step (or state) 414, step (or state) 416, decision step (or state) 418, step (or state) 420, and step (or state) 422. Process 400 may begin in step 402 and move to step 404.

[0075] In step 404, a local search range 206 may be estimated for the grid pixels in the target image 204 (e.g., by Figure 8 ). In an example, a rectangular search range may be suggested. In step 406, process 400 may begin determining a disparity angle for each of the grid pixels in target image 204. In step 408, process 400 may crop a pattern block around the grid pixels in reference image 202. In an example, the pattern block may be cropped around the grid pixel located at (x0, y0) in reference image 202. In step 410, process 400 may shift the pattern block across pixels in the local search range 206 in target image 204 and perform a block matching operation to match the pattern block with pixels in the local search range 206 to find a block 208 with the best pattern match (e.g., located at (x1, y1) in target image 204). In step 412, process 400 may determine a line connecting a point located at coordinates (x0, y0) and a point located at coordinates (x1, y1). The line connecting the two points is a disparity line. In step 414, process 400 may determine the tangent of the disparity angle as (y1-y0) / (x1-x0). In step 416, process 400 may populate the tangent value calculated for the grid pixel into a two-dimensional (2D) vector representing the disparity angle map for the grid pixel. In step 418, process 400 checks to determine whether the tangent value has been calculated for all grid pixels in the reference picture. When the tangent value has not been calculated for all grid pixels in the reference picture, process 400 moves to step 420 to select the next grid pixel and then returns to step 408. When the tangent value has been calculated for all grid pixels in the reference picture, process 400 moves to step 422 and terminates.

[0076] refer to Figure 11, a flow chart of process 500 is shown. Process (or method) 500 generally illustrates a process for generating a disparity angle map for all pixels in a target image according to an example embodiment of the present invention. Process 500 generally includes step (or state) 502, step (or state) 504, step (or state) 506, step (or state) 508, and step (or state) 510. Process 500 may start at step 502 and move to step 504. In step 504, process 500 may determine coefficients of a parametric surface formula of a disparity angle map matrix. In an example, the disparity angle map matrix may be considered a parametric surface. In an example, the surface formula of the parametric surface may be represented using the following equation 1:

[0077] z=ax^3+bx^2y+cxy^2+dy^3+ex^2+fxy+gy^2+hx+my+n Equation 1

[0078] where z is the tangent of the parallax angle and (x, y) are the pixel coordinates in the sensor plane 150. When the ten coefficients a, b, c, d, e, f, g, h, m, and n are determined, the surface is typically determined (fixed) for the parallax angle map.

[0079] In step 506, process 500 can determine the best-fit surface formula using a regression method. Based on the measurements of the grid pixels, a regression method such as least squares regression (LSR) can be used to easily obtain the best-fit surface. In an example with N measurement pairs (xi, yi)->zi (where (xi, yi) are the axes of the grid pixels and zi is the tangent value of the grid pixels), the coefficients a, b, c, d, ..., m, and n can be calculated by solving the following formula:

[0080]

[0081] When the coefficients have been calculated, the surface formula is fixed. The solved formula can be used to predict the updated disparity angle map for all grid pixels. The updated disparity angle map for all grid pixels includes corrections for any erroneous disparity lines and any missing disparity lines. When the updated disparity angle map for all grid pixels has been predicted, process 500 can move to step 508. In step 508, process 500 can use the best fit surface formula determined in step 506 to predict the updated disparity angle map for all pixels of the image. In an example, the disparity angle of non-grid pixels can be predicted by the following formula:

[0082]

[0083] When the disparity angles for the non-grid pixels have been predicted, a complete map of disparity angles has been obtained. Process 500 may then move to step 510 and terminate.

[0084] refer to Figure 12 , a graph 600 is shown illustrating example results between performing a typical block matching routine and performing a directional block matching routine according to an example embodiment of the present invention. In the example, graph 602 generally illustrates the results of running a typical block matching algorithm without calibrating for disparity angle when the structured light projector 104 is not accurately mounted in the monocular camera stereo system. Without calibrating for disparity angle, the results of the typical block matching algorithm are very noisy. Graph 604 generally illustrates the improved results of running a directional block matching algorithm according to an example embodiment of the present invention in conjunction with calibrating for disparity angle according to an example embodiment of the present invention when the structured light projector 104 is not accurately mounted in the monocular camera stereo system.

[0085] refer to Figure 13 , a block diagram illustrating an example implementation of a monocular camera stereo device 800 is shown. In an example, the monocular camera stereo device 800 may include blocks (or circuits) 802, 804, 806, 808, 810, 812, 814, 816, 818, and / or 820. Circuit 802 may be implemented as a processor and / or a system on a chip (SoC). Circuit 804 may be implemented as a capture device. Circuit 806 may be implemented as a memory. Block 808 may be implemented as an optical lens. Circuit 810 may be implemented as a structured light projector. Block 812 may be implemented as a structured light pattern lens. Circuit 814 may be implemented as one or more sensors (e.g., motion, ambient light, proximity, sound, etc.). Circuit 816 may be implemented as a communication device. Circuit 818 may be implemented as a wireless interface. The circuit 820 may be implemented as a battery 820 .

[0086] In some embodiments, the monocular camera stereo device 800 may include a processor / SoC 802, a capture device 804, a memory 806, a lens 808, an IR structured light projector 810, a lens 812, a sensor 814, a communication module 816, a wireless interface 818, and a battery 820. In another example, the monocular camera stereo device 800 may include the capture device 804, the lens 808, the IR structured light projector 810, the lens 812, and the sensor 814, and the processor / SoC 802, the memory 806, the communication module 816, the wireless interface 818, and the battery 820 may be components of a single device. The implementation of the monocular camera stereo device 800 may vary depending on the design criteria of a particular implementation.

[0087] Lens 808 can be attached to capture device 804. In an example, capture device 804 can include block (or circuit) 822, block (or circuit) 824, and block (or circuit) 826. Circuit 822 can implement an image sensor. In an example, the image sensor of circuit 822 can be an IR image sensor or an RGB-IR image sensor. Circuit 824 can be a processor and / or logic. Circuit 826 can be a memory circuit (e.g., a frame buffer).

[0088] The capture device 804 can be configured to capture video image data (e.g., light collected and focused by the lens 808). The capture device 804 can capture data received through the lens 808 to generate a video bitstream (e.g., a sequence of video frames). In various embodiments, the lens 808 can be implemented as a fixed-focus lens. Fixed-focus lenses generally promote smaller size and low power. In an example, a fixed-focus lens can be used for battery-powered doorbells and other low-power camera applications. In some embodiments, the lens 808 can be oriented, tilted, translated, zoomed, and / or rotated to capture the environment surrounding the monocular camera stereo device 800 (e.g., capture data from the field of view). In an example, an active lens system can be utilized to implement a specialized camera model for enhanced functionality, remote control, and the like.

[0089] The capture device 804 can convert the received light into a digital data stream. In some embodiments, the capture device 804 can perform analog-to-digital conversion. For example, the image sensor 822 can perform photoelectric conversion on the light received by the lens 808. The processor / logic 824 can convert the digital data stream into a video data stream (or bit stream), a video file, and / or multiple video frames. In an example, the capture device 804 can present the video data as a digital video signal (e.g., RAW VIDEO). The digital video signal can include video frames (e.g., sequential digital images and / or audio).

[0090] The video data captured by the capture device 804 can be represented as a signal / bitstream / data conveyed by a digital video signal RAW VIDEO. The capture device 804 can present the signal RAW VIDEO to the processor / SoC 802. The signal RAW VIDEO can represent video frames / video data. The signal RAW VIDEO can be a video stream captured by the capture device 804.

[0091] The image sensor 822 can receive light from the lens 808 and convert the light into digital data (e.g., a bit stream). For example, the image sensor 822 can perform a photoelectric conversion on the light from the lens 808. In some embodiments, the image sensor 822 can have additional margin that is not used as part of the image output. In some embodiments, the image sensor 822 may not have additional margin. In various embodiments, the image sensor 822 can be configured to generate an RGB-IR video signal. In a field of view illuminated only by infrared light, the image sensor 822 can generate a monochrome (B / W) video signal. In a field of view illuminated by both IR light and visible light, the image sensor 822 can be configured to generate color information in addition to the monochrome video signal. In various embodiments, the image sensor 822 can be configured to generate a video signal in response to visible light and / or infrared (IR) light.

[0092] The processor / logic 824 may convert the bitstream into human-viewable content (e.g., video data (regardless of image quality) that an average person can understand, such as video frames). For example, the processor / logic 824 may receive pure (e.g., raw) data from the image sensor 822 and generate (e.g., encode) video data (e.g., a bitstream) based on the raw data. The capture device 804 may have a memory 826 to store the raw data and / or the processed bitstream. For example, the capture device 804 may implement a frame memory and / or buffer 826 to store (e.g., provide temporary storage and / or caching) one or more of the video frames (e.g., digital video signals). In some embodiments, the processor / logic 824 may perform analysis and / or correction on the video frames stored in the memory / buffer 826 of the capture device 804.

[0093] The sensor 814 can implement multiple sensors, including but not limited to a motion sensor, an ambient light sensor, a proximity sensor (e.g., ultrasound, radar, lidar, etc.), an audio sensor (e.g., a microphone), etc. In an embodiment implementing a motion sensor, the sensor 814 can be configured to detect motion anywhere in the field of view monitored by the monocular camera stereo device 800. In various embodiments, the detection of motion can be used as a threshold for activating the capture device 804. The sensor 814 can be implemented as an internal component of the monocular camera stereo device 800 and / or as a component external to the monocular camera stereo device 800. In an example, the sensor 814 can be implemented as a passive infrared (PIR) sensor. In another example, the sensor 814 can be implemented as an intelligent motion sensor. In an embodiment implementing an intelligent motion sensor, the sensor 814 can include a low-resolution image sensor configured to detect motion and / or people.

[0094] In various embodiments, the sensor 814 may generate a signal (e.g., SENS). The SENS signal may include various data (or information) collected by the sensor 814. In an example, the SENS signal may include data collected in response to motion being detected in the monitored field of view, the ambient light level in the monitored field of view, and / or sounds picked up in the monitored field of view. However, other types of data may be collected and / or generated based on the design criteria of a particular application. The SENS signal may be presented to the processor / SoC 802. In an example, when motion is detected in the field of view monitored by the monocular camera stereo device 800, the sensor 814 may generate (assert) the SENS signal. In another example, when the sensor 814 is triggered by audio in the field of view monitored by the monocular camera stereo device 800, the sensor 814 may generate (assert) the SENS signal. In yet another example, the sensor 814 may be configured to provide directional information regarding motion and / or sound detected in the field of view. The directional information may also be transmitted to the processor / SoC 802 via the SENS signal.

[0095] The processor / SoC 802 may be configured to execute computer-readable code and / or process information. In various embodiments, the computer-readable code may be stored within the processor / SoC 802 (e.g., microcode, etc.) and / or in the memory 806. In an example, the processor / SoC 802 may be configured to execute one or more artificial neural network models (e.g., facial recognition CNN, object detection CNN, object classification CNN, etc.) stored in the memory 806. In an example, the memory 806 may store one or more directed acyclic graphs (DAGs) and one or more weight sets defining the one or more artificial neural network models. The processor / SoC 802 may be configured to receive inputs from the memory 806 and / or present outputs to the memory 806. The processor / SoC 802 may be configured to present and / or receive other signals (not shown). The number and / or type of inputs and / or outputs of the processor / SoC 802 may vary according to the design criteria of a particular implementation. The processor / SoC 802 may be configured for low-power (e.g., battery) operation.

[0096] The processor / SoC 802 may receive a signal RAW VIDEO and a signal SENS. In an example, the processor / SoC 802 may generate one or more video output signals (e.g., IR, RGB, etc.) and one or more data signals (e.g., DISPARITY MAP) based on the signal RAW VIDEO, the signal SENS, and / or other inputs. In some embodiments, the signals IR, RGB, and DISPARITY MAP may be generated based on analysis of the signal RAW VIDEO and / or objects detected in the signal RAW VIDEO. In an example, the signal RGB may typically include a color image (frame) in an RGB or YUV color space. In an example, the signal RGB may be generated when the processor / SoC 802 operates in daylight mode. In an example, the signal RGB may be generated when the processor / SoC 802 operates in daylight mode. In an example, the signal IR may typically include a monochrome IR image (frame). In one example, when the processor / SoC 802 operates in daylight mode, the signal IR may include an uncontaminated IR image (e.g., without a structured light pattern) using ambient IR light. In another example, when the processor / SoC 802 operates in night mode, the signal IR may include an uncontaminated IR image (e.g., without a structured light pattern) illuminated by an IR LED. In yet another example, when the IR projector is turned on and the processor / SoC 802 operates in day mode or night mode, the signal IR may include a contaminated IR image (e.g., a structured light pattern is present in at least a portion of the image).

[0097] In various embodiments, the processor / SoC 802 may be configured to perform one or more of feature extraction, object detection, object tracking, and object recognition. For example, the processor / SoC 802 may determine motion information and / or depth information by analyzing a frame from the signal RAW VIDEO and comparing the frame to a previous frame. The comparison may be used to perform digital motion estimation. In some embodiments, the processor / SoC 802 may be configured to generate video output signals RGB and IR comprising video data from the signal RAW VIDEO. The video output signals RGB and IR may be presented to the memory 806, the communication module 816, and / or the wireless interface 818. The signal DISPARITY MAP may be configured to indicate depth information associated with an object present in an image transmitted by the signals RGB and IR. In an example, when a structured light pattern is present, the image data carried by the signal RGB may be ignored (discarded).

[0098] The memory 806 can store data. The memory 806 can implement various types of memory, including but not limited to cache, flash memory, memory cards, random access memory (RAM), dynamic RAM (DRAM) memory, etc. The type and / or size of the memory 806 can vary according to the design standards of a particular implementation. The data stored in the memory 806 can correspond to video files, motion information (e.g., readings from the sensor 814), video fusion parameters, image stabilization parameters, user input, computer vision models, and / or metadata information.

[0099] The lens 808 (e.g., a camera lens) can be oriented to provide a view of the environment surrounding the monocular camera stereo device 800. The lens 808 can be designed to capture environmental data (e.g., light). The lens 808 can be a wide-angle lens and / or a fisheye lens (e.g., a lens capable of capturing a wide field of view). The lens 808 can be configured to capture and / or focus light for the capture device 804. Typically, the image sensor 822 is located behind the lens 808. Based on the light captured from the lens 808, the capture device 804 can generate a bitstream and / or video data.

[0100] The communication module 816 may be configured to implement one or more communication protocols. For example, the communication module 816 and the wireless interface 818 may be configured to implement one or more of the following: IEEE 802.11, IEEE 802.15, IEEE 802.15.1, IEEE 802.15.2, IEEE 802.15.3, IEEE 802.15.4, IEEE 802.15.5, IEEE 802.20, and / or In some embodiments, the wireless interface 818 may also implement one or more protocols associated with a cellular communication network (e.g., GSM, CDMA, GPRS, UMTS, CDMA2000, 3GPP LTE, 4G / HSPA / WiMAX, SMS, etc.). In embodiments where the monocular camera stereo device 800 is implemented as a wireless camera, the protocol implemented by the communication module 816 and the wireless interface 818 may be a wireless communication protocol. The type of communication protocol implemented by the communication module 816 may vary depending on the design criteria of a particular implementation.

[0101] The communication module 816 and / or the wireless interface 818 can be configured to generate a broadcast signal as an output from the monocular camera stereo device 800. The broadcast signal can transmit the video data RGB and / or IR and / or the signal DISPARITY MAP to an external device. For example, the broadcast signal can be sent to a cloud storage service (e.g., a storage service that can be expanded on demand). In some embodiments, the communication module 816 may not transmit data until the processor / SoC 802 has performed video analysis to determine that an object is in the field of view of the monocular camera stereo device 800.

[0102] In some embodiments, the communication module 816 can be configured to generate a manual control signal. The manual control signal can be generated in response to a signal received by the communication module 816 from a user. The manual control signal can be configured to activate the processor / SoC 802. The processor / SoC 802 can be activated in response to the manual control signal regardless of the power state of the monocular camera stereo device 800.

[0103] In some embodiments, the monocular camera stereo device 800 may include a battery 820 configured to provide power to the various components of the monocular camera stereo device 800. A multi-step method for activating and / or deactivating the capture device 804 based on the output of the motion sensor 814 and / or any other power-consuming features of the monocular camera stereo device 800 may be implemented to reduce the power consumption of the monocular camera stereo device 800 and extend the operational life of the battery 820. The motion sensor of the sensor 814 may have a very low power drain on the battery 820 (e.g., less than 10 μW). In an example, the motion sensor of the sensor 814 may be configured to remain on (e.g., always active) unless disabled in response to feedback from the processor / SoC 802. The video analysis performed by the processor / SoC 802 may have a significant power drain on the battery 820 (e.g., greater than the motion sensor 814). In an example, the processor / SoC 802 may be in a low-power state (or powered down) until certain motion is detected by the motion sensor of the sensor 814.

[0104] The monocular camera stereo device 800 can be configured to operate using various power states. For example, in a power-down state (e.g., a sleep state, a low-power state), the motion sensor of the sensor 814 and the processor / SoC 802 can be turned on, and other components of the monocular camera stereo device 800 (e.g., the image capture device 804, the memory 806, the communication module 816, etc.) can be turned off. In another example, the monocular camera stereo device 800 can operate in an intermediate state. In this intermediate state, the image capture device 804 can be turned on, and the memory 806 and / or the communication module 816 can be turned off. In yet another example, the monocular camera stereo device 800 can operate in a powered (or high-power) state. In the powered state, the sensor 814, the processor / SoC 802, the capture device 804, the memory 806, and / or the communication module 816 can be turned on. The monocular camera stereo device 800 can consume some power (e.g., a relatively small and / or minimal amount of power) from the battery 820 in the power-down state. In the powered-on state, the monocular camera stereo device 800 may consume more power from the battery 820. The number of power states and / or the number of components of the monocular camera stereo device 800 that are powered on while the monocular camera stereo device 800 operates in each of the power states may vary according to the design criteria of a particular implementation.

[0105] In some embodiments, the monocular camera stereo device 800 may include a keypad, a touchpad (or screen), a doorbell switch, and / or other human interface device (HID) 828. In an example, the sensor 814 may be configured to determine when an object is near the HID 828. In an example where the monocular camera stereo device 800 is implemented as part of an access control application, the capture device 804 may be activated to provide an image for identifying a person attempting access and illumination of the lock area, and / or the access touchpad may be activated.

[0106] In various embodiments, a low-cost 3D sensing platform can be provided. The low-cost 3D sensing platform can facilitate the development of smart access control systems and smart security products (e.g., smart video doorbells and door locks, payment systems, alarm systems, etc.). In various embodiments, the low-cost 3D sensing platform can include a vision system-on-chip (SoC), a structured light projector, and an IR image sensor or an RGB-IR image sensor. In various embodiments, an RGB-IR CMOS image sensor can be utilized to obtain both visible light images and infrared (IR) images for viewing and facial recognition, and infrared (IR) images can also be utilized for depth sensing. In an example, a vision SoC can provide depth processing, anti-spoofing algorithms, 3D facial recognition algorithms, and video encoding on a single chip.

[0107] In various applications, a low-cost 3D sensing platform according to an embodiment of the present invention can significantly reduce system complexity while improving performance, reliability and safety. In an example, a vision SoC according to an embodiment of the present invention can include, but is not limited to, a powerful image signal processor (ISP), native support for RGB-IR color filter arrays, and advanced high dynamic range (HDR) processing, which can produce extraordinary image quality in low-light and high-contrast environments. In an example, a vision SoC according to an embodiment of the present invention can provide an architecture that delivers computing power for presence detection and 3D facial recognition while running multiple artificial intelligence (AI) algorithms for advanced features such as people counting and anti-rear-end collision.

[0108] In various embodiments, system cost can be reduced by using an RGB-IR sensor (e.g., one sensor and one lens versus two sensors and two lenses). In some embodiments, system cost can be further reduced by using an RGB-IR rolling shutter sensor (e.g., rolling shutter versus global shutter). By controlling the structured light projector via software, the timing sequence can be easily adjusted, providing improved flexibility. Because the structured light projector can be used briefly by software, power savings can be achieved.

[0109] In various embodiments, a low-cost structured light-based 3D sensing system can be implemented. In one example, 3D information can be used for 3D modeling and presence determination. In one example, a low-cost structured light-based 3D sensing system can be used to unlock doors, disarm alarm systems, and / or allow "tripwire" access to restricted areas (e.g., a garden, garage, house, etc.). In one example, a low-cost structured light-based 3D sensing system can be configured to identify gardeners / pool maintenance personnel and prevent the alarm from being triggered. In another example, a low-cost structured light-based 3D sensing system can be configured to limit access to certain times and days of the week. In another example, a low-cost structured light-based 3D sensing system can be configured to trigger an alarm when certain objects are recognized (e.g., if a restraining order is out against an ex-spouse, call 911 if that person is detected). In another example, a low-cost structured light-based 3D sensing system can be configured to allow the alarm system to reprogram privileges based on video / audio recognition (e.g., only allowing person X or Y to change access levels or policies, add users, etc., even if the correct password is entered).

[0110] Various features (e.g., dewarping, digital scaling, cropping, etc.) can be implemented as hardware modules in processor 802. Implementing hardware modules can increase the video processing speed of processor 802 (e.g., faster than software implementations). Hardware implementations can enable video processing while reducing latency. The hardware components used can vary depending on the design criteria of a particular implementation.

[0111] The illustrated processor 802 includes multiple blocks (or circuits) 809a-809n. Blocks 809a-809n can implement various hardware modules implemented by the processor 802. Hardware modules 809a-809n can be configured to provide various hardware components to implement a video processing pipeline. Circuits 809a-809n can be configured to receive pixel data RAW VIDEO, generate video frames based on the pixel data, perform various operations on the video frames (e.g., dewarping, rolling shutter correction, cropping, upscaling, image stabilization, etc.), prepare the video frames for communication with external hardware (e.g., encoding, packaging, color correction, etc.), parse feature sets, implement various computer vision operations, etc. Various implementations of the processor 802 may not necessarily utilize all features of the hardware modules 809a-809n. The features and / or functionality of the hardware modules 809a-809n may vary depending on the design criteria of a particular implementation. Details of the hardware modules 809a-809n and / or other components of the monocular camera stereo device 800 may be described in association with U.S. patent application Ser. No. 15 / 931,942 filed on May 14, 2020, U.S. patent application Ser. No. 16 / 831,549 filed on March 26, 2020, U.S. patent application Ser. No. 16 / 288,922 filed on February 28, 2019, and U.S. patent application Ser. No. 15 / 593,493 filed on May 12, 2017 (now U.S. Patent No. 10,437,600), appropriate portions of which are incorporated herein by reference in their entireties.

[0112] The hardware modules 809a-809n can be implemented as dedicated hardware modules. Compared to software implementations, using dedicated hardware modules 809a-809n to implement the various functions of the processor 802 can enable the processor 802 to be highly optimized and / or customized to limit power consumption, reduce heat generation, and / or increase processing speed. The hardware modules 809a-809n can be customizable and / or programmable to implement a variety of operations. Implementing dedicated hardware modules 809a-809n can enable the hardware used to perform each type of calculation to be optimized for speed and / or efficiency. For example, the hardware modules 809a-809n can implement multiple relatively simple operations frequently used in computer vision operations, which together can enable computer vision algorithms to be executed in real time. The video pipeline can be configured to recognize objects. Objects can be recognized by interpreting digital and / or symbolic information to determine that the visual data represents a specific type of object and / or feature. For example, the number of pixels in the video data and / or the color of the pixels can be used to identify a portion of the video data as an object. The hardware modules 809a - 809n may enable computationally intensive operations (eg, computer vision operations, video encoding, video transcoding, etc.) to be performed locally on the monocular camera stereo device 800 .

[0113] One of the hardware modules 809a-809n (e.g., 809a) can implement a scheduler circuit. The scheduler circuit 809a can be configured to store a directed acyclic graph (DAG). In an example, the scheduler circuit 809a can be configured to generate and store a directed acyclic graph. The directed acyclic graph can define video operations to be performed to extract data from a video frame. For example, the directed acyclic graph can define various mathematical weights (e.g., neural network weights and / or biases) to be applied when performing computer vision operations to classify various pixel groups as specific objects.

[0114] The scheduler circuit 809a can be configured to parse the acyclic graph to generate various operators. The operators can be scheduled by the scheduler circuit 809a in one or more of the other hardware modules 809a-809n. For example, one or more of the hardware modules 809a-809n can implement a hardware engine configured to perform a specific task (e.g., a hardware engine designed to perform specific mathematical operations repeatedly used to perform computer vision operations). The scheduler circuit 809a can schedule operators based on when the operators are ready to be processed by the hardware engines 809a-809n.

[0115] The scheduler circuit 809a can time-multiplex tasks to hardware modules 809a-809n based on their availability to perform work. The scheduler circuit 809a can parse the directed acyclic graph into one or more data streams. Each data stream can include one or more operators. Once the directed acyclic graph is parsed, the scheduler circuit 809a can assign the data streams / operators to the hardware engines 809a-809n and send the relevant operator configuration information to start the operators.

[0116] The binary representation of each directed acyclic graph can be an ordered traversal of the directed acyclic graph, where descriptors and operators are interwoven based on data dependencies. Descriptors typically provide registers that link data buffers to specific operands in dependent operators. In various embodiments, an operator may not appear in the directed acyclic graph representation until all dependency descriptors have been declared for the operands.

[0117] One of the hardware modules 809a-809n (e.g., 809b) can implement a convolutional neural network (CNN) module. The CNN module 809b can be configured to perform computer vision operations on the video frames. The CNN module 809b can be configured to perform object and / or event recognition through multi-layer feature detection. The CNN module 809b can be configured to calculate descriptors based on the feature detection performed. The descriptors can enable the processor 802 to determine the likelihood that a pixel of a video frame corresponds to a specific object (e.g., a person, a pet, an item, text, etc.).

[0118] The CNN module 809b can be configured to implement convolutional neural network capabilities. The CNN module 809b can be configured to use deep learning techniques to implement computer vision. The CNN module 809b can be configured to use a training process through multi-layer feature detection to implement pattern and / or image recognition. The CNN module 809b can be configured to perform inference for a machine learning model.

[0119] The CNN module 809b can be configured to perform feature extraction and / or matching only in hardware. Feature points typically represent areas of interest (e.g., corners, edges, etc.) in a video frame. By temporarily tracking feature points, an estimate of the self-motion of the capture platform or a motion model of an object observed in the scene can be generated. In order to track feature points, a matching algorithm is typically incorporated into the CNN module 809b by hardware to find the most likely correspondence between the feature points in the reference video frame and the target video frame. In the process of matching pairs of reference feature points and target feature points, each feature point can be represented by a descriptor (e.g., image block, SIFT, BRIEF, ORB, FREAK, etc.). Using dedicated hardware circuits to implement the CNN module 809b can achieve real-time calculation of descriptor matching distances.

[0120] The CNN module 809b can be a dedicated hardware module configured to perform feature detection of video frames. The features detected by the CNN module 809b can be used to calculate descriptors. The CNN module 809b can determine the possibility that the pixels in the video frame belong to a specific object and / or multiple objects in response to the descriptors. For example, using descriptors, the CNN module 809b can determine the possibility that the pixels correspond to specific objects (e.g., a person, a piece of furniture, a photo of a person, a pet, etc.) and / or the characteristics of the object (e.g., a person's mouth, a person's hand, a screen of a television, an armrest of a sofa, a clock, etc.). Implementing the CNN module 809b as a dedicated hardware module of the processor 802 can enable the monocular camera stereo device 800 to perform computer vision operations locally (e.g., on-chip) without relying on the processing power of a remote device (e.g., transmitting data to a cloud computing service).

[0121] The computer vision operations performed by the CNN module 809b can be configured to perform feature detection on the video frame to generate descriptors. The CNN module 809b can perform object detection to determine areas in the video frame with a high probability of matching a specific object. In one example, an open operand stack can be used to customize the type of object to be matched (e.g., a reference object) (thereby enabling the programmability of the processor 802 to implement various directed acyclic graphs, each of which provides instructions for performing various types of object detection). The CNN module 809b can be configured to perform local masking to detect objects in areas with a high probability of matching (multiple) specific objects.

[0122] In some embodiments, the CNN module 809b can determine the positions (e.g., 3D coordinates and / or position coordinates) of various features (e.g., characteristics) of the detected object. In one example, the 3D coordinates can be used to determine the positions of the arms, legs, chest, and / or eyes. A position coordinate for the vertical position of the body part in 3D space on the first axis and another coordinate for the horizontal position of the body part in 3D space on the second axis can be stored. In some embodiments, the distance from the lens 808 can represent a coordinate for the depth position of the body part in 3D space (e.g., a position coordinate on the third axis). Using the positions of the various body parts in 3D space, the processor 802 can determine the body position and / or physical characteristics of a person in the field of view of the monocular camera stereo device 800.

[0123] The CNN module 809b can be pre-trained (e.g., the CNN module 809b is configured to perform computer vision to detect objects based on the training data received for training the CNN module 809b). For example, the results of the training data (e.g., a machine learning model) can be pre-programmed and / or loaded into the processor 802. The CNN module 809b can perform inferences for the machine learning model (e.g., to perform object detection). Training can include determining weight values (e.g., neural network weights) for each of a plurality of layers. For example, weight values can be determined for each of a layer for feature extraction (e.g., a convolutional layer) and / or a layer for classification (e.g., a fully connected layer). The weight values learned by the CNN module 809b can vary according to the design criteria of a particular implementation.

[0124] A convolution operation can include sliding a feature detection window along a layer while performing calculations (e.g., matrix operations). The feature detection window can apply filters to pixels and / or extract features associated with each layer. The feature detection window can be applied to a pixel and multiple surrounding pixels. In an example, a layer can be represented as a matrix representing the values of a pixel and / or a feature in a layer, and the filter applied by the feature detection window can be represented as a matrix. The convolution operation can apply matrix multiplication between regions of the current layer covered by the feature detection window. The convolution operation can slide the feature detection window along regions of the layer to generate a result representing each region. The size of the region, the type of operation applied by the filter, and / or the number of layers can vary according to the design criteria of a particular implementation.

[0125] Using convolution operations, the CNN module 809b can calculate multiple features for the pixels of the input image in each extraction step. For example, each of the layers can receive input from a feature set in a small neighborhood (e.g., region) located in the previous layer (e.g., local receptive field). The convolution operation can extract basic visual features (e.g., directional edges, endpoints, corners, etc.), which are then combined by higher layers. Since the feature extraction window operates on pixels and nearby pixels (or sub-pixels), the result of the operation can have position invariance. The layer can include a convolution layer, a pooling layer, a nonlinear layer, and / or a fully connected layer. In an example, the convolution operation can learn to detect edges from raw pixels (e.g., the first layer), then use features from the previous layer (e.g., detected edges) to detect shapes in the next layer, and then use the shapes to detect higher-level features in the higher layer (e.g., facial features, pets, furniture, etc.), and the last layer can be a classifier using higher-level features.

[0126] The CNN module 809b can execute data flows directed to computer vision, feature extraction, and feature matching, including two-stage detection, warping operators, component operators that manipulate lists of components (e.g., components can be regions of vectors that share common attributes and can be grouped together with bounding boxes), matrix inversion operators, dot product operators, convolution operators, conditional operators (e.g., multiplexing and demultiplexing), remapping operators, minimum and maximum reduction operators, pooling operators, non-minimum and non-maximum suppression operators, non-maximum suppression operators based on scanning windows, gather operators, scatter operators, general operators, classifier operators, integral image operators, comparison operators, indexing operators, pattern matching operators, feature extraction operators, feature detection operators, two-stage object detection operators, score generation operators, block reduction operators, and upsampling operators. The types of operations performed by the CNN module 809b to extract features from training data can vary according to the design criteria of a particular implementation.

[0127] Each of the hardware modules 809a-809n can implement a processing resource (or hardware resource or hardware engine). The hardware engines 809a-809n can be operable to perform specific processing tasks. In some configurations, the hardware engines 809a-809n can operate in parallel and independently of each other. In other configurations, the hardware engines 809a-809n can operate collectively with each other to perform assigned tasks. One or more of the hardware engines 809a-809n can be homogeneous processing resources (all circuits 809a-809n can have the same capabilities) or heterogeneous processing resources (two or more circuits 809a-809n can have different capabilities).

[0128] The present invention may be implemented using one or more of a conventional general-purpose processor, a digital computer, a microprocessor, a microcontroller, a RISC (Reduced Instruction Set Computer) processor, a CISC (Complex Instruction Set Computer) processor, a SIMD (Single Instruction Multiple Data) processor, a signal processor, a central processing unit (CPU), an arithmetic logic unit (ALU), a video digital signal processor (VDSP), and / or similar computing machines programmed according to the teachings of the specification. Figure 1-13 The functions illustrated in the diagrams will be readily apparent to those skilled in the relevant art(s). Skilled programmers can readily prepare appropriate software, firmware, coding, routines, instructions, opcodes, microcodes, and / or program modules based on the teachings of this disclosure, as will also be readily apparent to those skilled in the relevant art(s). The software is typically executed from one or more media by one or more of the processors of the machine implementation.

[0129] The invention may also be implemented by preparing an ASIC (application-specific integrated circuit), a platform ASIC, an FPGA (field programmable gate array), a PLD (programmable logic device), a CPLD (complex programmable logic device), a sea-of-gates, an RFIC (radio frequency integrated circuit), an ASSP (application-specific standard product), one or more monolithic integrated circuits, one or more chips or dies arranged as a flip-chip module and / or a multi-chip module, or by interconnecting an appropriate network of conventional component circuits, as described herein, modifications of which will be apparent to those skilled in the art.

[0130] Thus, the present invention may also include a computer product, which may be one or more storage media and / or one or more transmission media, comprising instructions that can be used to program a machine to perform one or more processes or methods according to the present invention. The execution of the instructions contained in the computer product by the machine and the operation of the surrounding circuitry may convert input data into one or more files on the storage medium and / or one or more output signals representing a physical object or substance, such as audio and / or visual depictions. The storage medium may include, but is not limited to, any type of disk, including floppy disks, hard drives, magnetic disks, optical disks, CD-ROMs, DVDs, and magneto-optical disks, as well as circuits, such as: ROM (read-only memory), RAM (random access memory), EPROM (erasable programmable ROM), EEPROM (electrically erasable programmable ROM), UVPROM (ultraviolet erasable programmable ROM), flash memory, magnetic cards, optical cards, and / or any type of medium suitable for storing electronic instructions.

[0131] The elements of the present invention can form part or all of one or more devices, units, components, systems, machines and / or apparatuses. Devices can include, but are not limited to, servers, workstations, storage array controllers, storage systems, personal computers, laptop computers, notebook computers, palmtop computers, cloud servers, personal digital assistants, portable electronic devices, battery-powered devices, set-top boxes, encoders, decoders, transcoders, compressors, decompressors, preprocessors, postprocessors, transmitters, receivers, transceivers, cryptographic circuits, cellular phones, digital cameras, positioning and / or navigation systems, medical equipment, heads-up displays, wireless devices, audio recordings, audio storage and / or audio playback devices, video recordings, video storage and / or video playback devices, gaming platforms, peripheral devices and / or multi-chip modules. (Multiple) related art technicians will understand that the elements of the present invention can be implemented in other types of devices to meet the standards of specific applications.

[0132] When used herein in conjunction with the verb "is," the terms "may" and "generally" are intended to convey the intention that the description is exemplary and is to be considered broad enough to encompass both the specific examples presented in this disclosure and alternative examples that can be derived based on this disclosure. The terms "may" and "generally" as used herein should not be interpreted as necessarily implying the desirability or possibility of omitting the corresponding element.

[0133] While the invention has been particularly shown and described with reference to embodiments thereof, it will be understood by those skilled in the art that various changes in form and details may be made therein without departing from the scope of the invention.

Claims

1. A monocular camera stereo system comprising: an interface configured to receive pixel data; as well as A processor configured to: (i) generate a reference image and a target image based on the pixel data, (ii) perform a parallax operation on the reference image and the target image, and (iii) construct a parallax angle map in response to the parallax operation, wherein the parallax operation includes: (a) selecting a plurality of grid pixels, (b) measuring a parallax angle for each grid pixel, (c) calculating a plurality of coefficients by solving a surface formula of the parallax angle map for the grid pixels, and (d) generating values in the parallax angle map for the pixel data using the coefficients.

2. The monocular camera stereo system according to claim 1, wherein: The processor is further configured to construct a disparity map by performing pattern shift matching using the disparity angle map.

3. The monocular camera stereo system according to claim 2, wherein: The processor is further configured to perform the pattern shift matching by applying a directional block matching process using the disparity angle map.

4. The monocular camera stereo system according to claim 1, wherein: The processor is further configured to construct the disparity angle map by executing an offline calibration procedure, the offline calibration procedure comprising capturing a first image comprising a speckle pattern projected onto a wall at a first distance, and capturing a second image comprising the speckle pattern projected onto the wall at a second distance.

5. The monocular camera stereo system according to claim 1, wherein: The processor is further configured to construct the disparity angle map by executing an online calibration procedure, the online calibration procedure comprising capturing a first image comprising a speckle pattern projected onto an object at a first distance, and capturing a second image comprising the speckle pattern projected onto the object at a second distance.

6. The monocular camera stereo system according to claim 1, wherein: The processor is further configured to calculate the plurality of coefficients by solving the surface formula of the disparity angle map of the grid pixels by applying a regression algorithm to a parametric surface including the grid pixels.

7. The monocular camera stereo system according to claim 6, wherein: The regression algorithm includes a least squares regression algorithm.

8. The monocular camera stereo system according to claim 6, wherein: The parametric surface includes a cubic parametric surface.

9. The monocular camera stereo system according to claim 1, further comprising: a camera configured to generate the pixel data; as well as A structured light projector is configured to project a speckle pattern.

10. The monocular camera stereo system according to claim 9, wherein: The processor is further configured to: generating the reference image comprising the speckle pattern projected onto an object at a first distance from the camera; generating the target image including the speckle pattern projected onto the object at a second distance from the camera; as well as The disparity angle is measured for each grid pixel by determining a pattern shift between the speckle pattern in the reference image and the speckle pattern in the target image.

11. A method for calibrating a structured light projector in a monocular camera stereo system, comprising: Receiving pixel data at the interface; generating a reference image and a target image based on the pixel data using a processor; performing a disparity operation on the reference image and the target image; as well as A disparity angle map is constructed in response to the disparity operation, wherein the disparity operation includes: (a) selecting a plurality of grid pixels, (b) measuring a disparity angle for each grid pixel, (c) calculating a plurality of coefficients by solving a surface formula of the disparity angle map for the grid pixels, and (d) generating values in the disparity angle map for the pixel data using the coefficients. 12 . The method of claim 11 , further comprising constructing a disparity map using the processor by performing pattern shift matching using the disparity angle map.

13. The method according to claim 12, further comprising: The pattern shift matching is performed by applying a directional block matching process using the disparity angle map.

14. The method according to claim 11, further comprising: The disparity angle map is constructed by performing an offline calibration procedure, which includes: capturing a first image including a speckle pattern projected onto a wall at a first distance, and capturing a second image including the speckle pattern projected onto the wall at a second distance.

15. The method according to claim 11, further comprising: The disparity angle map is constructed by performing an online calibration procedure, which includes: capturing a first image including a speckle pattern projected onto an object at a first distance, and capturing a second image including the speckle pattern projected onto the object at a second distance.

16. The method according to claim 11, further comprising: The plurality of coefficients are calculated by applying a regression algorithm to a parametric surface including the grid pixels to solve the surface formula for the disparity angle map of the grid pixels.

17. The method according to claim 16, wherein The regression algorithm includes a least squares regression algorithm.

18. The method according to claim 16, wherein The parametric surface includes a cubic parametric surface.

19. The method according to claim 11, further comprising: generating the pixel data using a camera; as well as A speckle pattern is projected onto the object using a structured light projector.

20. The method according to claim 19, further comprising: generating the reference image comprising the speckle pattern projected onto an object at a first distance from the camera; generating the target image including the speckle pattern projected onto the object at a second distance from the camera; as well as The disparity angle is measured for each grid pixel using the processor to determine a pattern shift between the speckle pattern in the reference image and the speckle pattern in the target image.

21. A machine-readable medium embodying a program which, when executed by one or more processors, causes the one or more processors to perform the method according to any one of claims 11 to 20.

Citation Information

Patent Citations

  • Memory hierarchy to transfer vector data for operators of a directed acyclic graph

    US10437600B1

  • Using camera data to manage a vehicle parked outside in cold climates

    US11001231B1

  • Generating training data for speed bump detection

    US11586843B1

  • Generating detection parameters for a rental property monitoring solution using computer vision and audio analytics from a rental agreement

    US11645706B1

  • Stereo dense matching method and system

    CN106780442A