Data-driven 3D scanning
The data-driven structured light scanning method using RGB patterns for direct depth encoding addresses the inefficiencies of triangulation-based methods, achieving high-speed and accurate 3D scanning with a custom analog projector.
Patent Information
- Application Number
- PCT/US2025/012626
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-23
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-31
AI Technical Summary
Existing structured light scanning methods are limited by the need for triangulation and require multiple patterns for depth reconstruction, which is inefficient for fast-moving objects and leads to reconstruction errors.
A data-driven approach using a single pattern with three color channels (RGB) to directly encode per-pixel depth, eliminating the need for triangulation and allowing high-speed scanning with a custom analog projector.
Enables high-framerate (up to 800 fps) and high-resolution 3D scanning with reduced hardware costs and improved accuracy, while being resilient to shadows and noise.
Smart Images

Figure US2025012626_31072025_PF_FP_ABST
Abstract
Description
DATA-DRIVEN 3D SCANNINGCROSS-REFERENCE TO RELATED PATENT APPLICATION
[0001] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 624,230, filed on January 23, 2024, the entirety of which is incorporated by reference herein.STATEMENT OF GOVERNMENT INTEREST
[0002] This invention was made with government support under 1835712, and 1828576 awarded by the National Science Foundation. The government has certain rights in the invention.TECHNICAL FIELD
[0003] The present disclosure relates generally to data-driven 3D scanning.BACKGROUND
[0004] The acquisition of 3D geometry is ubiquitous in our lives, with its use in face recognition, car navigation, manufacturing, and virtual reality. Various hardware solutions have been introduced to tackle this problem, using many sensing mechanisms including touch, time-of-flight, stereo imaging, structured light, x-rays, or magnetic resonance.
[0005] Structured light (using visible or non-visible frequencies) is popular due to its high resolution and possibility for acquisition of moving objects, with frame rates up to 90 fps. Structured light scanning is a non-contact active technique where light is emitted and its reflection is detected to probe the shape of the scene. A structured light scanner usually includes a projector and a camera. An initial calibration is used to create a digital twin model of the scanner by estimating the intrinsic and extrinsic parameters of the camera and projector. The digital twin is then used to triangulate points in the scene by projecting a unique code per pixel and decoding it on the image plane. The quality of the reconstruction heavily depends on the accuracy of the digital twin used to simulate the camera and projector, usually leading to the use of expensive electronics and high-quality lenses in the hardware setup to achieve high resolution and accuracy. The calibration of the scanner leads to a small set of parameters describing the geometry of the scanner and lens models for the camera and projector. Whenmany patterns are used, this approach is inherently limited to objects moving at low speed due to the necessity of projecting and acquiring multiple coded patterns for each depth reconstruction: for fast-moving objects, the different patterns are projected on a different, moving geometry and thus create large reconstruction errors.
[0006] As described herein, accuracy of reconstructions is inherently limited by triangulation, and additionally an involved procedure to reconstruct the projector from the image is necessary. We propose a different approach that uses the projected colors to directly encode a per-pixel depth, avoiding the need for triangulation. We perform a detailed comparison against the more common patterns proposed by these approaches (binary and continuous). Additionally, we propose an analog projector that allows an acquisition a range scan with only two patterns, one of the two patterns with three color channels, which enables for the design of efficient high-speed projectors given adequate engineering resources.SUMMARY
[0007] One embodiment relates to a method for structured light scanning. The method includes projecting, by a projector, at least one pattern on a calibration plane at a first location and capturing, by a camera, a first image of the pattern on the calibration plane at the first location. The method also includes determining, by a controller, a color of a pixel from the first image, the pixel at a known depth value from the camera and storing, by the controller, a function mapping the color to the depth value.
[0008] Another embodiment relates to a calibration device. The calibration device includes a calibration board and a positioning device. At least a portion of the calibration board is at a known depth away from a camera. A pattern and a normalizing pattern are projected on the calibration board in a switchable manner and the camera is configured to capture an image of the pattern projected on the calibration board. The positioning device moves the pattern and the normalizing pattern along a trajectory. The calibration device is removably coupled to a light scanning system that includes the camera. The calibration device determines a function that includes the pattern and a plurality of depth value mapped to a plurality of colors from pixels of the image.
[0009] Another embodiment relates to a light scanning system. The light scanning system includes a projector to emit structured light and project a pattern and a normalizing pattern on an object. The pattern and the normalizing pattern move along a trajectory. The lightscanning system also includes a camera discretized as a collection of camera rays. The camera obtains a first image of the object with the pattern and a second image of the object with the normalizing pattern. The first image is normalized by the second image. The light scanning system is to capture information for a camera ray, the information to include a color for a pixel of the normalized first image. The light scanning system is to use a function that includes the pattern and a plurality of depth value mapped to a plurality of colors to determine a depth value of the pixel from the camera.
[0010] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the subject matter disclosed herein. In particular, all combinations of claimed subj ect matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein.BRIEF DESCRIPTION OF THE FIGURES
[0011] The foregoing and other features of the present disclosure will become more fully apparent from the following description and appended claims, taken in conjunction with the accompanying drawings. Understanding that these drawings depict only several implementations in accordance with the disclosure and are therefore not to be considered limiting of its scope, the disclosure will be described with additional specificity and detail through use of the accompanying drawings.
[0012] FIG. 1 is a schematic of a reconstruction, according to an example embodiment. A projector illuminates a scene with a specific pattern, whereas we highlight a single camera array in the scene. A simple L2-norm for the camera ray is run to identify its depth.
[0013] FIG. 2 shows normalized colors (right) obtained by dividing each pixel of an image acquired while projecting pattern Pt(left) with the corresponding pixel of an image obtained by projecting a white pattern (middle).
[0014] FIG. 3 shows active calibration hardware in an example embodiment, composed of a linear stage and a custom ChArUco board.
[0015] FIGS. 4A-4D show comparisons of different RGB patterns tested: Random (FIG. 4A), Lissajous (FIG. 4B), Stairs (FIG. 4C), and Spiral (FIG. 4D). The confusion matricesare normalized on per pattern basis to better represent the confusion structure within each pattern.
[0016] FIGS. 5A-5C show qualitative comparisons comparison of pawn reconstructions: triangulation vs. LookUp (FIG. 5A), gray pattern vs. color pattern (FIG. 5B), and RGB patterns (FIG. 5C). The pawn is plotted in the same position and camera view using different reconstruction methods.
[0017] FIG. 6A-6C shows reconstructions of a dice using different methods: Triangulation+ (FIG. 6A), Gray Lookup3D (FIG. 6B), and Colors LookUp3D (FIG. 6C). FIG 6D shows a white pattern image of the dice.
[0018] FIGS. 7A-7B show principal component analyses of a flat plane. Reconstruction of the flat plane with different methods: triangulation and LookUp (with gray and RGB patterns) (FIG. 7A). The histograms show the count of deviation (in millimeters) of each reconstructed point when compared to a perfectly flat surface. FIG. 7B shows random, Lissajous, stairs, and spiral patterns.
[0019] FIG. 8A is a schematic of a specialized analog projector used for high-speed 3D scanning with alternating color and white patterns. FIG. 8B is a corresponding cross-section of the projector’s CAD model.
[0020] FIGS. 9A-9F show a dynamic scene: a fan spinning at 300 rpm. FIGS. 9A-9C shows the fan two frames apart from the previous frame when recorded at 270 fps. Three commercial devices were used to reconstruct and qualitatively compare a similar scene of a fan spinning: Microsoft Azure Kinect (FIG. 9D), Intel Realsense 455 (FIG. 9E), and Photoneo (FIG. 9F).
[0021] FIGS. 10A-10B show a tall rook figure under white pattern illumination (FIG. 10A) and its static reconstruction (FIG. 10B).
[0022] FIG. 11 illustrates a computer system for use with certain implementations.
[0023] FIG. 12 illustrates a computer system for use with certain implementations.
[0024] Reference is made to the accompanying drawings throughout the following detailed description. In the drawings, similar symbols typically identify similar components,unless context dictates otherwise. The illustrative implementations described in the detailed description, drawings, and claims are not meant to be limiting. Other implementations may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented here. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, and designed in a wide variety of different configurations, all of which are explicitly contemplated and made part of this disclosure.DETAILED DESCRIPTION
[0025] Embodiments described herein relate generally to systems and methods for structured light 3D scanning. This includes a calibration hardware and an effective reconstruction algorithm. The configuration of the systems can be applied to a multitude of applications requiring the acquisition of 3D data, as it allows for high-quality and high- framerate scanning with low cost hardware. A standard structured light 3D scanner composed of a digital camera and a digital projector. Additionally, a hybrid scanner hardware can be used, including a digital camera and an analog projector. The hybrid scanner hardware with the calibration method allows for depth acquisition at high-framerate (800 fps) and 1 megapixel resolution.
[0026] The systems and methods, which can be referred to as LookUp3D (LU3D) or LookUp, eliminate the need for a digital twin model. In some embodiments, this approach relies entirely on raw calibration data for 3D reconstruction. Our core observation is that it is not necessary nor beneficial to use a calibration procedure to fit a low-parametric scanner model, as the model’s accuracy will be an upper bound on the reconstruction accuracy. Instead, the calibration procedure is extended to acquire a per-pixel function mapping colors to depth. During calibration, normalized colors for all possible depth values are acquired. With this data, the reconstruction is a search on a per-pixel color lookup table. While at a first glance, this might seem unwieldy in terms of calibration complexity, time, and storage, this is not the case. In various embodiments, the calibration can be performed in 10 minutes with a motorized uni-axial linear stage and a large calibration board mounted on top, the storage of the calibration data is less than 10 GB, and the reconstruction time of our CPU-only implementation is around 5 minutes per frame. The benefit of the approach is a major increase in precision, similar accuracy, and a surprising resistance to shadows.
[0027] Additionally, a key benefit of the approach is that a specific sequence of patterns is not required. Any pattern sequence can be used in the calibration and reconstruction procedure, even a sequence of random patterns. Described herein are effects of using different patterns and shown is that a single pattern with three channels (RGB image) is sufficient for high-quality reconstruction. The use of a single pattern enables a digital projector commonly used for light scanning to be replaced with a custom analog projector. The analog projector can include two LED lights, a beam splitter, a camera lens, and a 35 mm film slide. The analog projector has the significant benefits over the digital one of supporting a higher framerate and a lower cost and complexity. When the analog projector is paired with a highspeed camera, a high-speed acquisition system for 3D geometry at framerates up to 800 fps at 1 megapixel resolution can be obtained. The scanning and the reconstruction algorithm are identical to the setup above, making the approach very flexible and adaptable to specific 3D acquisition requirements.Calibration.
[0028] The systems and methods, as described herein, eliminate a need for a digital twin model and a use of triangulation. Rather, a normalized color for all points within a scanning volume is found. Using the normalized colors, given the color is unique, a 3D reconstruction for a ray is a color lookup for a table of the normalized color for all points.
[0029] A light scanning system includes a camera and a projector. In various embodiments, the camera is a digital camera. The projector is configured to emit a structured light pattern on an object. For example, the pattern can be striped, grids, dots, etc. A sequence of n user-provided mono-chromatic patterns Pk, k E (l,n) can be used. For example, a sequence of binary patterns or gray patterns can be used. The camera captures images of the object with the pattern projected on it and deformations / bending on the pattern on the object can be used for both calibration and reconstruction. Intrinsic parameters of the camera and a response function of the projection is used. In various embodiments, the response function of the projection is acquired with an open-source implementation.
[0030] A frustum in front of the camera is discretized as a discrete collection of camera rays to efficiently process and analyze a view of the camera. The collection of camera rays is represented as Rl,i(d) : R -> R3, where i,j are the integer coordinates of the pixel the ray intersects, and d is the distance to the camera, shown in FIG. 1. A normalized intensity, asdescribed herein, is denoted as Cpl J(d) : R -> R, measured at distance d along the corresponding ray Rtwhen projecting the pattern Pk. Normalized intensity is used to factor out ambient light, object texture, and lens vignetting. The color at each pixel is normalized against a projected white image, as shown in FIG. 2. For each series of projected patterns Pk, FIG. 2 (left image), a full white pattern Pw, FIG. 2 (middle image) is added. IPk(i, j) is denoted as an image acquired by the camera with the pattern Pk. The normalized color image, FIG. 2 (right image) is represented as:
[0031] FIG. 1 also depicts a graph that represents a collection of normalized colors for a range of depths. In various embodiments, the normalized color can use three channels (RGB, as is shown in FIG. 1) such that the normalized color is a combination of the channels. The purpose of the calibration, as described herein, is to acquire, for each pattern Pka complete volume of normalized colors. In other words, a per-ray and per-pattem ID function Cpl^ (d) storing the normalized intensity that an object placed at d distance on the ray would have if the pattern Pnis projected is found. In various embodiments, the light scanning system includes a controller to perform calculations and receive information. The controller, as described herein with reference to FIG. 11 and 12, can be communicatively coupled to a computer, the camera, and / or the projector.
[0032] An active calibration hardware can be used to perform the calibration. The active calibration hardware can include a positioning device and a calibration board, as shown in FIG. 3. In various embodiments, the positioning device is a motorized linear stage (e.g., an off-the- shelf motorized linear stage). The positioning device allows for precise and controlled movements of the light scanning system (e.g., the camera, the projector) to ensure positions, such as a position of the camera relative to the object are known. In various embodiments, the calibration board is custom. The calibration board can be fabricated using aluminum honeycomb and is substantially flat. The flatness of the calibration is a reasonable assumption given that the fabrication tolerance is negligible compared to our target resolution. As shown in FIG. 3, the pattern on the calibration board is a custom pattern with ChArUco tiles on a boundary of the calibration board, and a uniform white in the center. During calibration, the pattern is moved in a trajectory (e.g., linear), stopped every 150 microns, and a stack of images(one for each input pattern Pkplus one for the white pattern Pw) is acquired with the camera while the patterns are projected. When using a color camera and projector it is possible to pack up to 3 patterns in a single image by using three channels (e.g., RGB) of the color camera. A position of the board is reconstructed in each frame relying on the ChArUco tiles using a library, software, and / or database (e.g., OpenCV). A reference frame of each board plane is then leastsquare fitted to a linear trajectory, due to the linear stage fabrication tolerances being low, to further reduce measurement errors.
[0033] As a result, a set of measurements on each camera ray is obtained. The set of measurements include one for each intersection between the calibration board and the camera ray. To obtain a smooth per-ray color function, as desired, a ID cubic B-spline is fitted, as shown in FIG. 3. In various embodiments, the spline fitting is done using a function, such as scipy. interpolate. splrep in SciPy. The fitting is done independently for each projected pattern. In embodiments where the camera has multiple channels, for example RGB, the fitting is done independently on each channel. After fitting, at least a portion of the calibration data can be discarded, such that the output of this stage is a set of ID splines, one for each pixel of the camera and for each pattern to store the normalized color. As such, the calibration systems and method, as described herein, allows for a function containing information on the color and the depth value to be stored for convenient lookup during scanning and reconstruction of an object.Scanning and Reconstruction.
[0034] The scanning procedure is substantially similar to calibration. The patterns are projected in sequence, and the corresponding images are acquired and normalized the images against the projected white pattern, leading to a stack of normalized images Tpn(i, j) . The reconstruction is parallel and done independently on each pixelis defined as the n-vector stacking the intensity for the pixel (i, j) of all acquired images, and C(d) : R -> Rnis the multivariate function returning the stacked intensity of all calibration rays corresponding to the pixel (i,]")
[0035] The reconstruction algorithm is a lookup on the ray to find the color closer to the one acquired:x = arg Eq.4~)
[0036] The reconstruction algorithm is implemented by sampling with a determined spacing (e.g., 10 microns), each ray and performing a search to find the color with the smaller L2 distance. The search can be a brute force search (e.g., linear search) or use other searching algorithms such as binary searching, tree-based searching, etc.
[0037] For optimal reconstruction with the algorithm, it is relied on that each depth value in a ray should have a corresponding distinct color to ensure that the lookup is unique. This is difficult to achieve, and figuring out which pattern gets closer to this idealized desired data can be a challenge. Due to this, in order to not change the projector or the camera focus for different depths, the acquired color will depend not only on the pattern, but also on how the projector and the camera blur and distort it. As described herein, pattern design plays a role in determining depth.
[0038] Pattern design has a huge impact on our ability to estimate the depth from the acquired intensities. Designing pattern sequences is a delicate balance as the use of many patterns makes the depth values easier to differentiate and makes the reconstruction more resilient to noise, but at the same time increases the time required for acquisition. While for static scanning this may not be an issue, for dynamic scanners the use of many patterns negatively affects the acquisition speed and also makes the acquisition of fast objects noisier, as the object will look different as the system cycles through the different patterns.
[0039] Structured light 3D scanning can be viewed as an information transfer problem. For instance, when a scene illuminated by a structured light is imaged with a monochrome camera, there are typically around 12 bits of dynamic range for the light intensity information captured by the camera for each pixel. With no camera sensor noise or distortions, that would allow one to decode 212= 4096 possible depth locations where the scanned object might be along said camera ray. Realistically, a few least significant bits will likely be too noisy to be useful and a few most significant bits must be reserved to account for the dynamic range of object illumination / texture, leaving a couple hundred depth gradations at most which would be insufficient for a high-quality reconstruction.
[0040] One way to improve the situation is to use multiple readings (i.e., use multiple monochrome patterns). Acquiring the patterns sequentially is a great choice for static scanning but has a negative impact on acquisition speed. We can, however, trade spatialresolution for additional dynamic range by using a color Bayer filter on top of the sensors. In this way, three monochromatic patterns can be acquired in one image at a quarter of the image resolution. One bit of spatial resolution is effectively traded to triple the amount of information captured for each pixel. This makes it more feasible to perform reliable scanning with more than a thousand of depth gradations for each pixel even in the presence of image noise and specular material. The systems and methods as described herein, directly encode the per-ray depth and determine a number of patterns that are optimal for this purpose.
[0041] As described herein, a number of pattern types can be used such as monochrome patterns. For example, binary patterns, such as ubiquitous gray code patterns, are very robust to noise and textures by using an entire intensity range (12 bits) to encode only 1 bit. This makes thresholding very reliable, especially when the lighting conditions cannot be controlled. However, this pattern can be inefficient as it is usually necessary to use log (Vr) patterns where r is the image resolution, leading to sequences of over 10 patterns for reasonable resolutions. In another example, a linear ramp can be used in the intensity ratio. This pattern may be very imprecise but uses the full dynamic range. Additionally, this pattern also requires a linear response from both the project and camera, which is very hard to achieve, even after calibration. Another group of pattern, phase shifting (e.g., fringe patterns, sinusoidal patterns, etc.) can be smooth by design and consist of multiple sine waves of different frequencies and amplitudes. These patterns have a property of having constant precision, meaning that if two or more patterns are used, in any point on the domain there is at least one pattern that has a non-zero gradient. This is useful, as it implies that, at least one of the channels, is resilient to noise everywhere in the domain.
[0042] Three channel patterns (e.g., RGB patterns) were used for the benefits as described herein. Three different visualizations are used to compare different patterns to better understand their performance. Although patterns made of stripes were used as they are easier to use in a physical setup, any pattern type may be used. With patterns made of stripes, to maximize effectiveness, the stripes were positioned to be as parallel as possible to a plane perpendicular to the camera rays (i.e., the ID color pattern is mapped to the ray). The usage of stripes is just a practical convenience to make it easy to align the projector in a physical setup and the reconstruction method does not require or use this structure.
[0043] FIGS. 4A-4D show four of our designs in the form of an RGB image (1strow), intensity plots of individual channels labelled with R, G, and B (2ndrow), a plot of the patternin a space where the three coordinates are mapped to colors (3rdrow), and as a confusion matrix (4throw), where a pixelhas the value of the L2 difference between the color at depth i and j. The purpose of the design is to achieve maximal separation. In the 1strow, the RGB image includes a plurality of colors in a pattern. In the 3rdrow (the plot of the pattern in the space), a curve that spans the space without getting too close to itself is desired. In other words, if points on the curve are far, they should be far in space too. In the 4throw (the confusion matrix), a single bright diagonal, denoted as diagonal 400, is desired as that denotes minimal confusion. In various embodiments, the three channels are not symmetric due to the Bayer pattern. The green channel has less noise than the other two channels, thus a signal with higher resolution is placed on that channel.
[0044] FIG. 4A depicts a Random pattern. The Random pattern is generated by sampling random colors and interpolating them with a cubic B-spline with uniform knots. The line plot shows that the Random pattern does not have an optimal performance, with the curve almost intersecting in multiple regions, and the bright confusion matrix.
[0045] FIG. 4B depicts a Lissajous pattern inspired by the Lissajous curve (i.e., the Bowditch curve). The Lissajous pattern is better than the random pattern, as the curve “threads” are more uniformly the space of the color cube.
[0046] FIG. 4C depicts a Stairs pattern. The Stairs pattern uses linear ramps at different frequencies. Disappointingly, the space coverage, shown in the 3rdrow of FIG. 4C, is still not ideal. In addition, the pattern introduces color discontinuities which are difficult to fabricate and will be lost due to camera and projector blur.
[0047] FIG. 4D depicts a Spiral pattern. The Spiral pattern combines the strengths of both Lissajous and Stairs patterns. It uses a linear ramp and a pair of sine and cosine patterns with modulated amplitude to further reduce similarities.
[0048] From the confusion matrices for each pattern (4throw in FIGS. 4A-4D), it is clear that the spiral pattern is the most desirable pattern in this set. It has a consistently lower confusion score the farther away we move along the pattern thanks to the sampling of the color cube space (3rdrow in FIGS. 4A-4D) with maximized distances between consecutive turns. In various embodiments, an optimization method can be applied to further optimize the pattern for a given color response of both camera and projector.Static Analysis and Experiments.
[0049] To evaluate the effect of the calibration and reconstruction method, as described herein, in isolation, a structured light scanner proposed in [Koch et al. 2021] was built and their software for calibration and reconstruction was used. A comparison between reconstructions using a previous method (i.e., a method used in [Koch et al. 2021]) and using the reconstruction method as described herein was made. The previous method is a triangulation-based structure light approach with gray code patterns. The reconstruction method uses a variety of patterns.Qualitative Comparison.
[0050] FIG. 5A-5C shows reconstructions of a pawn figure coated with white specular paint. The reconstructions show when and how each method / pattern combination behaves for a specular material. FIG. 5A (left) shows a triangulation-based reconstruction using gray code patterns with (top-left) and without (bottom-left) grouping of camera rays corresponding to the same projector rays. This is a refinement stage implemented by [Koch et al. 2021] that accounts for the camera’s resolution being higher than that of the projector, and it yields cleaner point clouds but with distinct discretization of the object’s surface. This refinement stage is referred to as triangulation+. Our method, in contrast, offers smooth point clouds at the full resolution of the camera instead of being bound by the projector resolution (FIG. 5A, right). However, our method is more sensitive to the specularities, which are captured as bumps on the object’s surface. Due to memory limitations, our method uses only one quarter of the patterns used by the triangulation method (we use vertical stripes without inverse images).
[0051] FIG. 5C shows the split-rendered comparison of reconstructions with our method using four different RGB patterns, as described with reference to FIGS. 4A-4D. Top-left of FIG. 5C shows Lissajous patten, top-right shows Spiral pattern, bottom-left shows Random pattern, and bottom-right shows Stairs pattern. Due to the constant prevision smooth design of the Lissajous and Spiral patterns, the reconstructions with those patterns are more desirable. Random pattern, as expected, has issues in random locations due to color confusion at different depths. Similarly, the Stairs pattern is a non- smooth attempt at constructing a constant precision pattern and, thus, suffers from loss of precision at regular intervals when its high- frequency component has a rapid change in the gradient direction (see FIG. 4C).
[0052] FIG. 5B shows a split-rendered comparison of, using our method, a standard stack of 11 gray code patterns (left) against a stack of 16 monochrome patterns (right), grouped into 4 RGB images. The performance of the color pattern is superior due to the smoother design, making it a better choice than the gray code patterns.Quantitative Comparison.
[0053] For quantitative comparison, reconstructions for different method / pattern combinations are performed. FIG. 6A-6C shows reconstructions of a dice using different methods: Triangulation+ (FIG. 6A), Gray Lookup3D (FIG. 6B), and Colors LookUp3D (FIG. 6C) FIG 6D shows a white pattern image of the dice. Gray Lookup3D includes a use of gray code patterns while Colors LookUp3D includes color (e.g., RGB) patterns.
[0054] FIGS. 7A and 7B shows reconstruction noise for different method / pattern combinations by scanning a flat plane object. Principal Component Analysis (PCA) is then performed on the cropped portion of the said plane to obtain a standard deviation of the reconstructed point from the fitted planes, as shown in FIGS. 7A-7B. The baseline triangulation method, labelled as triangulation 702, performed the worst with a standard deviation of ~ 75 pm and the refining stage almost halved this number to ~ 40 pm. Our method, labelled as LookUp Gray 706 and LookUp Colors 708, has similar performance with a- ~ 35 pm regardless of patterns stack. This is likely due to the limited flatness / precision of the scanned object itself. Additionally, the Random and Staircase patterns, labelled as random 712 and stairs 716, respectively, performed the worst among RGB patterns (a- ~ 50 pm), with Lissajous, labelled as Lissajous 714, and Spiral, labelled as spiral 718, patterns offering the same degree of precision as the triangulation+ method but with a single image (a- ~ 40 pm). The results show the benefits of different patterns that can be used along the systems and methods for calibration and scanning.Dynamic Analysis and Experiments.
[0055] With our approach, scanning at x fps requires a synchronized camera and projector able to capture / project at 2x fps. The projector will alternate projecting a pattern (e.g., the Spiral pattern) and a white pattern for normalization. While off-the-shelves high-framerate camera are available (for example, iPhones can do 240 fps, and the Kronos HD camera used in testing can do Ik fps), the same is not true for projectors. The best color projectors conveniently available off-the-shelves are based on DPL technology and max out at 120 fps.Analog Projector.
[0056] A purely analog projector that can project the two alternating patterns at over Ik fps and at high-resolution is proposed to address the problems with current off-the-shelves color projectors. FIGS. 8A-8B shows a diagram of the purely analog projector. The projector includes two high power (100 Watt) LEDs Lyrontand Lsideilluminating the corresponding diffusers. The LEDs are located at 90 degrees angle to each other. A color film slide with the color pattern exposed onto it is placed in front of one of the diffusers and a light trap against another with a beam splitter in between. In various embodiments, the film slide is 35 mm. The LEDs are individually controlled with power transistors and a microcontroller, which also triggers the camera in sync with the projector. With tis configuration, we can alternate between the color pattern and white at frequencies greater than 1 kHz. A Chronos 2.1 Full HD camera is used with a Sigma 35 mm fl.4 that can record bursts of frames at Ik fps. The projector uses a Sigma 50 mm fl.4. To expose the film, a calibrated static 4k projector and a Canon Al analog camera with a Tamron 90 mm f / 2.8 are used.
[0057] The analog projector can be used with our methods, as described herein, and a single pattern. Adding additional patterns to this setup, may require additional beamsplitters, higher power sources, careful alignment, and additional lenses.High FPS.
[0058] For high-speed capture, the camera and the projector are configured to use 1.25 ms exposure for the pattern frame and 1 ms for white. This is done to compensate for the fact that the film does not transmit the light perfectly, as part of it may get inevitably absorbed by the substrate, and, thus, boost the signal -to-noise ratio during color normalization. This results in scanning speed of 400 pattern pairs per second, which we will refer to as 400 fps.
[0059] FIGS. 9A-9C show a progression of a dynamic scene where a silicon dice-shaped cube falls and bounces off an air balloon, causing it to squish and bounce off in the opposite direction. The entire action (T2-T0) took 38 ms (i.e., a single frame at video speed (26 fps)). Of note is the partial detection of the falling cube before (Toshown in FIG. 9A) and after (T2shown in FIG. 9C) the still moment Tx, shown in FIG. 9B. This is due to the increased noise / error during its reconstruction because of the motion blur of the pattern frame, leading to depth bias / confusion. As such, some points do not pass the temporal filtering stage. Thefiltering is implemented as a basic moving average on three consecutive frames with an additional condition on points not differing in depth by more than 1 cm from frame to frame.Low FPS.
[0060] The high-speed scanner prototype is compared with three commercial depth sensing devices: Photoneo MotionCAM 3D M+, Intel Realsense 455, and Microsoft Azure Kinect for the dynamic scene as described with reference to FIGS. 9A-9C. The exact same scenes of FIGS. 9A-9C cannot be compared as all three scanners use infrared and do interfere with each other if used at the same time. Thus each scene was repeated for each scanner, striving to keep the initial conditions as similar as possible. There is a discrepancy in the framerate of the scanners (8 fps Photoneo, 90 fps Real Sense, 30 fps Azure Kinect) compared to 400 fps for our scanner.
[0061] Photoneo offers cleaner point clouds, although at a very low rate (8 fps) and high resolution. However, fast-moving objects are not captured at all. Real Sense has the highest frame rate (90 fps), but this is compensated by a low resolution and the point cloud tends to blur / smooth out a lot for fast-moving objects. Kinect strikes a balance between the previous two scanners in terms of scanning speed and resolution. However, fast-moving objects tend to have a lot of missing points in them, which look like holes in the reconstruction.
[0062] Our scanner, in contrast, compromises on neither speed of scanning nor point cloud resolution. However, the 35 mm film limits the power of the LEDs we can use as increasing the power further will lead to too much heat (even with our active cooling system) that would melt the film. The current 100 W power is plentiful for low fps, but it is the limiting factor at 800 fps, resulting in more image noise. This limitation could be lifted with a more hardware engineering, for example using a large format film.
[0063] FIG. 10B shows a reconstruction of the static rook figure as captured by the Chronos camera in the analog setup, shown in FIG. 10A. We replace our temporal filtering stage with an average of five pattern pairs. This greatly reduced the noise in the reconstructed point cloud revealing the great potential of our approach if engineered thoroughly to achieve higher brightness of the projected pattern. In order to accommodate exposure times as short as 1 ms, bright lenses (f / 2.0 for the camera and f / 1.4 for the projector) were used, resulting in a very shallow depth of field. However, our method performed well on all regions of such a large object (23.5 cm tall). An additional surprising observation is that some of the horizontalledges of the object, shown in FIG. 10A, have been reconstructed reasonably well despite being in shadow to the direct projector sight. This may be due to the secondary scattering of the light from the overhang above the ledge.
[0064] Advantageously, the calibration and reconstruction approach for structured light scanning is introduced and a high-framerate depth scanning using the approach is demonstrated. Our approach can be used with existing structured light scanners and is simple to implement. Additionally, with the analog projector, slow-motion 3D scanning may be introduced.
[0065] The possibility of using arbitrary patterns opens many interesting avenues for future work. For example, it is possible to use multiple synchronized projectors and cameras to increase scanning view without worrying about interference as long as the projectors do not overexpose. In another example, the use of multiple analog projectors multiplexed in time to reduce noise at the cost of lower temporal resolution and the use of invisible light patterns (infrared) potentially combined with visible light is possible. The arbitrary patterns also offer alternatives to the use of a film in the analog project to support higher-power projectors.
[0066] The system can also be used to acquire reference geometry for deformable objects in high-speed contact scenarios, and in robotic state estimation for soft object manipulation. With an increase in a reconstruction time of objects scanned with the system, applications for the system can be expanded. The increase in the reconstruction time can be achieved by developing GPU / TPU accelerations or finding ways to compress our calibration data, among other. For example, neural networks can be used to accelerate the depth lookup from the function.Definitions.
[0067] As used herein, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Thus, for example, the term “a member” is intended to mean a single member or a combination of members, “a material” is intended to mean one or more materials, or a combination thereof.
[0068] As used herein, the terms “about” and “approximately” generally mean plus or minus 10% of the stated value. For example, about 0.5 would include 0.45 and 0.55, about 10 would include 9 to 11, about 1000 would include 900 to 1100.
[0069] It should be noted that the term “exemplary” as used herein to describe various embodiments is intended to indicate that such embodiments are possible examples, representations, and / or illustrations of possible embodiments (and such term is not intended to connote that such embodiments are necessarily extraordinary or superlative examples).
[0070] As used herein, the terms “coupled,” “connected,” and the like mean the joining of two additional intermediate members being integrally formed as a single unitary body with one another or with the two members or the two members and any additional intermediate members being attached to one another.
[0071] As shown in FIG. 11, e.g., a computer-accessible medium 120 (e.g., as described herein, storage members directly or indirectly to one another. Such joining may be stationary (e.g., permanent) or moveable (e.g., removable or releasable). Such joining may be achieved with the two members or the two members and any device such as a hard disk, floppy disk, memory stick, CD-ROM, RAM, ROM, etc., or a collection thereof) can be provided (e.g., in communication with the processing arrangement 110). The computer-accessible medium 120 may be a non-transitory computer-accessible medium. The computer-accessible medium 120 can contain executable instructions 130 thereon. In addition or alternatively, a storage arrangement 140 can be provided separately from the computer-accessible medium 120, which can provide the instructions to the processing arrangement 110 so as to configure the processing arrangement to execute certain exemplary procedures, processes and methods, as described herein, for example. The instructions may include a plurality of sets of instructions.
[0072] System 100 may also include a display or output device, an input device such as a keyboard, mouse, touch screen or other input device, and may be connected to additional systems via a logical network. Many of the embodiments described herein may be practiced in a networked environment using logical connections to one or more remote computers having processors. Logical connections may include a local area network (“LAN”) and a wide area network (“WAN”) that are presented here by way of example and not limitation. Such networking environments are commonplace in office-wide or enterprise- wide computer networks, intranets and the Internet and may use a wide variety of different communication protocols. Those skilled in the art can appreciate that such network computing environments can typically encompass many types of computer system configurations, including personal computers, hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers,and the like. Embodiments of the invention may also be practiced in distributed computing environments where tasks are performed by local and remote processing devices that are linked (either by hardwired links, wireless links, or by a combination of hardwired or wireless links) through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0073] FIG. 12 shows the system 100, according to another embodiment. The system 100 can include a processing unit 103 arrangement to execute certain exemplary procedures, processes and methods. The processing unit 103 can have a same configuration as the processing arrangement. The processing unit 103 can include a computer program 104 and a CPU / GPU 105. As described herein, calibration data 101 and frames 102 are communicatively coupled to the processing unit 110 to reconstruct an object, shown as object 600. In various embodiments, the reconstructed object, shown as reconstruction results 106, can be displayed on a display 300. As shown in FIG. 12, in various embodiments, a controller 200 can be remote of the system 100. The controller 200 can have a same configuration as the computer-accessible medium 120. In various embodiments, the controller 200 is coupled to the system 100 by a USB. In various embodiments, the controller 200 controls operation of at least one of a camera 400 and a projector 500.
[0074] Various embodiments are described in the general context of method steps, which may be implemented in one embodiment by a program product including computerexecutable instructions, such as program code, executed by computers in networked environments. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. Computer-executable instructions, associated data structures, and program modules represent examples of program code for executing steps of the methods disclosed herein. The particular sequence of such executable instructions or associated data structures represents examples of corresponding acts for implementing the functions described in such steps.
[0075] Software and web implementations of the present invention could be accomplished with standard programming techniques with rule-based logic and other logic to accomplish the various database searching steps, correlation steps, comparison steps and decision steps. It should also be noted that the words “component” and “module,” as used herein and in the claims, are intended to encompass implementations using one or more linesof software code, and / or hardware implementations, and / or equipment for receiving manual inputs.
[0076] It is important to note that the construction and arrangement of the various exemplary embodiments are illustrative only. Although only a few embodiments have been described in detail in this disclosure, those skilled in the art who review this disclosure will readily appreciate that many modifications are possible (e.g., variations in sizes, dimensions, structures, shapes and proportions of the various elements, values of parameters, mounting arrangements, use of materials, colors, orientations, etc.) without materially departing from the novel teachings and advantages of the subject matter described herein. Other substitutions, modifications, changes and omissions may also be made in the design, operating conditions and arrangement of the various exemplary embodiments without departing from the scope of the present invention.
[0077] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular implementations of particular inventions. Certain features described in this specification in the context of separate implementations can also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation can also be implemented in multiple implementations separately or in any suitable sub combination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
Claims
WHAT IS CLAIMED IS:
1. A method for structured light scanning, comprising: projecting, by a projector, at least one pattern on a calibration plane at a first location; capturing, by a camera, a first image of the pattern on the calibration plane at the first location; determining, by a controller, a color of a pixel from the first image, the pixel at a first known depth value from the camera; and storing, by the controller, a function mapping the color to the first known depth value.
2. The method of claim 1, further comprising: projecting, by the projector, a normalizing pattern on the calibration plane at the first location; capturing, by a camera, a second image of the normalizing pattern on the calibration plane at the first location; and normalizing, by the controller, the first image with the second image to create a normalized first image, the normalized first image used to determine the color of the pixel.
3. The method of claim 1, wherein at least a portion of calibration data, including the first image is deleted after completion of the function.
4. The method of claim 1, wherein the calibration plane includes a calibration board.
5. The method of claim 1, wherein at least one of the patterns is a color pattern and is a random pattern, a Lissajous pattern, a stairs pattern, or a spiral pattern.
6. The method of claim 1, further comprising: after determining the color of the pixel, moving the pattern along a trajectory to a second location, the first location along the trajectory; projecting, by the projector, the pattern on the calibration plane at the second location; capturing, by a camera, a second image of the pattern on the calibration plane at the second location; determining, by a controller, a second color of a second pixel from the second image, the second pixel at a second known depth value from the camera; andstoring, by the controller, the second color to the second known depth value in the function.
7. The method of claim 1, further comprising: projecting, by a projector, the pattern on an object after storing of the function; capturing, by a camera, an image of the pattern on the projection plane at the location; determining, by a controller, a color of a first pixel from the first image; and searching, by the controller, to determine a first depth value associated with the color.
8. The method of claim 7, wherein a portion of the object is reconstructed using the first depth value.
9. The method of claim 7, further comprising: determining, by the controller, a color for each of a plurality of pixels from the first image; searching, by the controller, to determine a depth value associated with the color for each of the plurality of pixels; and reconstructing at least a portion of the object using the depth values.
10. The method of claim 1, wherein the projector comprises at least one LED light, a cameras lens, and at least one of a beamsplitter or a film slide.
11. The method of claim 1, wherein the camera includes at least three channels and the pattern includes the at least three channels.
12. A calibration device, comprising: a calibration board, at least a portion of the calibration board is at a known depth away from a camera, a pattern and a normalizing pattern projected on the calibration board in a switchable manner, the camera to capture an image of the pattern projected on the calibration board; and a positioning device to move the pattern and the normalizing pattern along a trajectory;wherein the calibration device is removably coupled to a light scanning system that includes the camera, the calibration device to determine a mapping function that includes the pattern and a plurality of depth value mapped to a plurality of colors from pixels of the image.
13. The calibration device of claim 12, wherein the positioning device is a motorized linear stage.
14. The calibration device of claim 12, wherein the camera captures a plurality of images of the pattern projected on the calibration board at a number of locations along the trajectory using the positioning device, the plurality of images used to determine the mapping function.
15. The calibration device of claim 12, wherein the image is normalized with a second image of the normalizing pattern projected on the calibration board.
16. A light scanning system, comprising: a projector to emit structured light and project a pattern and a normalizing pattern on an object; and a camera discretized as a collection of camera rays, the camera to obtain a first image of the object with the pattern and a second image of the object with the normalizing pattern, the first image normalized by the second image; wherein the light scanning system is to capture information for a camera ray, the information to include a color for a pixel of the normalized first image, the light scanning system to use functions that includes the pattern and a plurality of depth value mapped to a plurality of colors to determine a depth value of the pixel from the camera.
17. The light scanning system of claim 16, wherein the functions is determined by a calibration device, the calibration device to include a calibration board, at least a portion of the calibration board at a known depth value away from the camera, the pattern and the normalizing pattern projected onto the calibration board to create an image, the functions determined by mapping determined colors for a plurality of pixels from the image to the depth value of the plurality of pixels.
18. The light scanning system of claim 16, wherein the information for the camera ray includes a color for each of a plurality of pixels of the normalized first image, the lightscanning system to use the functions to determine a depth value associated with the color for each of the plurality of pixels, the depth values used of the plurality of pixels used to reconstruct a portion of the object.
19. The light scanning system of claim 16, wherein the projector includes at least one LED light, a beamsplitter, a camera lens, and a film slide.
20. The light scanning system of claim 16, wherein the camera includes at least three channels and the pattern includes the at least three channels.
Citation Information
Patent Citations
Imaging apparatus with projector
JP2010016476A
Image recognition device
US20070076951A1
Scanned-beam depth mapping to 2d image
US20110279648A1
Display apparatus and calibration method therefor
US20120320221A1
Light field display metrology
US20170122725A1