Method for determining one or more groups of exposure settings for use in a 3D image acquisition process - Patents.com
The method addresses the challenges of 3D HDR imaging by optimizing exposure settings to improve SNR and accuracy in 3D coordinate calculation, resulting in higher quality 3D image acquisition across bright and dark areas.
Patent Information
- Application Number
- JP2021576112
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-06-21
- Filing Date
- 2020-06-19
- Publication Date
- 2025-05-07
- Estimated Expiration
- 2040-06-19
AI Technical Summary
Current 3D HDR imaging techniques face challenges in accurately determining exposure settings for capturing high-quality 3D images, particularly in balancing signal-to-noise ratio (SNR) across bright and dark areas, and in handling the increased computational complexity of 3D coordinate calculation.
A method for determining groups of exposure settings in a 3D image acquisition process, which involves identifying candidate groups, assessing signal reception at different pixels, determining exposure costs, and selecting optimal groups based on optimization criteria to maximize SNR and minimize exposure costs.
The method effectively improves the SNR in 3D HDR images by optimizing exposure settings, ensuring accurate calculation of 3D coordinates, and enhancing the overall quality of 3D image acquisition across varying light conditions.
Smart Images

Figure 0007671991000066 
Figure 0007671991000067 
Figure 0007671991000068
Abstract
Description
[Technical field]
[0001] The embodiments described herein relate to a method for determining one or more groups of exposure settings for use in a 3D image acquisition process. The embodiments described herein also relate to a system for generating a three-dimensional image of a subject. [Background technology]
[0002] Three-dimensional surface imaging (3D surface imaging) is a rapidly growing field of technology. As used herein, the term "3D surface imaging" can be understood to refer to the process of generating a 3D representation of a subject's surface by capturing spatial information all in three dimensions, in other words, by capturing depth information in addition to the two-dimensional spatial information present in a traditional image or photograph. This 3D representation can be visually displayed on a screen, for example, as a "3D image."
[0003] A number of different techniques may be used to acquire the data needed to generate a 3D image of the object's surface. These techniques include, but are not limited to, structured light illumination, time-of-flight imaging, holographic techniques, stereo imaging (both active and passive), and laser line triangulation. In each case, the data may be captured in the form of a "point cloud" in which intensity values are recorded for different points in three-dimensional space, with each point in the cloud having its own set of (x,y,z) coordinates and an associated intensity value I.
[0004] Figure 1 shows an example of how data in a point cloud may be stored in memory. As shown, the data is stored in the form of a two-dimensional point cloud matrix 101 with N rows and M columns. Each element in the matrix represents the coordinate system {x ij y ij ,z ij} coordinates, where i={1,2,...N}, j={1,2,...M}, and N and M are integer values. The data tuple also contains the intensity value I ij The intensity values and their respective spatial coordinates together define the shape of the exterior surface of the object under consideration. These points can be rendered in three dimensions to provide a 3D representation of that object.
[0005] In some cases, point cloud data may be calculated from one or more two-dimensional images of a subject collected using a 2D sensor array. In such cases, elements in the point cloud matrix may be mapped to pixel elements in the sensor array. For example, for a given matrix element, the i and j indexes may indicate the location of the respective pixel element in the sensor array. ij y ij ,z ij} coordinates then define the location in space of a point found within a pixel in one or more of the two-dimensional images.
[0006] As with all forms of imaging, the signal-to-noise ratio (SNR) in the final 3D image will be determined in part by the dynamic range of the sensor used to capture that image data. For example, if the surface of the object contains some very bright and very dark areas such that the signal strength received from different points on the object can vary considerably, a balance must be struck between (i) maximizing the illumination intensity and / or sensor exposure time to ensure that sufficient light is received from the darker areas of the object, and (ii) minimizing the illumination intensity and / or sensor exposure time to avoid saturating the sensor with signal from the brighter areas of the object. To address this issue, one proposed solution is to apply the concept of High Dynamic Range (HDR) imaging to 3D imaging. High Dynamic Range Imaging (HDR) is an established technique for increasing the dynamic range of light levels seen in a digital image. This technique involves capturing several images of the same scene with different exposure times and post-processing the data from the images to produce a single HDR image of the scene. Images captured with longer exposure times allow for the capture of details in darker areas of a scene that cannot be discerned in an image captured with a shorter exposure time due to insufficient signal reaching the camera, while images captured with shorter exposure times allow for the capture of details in brighter areas of a scene that cannot be discerned in an image captured with a longer exposure time due to camera saturation. By post-processing these images using a suitable HDR algorithm, it is possible to obtain a single high resolution image in which elements of detail are visible throughout both the bright and dark areas of the image.
[0007] The principles of HDR imaging are generally applicable to 3D imaging in the same way as conventional 2D imaging. Figure 2 shows an example of how this can be implemented. In a similar manner to 2D HDR imaging, multiple image datasets can be captured with different exposure settings, but in this case the image datasets contain 3D image data rather than 2D image data, i.e. the image datasets specify the three-dimensional locations of points in the scene being viewed. Each image dataset can be stored in the form of a respective point cloud matrix 201a, 201b, 201c. The point cloud matrices are then merged into a single point cloud matrix 203 from which the 3D HDR image can be rendered. 3D HDR imaging, however, presents additional challenges compared to 2D HDR imaging. One issue is that, in contrast to 2D imaging, an additional computational step is required to obtain the three-dimensional coordinates of each point, and the effect of any deficiency in SNR in the collected image data is magnified by this additional computation, which can significantly affect the accuracy with which the 3D coordinates of each point in the output image are calculated. It is also difficult to identify the appropriate exposure settings for capturing each set of image data that will be used in generating the final 3D image. [Prior art documents] [Non-patent literature]
[0008] [Non-Patent Document 1] O. Skotheim and F. Couweleers, "Structured light projection for accurate 3D shape determination", ICEM12 - 12th International Conference on Experimental Mechanics, August 29-September 2, 2004, Politecnico di Bari, Italy. [Non-Patent Document 2] Giovanna Sansoni, Matteo Carocci and Roberto Rodella, "Three-dimensional vision based on a combination of Gray-code and phase-shift light protection: analysis and compensation of the systematic errors", Applied Optics, 38, 6565-6573, 1999. [Non-Patent Document 3] Bouquet, G., et al.,"Design tool for TOF and SL based 3D cameras",Opt. Express 25,27758~27769,2017 Summary of the Invention [Problem to be solved by the invention]
[0009] As a result, there is a need to provide improved techniques for generating 3D HDR images. [Means for solving the problem]
[0010] According to a first aspect of the present invention there is provided a method for determining one or more groups of exposure settings for use in a 3D image acquisition process carried out using an imaging system, the imaging system comprising an image sensor, the 3D image acquisition process comprising capturing one or more sets of image data with the image sensor using respective groups of exposure settings, the one or more sets of image data such that one or more 3D point clouds defining three dimensional coordinates of points on a surface of one or more objects being imaged are generated, each group of exposure settings specifying values for one or more parameters of the imaging system that will affect an amount of signal reaching the image sensor, the method comprising: (i) identifying one or more candidate groups of exposure settings using image data captured by an image sensor; (ii) for each candidate group of exposure settings, determining an amount of signal that may be received at different pixels of an image sensor when the candidate group of exposure settings is used to capture a set of image data for use in a 3D image acquisition process; determining whether each pixel would be an overexposed pixel when using the candidate group of exposure settings based on the amount of signal likely to be received at different pixels, where an overexposed pixel is a pixel for which a value of a quality parameter associated with that pixel exceeds a threshold, the value of the quality parameter for a pixel reflecting a degree of uncertainty that would exist in the three dimensional coordinates of a point in the point cloud associated with that pixel if the point cloud were generated using a set of image data captured with the candidate group of exposure settings; determining an exposure cost, the exposure cost being derived from values of one or more parameters in a candidate group of exposure settings; (iii) selecting one or more groups of exposure settings to be used for the 3D image acquisition process from the one or more candidate groups of exposure settings, the selection being such that one or more optimization criteria are satisfied, the one or more optimization criteria comprising: (a) the number of pixels in set N such that a pixel belongs to set N if there is at least one selected group of exposure settings for which the pixel is determined to be an overexposed pixel; and (b) the cost of exposure to one or more selected groups of exposure settings; Steps and Includes.
[0011] In some embodiments, each of the one or more sets of image data is such that it enables generation of a respective 3D point cloud. In such embodiments, each set of image data may include image data captured with a respective group of exposure settings and then combined with previously captured image data to form a new set of image data from which a 3D point cloud can be generated.
[0012] In some embodiments, the distinct pixels include a subset of all pixels of the image sensor. In some embodiments, the distinct pixels all belong to a predefined region of the image sensor. In some embodiments, the distinct pixels include the entirety of the pixels in the image sensor.
[0013] In some embodiments, for each candidate group of exposure settings, the method includes: identifying one or more alternative candidate groups of exposure settings having different values for one or more parameters of the imaging system but the same amount of signal expected to be received at the image sensor; determining an exposure cost for each alternative candidate group of exposure settings, the exposure cost being derived from values of one or more parameters in the alternative candidate group of exposure settings; Including, One or more alternative candidate groups of exposure settings are available to be selected for use in the 3D image acquisition process.
[0014] The selection of the candidate group or groups of exposure settings may be such as to ensure that the ratio of the number of pixels in the set N and the exposure cost for the selected group or groups of exposure settings meets a criterion.
[0015] The selection of one or more candidate groups of exposure settings is (a) the number of pixels in the set N satisfies a first criterion; and (b) the exposure costs for one or more selected groups of exposure settings meet a second criterion. The method may be such as to ensure that:
[0016] The first criterion may be to minimize the number of pixels that belong to set N. The first criterion may be to ensure that the number of pixels that belong to set N exceeds a threshold.
[0017] The second criterion may be to minimize the sum of the exposure costs.The second criterion may be to ensure that the sum of the exposure costs for each of the selected group of exposure settings is below a threshold.
[0018] In some embodiments, for one or more of the candidate groups of exposure settings, determining an amount of signal that may be received at different pixels of the image sensor when the candidate group of exposure settings is used to capture the set of image data includes capturing the set of image data using the candidate group of exposure settings.
[0019] Sets of image data captured when using one or more candidate groups of exposure settings may be used to identify one or more other candidate groups of exposure settings.
[0020] In some embodiments, steps (i) through (iii) are repeated through one or more iterations, and for each iteration: One of the candidate groups of exposure settings identified in that iteration is selected; The optimization criteria include a first criterion which is to maximize the number of pixels in set N, and a second criterion which is to ensure that the sum of the exposure cost for the group of exposure settings selected in the current iteration and the respective exposure costs for the groups of exposure settings selected in all previous iterations is less than a threshold.
[0021] For each iteration, a selected group of exposure settings may be used to capture a set of imaging data with the imaging system; For each iteration after the second iteration, the set of image data captured in the previous iteration is used in determining a candidate group of exposure settings for the current iteration.
[0022] The step of determining whether each pixel will be an exposed pixel when using the candidate group of exposure settings may include determining a probability that each pixel will be exposed, the probability being determined based on the amount of signal received at those pixels in previous iterations of the method.
[0023] The exposure cost for each group of exposure settings may be a function of the exposure time used within that group of settings.
[0024] Identifying one or more candidate groups of exposure settings may include determining, for one or more pixels of the image sensor, a range of exposure times over which the pixel is likely to be an overexposed pixel.
[0025] The value of the quality parameter associated with a pixel may be determined based on the amount of ambient light in the scene being imaged.
[0026] In some embodiments, each group of exposure settings includes: The exposure time of the image sensor, The size of the aperture stop in the path between the subject and the sensor, The intensity of the light used to illuminate the subject, and The strength of the ND filter placed in the light path between the subject and the sensor Includes one or more of:
[0027] The imaging system may be an optical imaging system with one or more photosensors. The imaging system may include one or more light sources that are used to illuminate the subject being imaged.
[0028] The image data in each set of image data may include one or more 2D images of the object captured by the sensor.
[0029] The imaging system may be an imaging system that uses structured illumination to acquire each set of image data.
[0030] Each set of image data may include a sequence of 2D images of a subject captured by a photosensor, with each 2D image in the sequence being captured using a different illumination pattern.
[0031] Each set of image data may include a sequence of gray-coding images and a sequence of phase-shifted images.
[0032] Each set of image data may include color information.
[0033] According to a second aspect of the present invention there is provided a method for generating a 3D image of one or more objects using an imaging system comprising an image sensor, the method comprising: capturing one or more sets of image data with an image sensor using respective groups of exposure settings, the sets of image data being such that one or more 3D point clouds defining three-dimensional coordinates of points on a surface of one or more objects are generated, each group of exposure settings specifying values for one or more parameters of the imaging system that will affect the amount of signal reaching the image sensor; constructing a 3D point cloud using data from one or more of the captured sets of image data; Including, The exposure settings used to capture each set of image data are determined using a method according to the first aspect of the invention.
[0034] According to a third aspect of the present invention there is provided a computer readable storage medium having stored thereon computer executable code which, when executed by a computer, causes the computer to perform a method according to the first aspect of the present invention.
[0035] According to a fourth aspect of the present invention, there is provided an imaging system for performing a 3D image collection process by capturing one or more sets of image data using one or more groups of exposure settings, such that the one or more sets of image data enable generation of one or more 3D point clouds defining three-dimensional coordinates of points on a surface of one or more objects being imaged, the imaging system comprising an image sensor for capturing the one or more sets of image data, the imaging system being configured to determine the one or more groups of exposure settings to use for the 3D image collection process by performing a method according to the first aspect of the present invention.
[0036] The embodiments described herein provide a means for acquiring a 3D HDR image of an object. The 3D HDR image is acquired by capturing multiple sets of image data. The image data may be acquired using one of several techniques known in the art. These techniques may include, for example, structured light illumination, time-of-flight imaging, and holographic techniques, as well as stereo imaging (both active and passive) and laser line triangulation. Each set of image data may be used to calculate a respective input point cloud that specifies the three-dimensional coordinates of points on the object surface and their intensities. The information in each set of image data may be combined or merged in such a way as to provide a single output point cloud from which a 3D image of the object may be rendered.
[0037] The sets of image data may be collected with different groups of exposure settings such that the exposure of the imaging system is different for each set of image data. Here, "exposure setting" refers to a physical parameter of the imaging system that may be adjusted to change the exposure of the system. If an optical imaging system is used, the group of exposure settings may include parameters that directly affect the amount of light incident on the image sensor, examples of such parameters include one or more of the integration time of the camera / light sensor, the size of the camera / light sensor aperture, and the intensity of the light used to illuminate the scene or subject. The group of exposure settings may also include the strength of a neutral density filter placed in the optical path between the subject and the camera / light sensor. The strength of the neutral density filter may be changed each time a new set of image data is collected to change the amount of light reaching the camera or light sensor.
[0038] It will be appreciated that in addition to parameters that affect the amount of light incident on the sensor, other parameters of the system may also vary between exposures. These parameters may include, for example, the sensitivity of the camera or light sensor, which may be changed by adjusting the gain applied to the device.
[0039] In some embodiments, only one of the exposure settings may change for each exposure, while the other exposure settings remain constant each time a new set of image data is collected, in other embodiments, two or more settings may change during the capture of each new set of image data.
[0040] If image data is collected with a greater exposure, this will maximize the signal detected from the darker regions of the object or scene. The SNR for points in the darker regions may thereby be increased, allowing the spatial coordinates of those points to be calculated with greater precision when compiling a point cloud from the set of image data. Conversely, adjusting the exposure settings in such a way as to reduce the exposure allows the signal captured from the brighter regions of the object or scene to be maximized without saturating the camera or light sensor, and thus, by reducing the exposure, it is possible to obtain a second set of image data in which the SNR for points in the brighter regions may be increased. It is then possible to generate a second point cloud in which the spatial coordinates of points in those brighter regions may be calculated with greater precision. Image data from these different exposures may then be combined in such a way as to ensure that the SNR in the output 3D image is improved for both the brighter and darker regions of the object.
[0041] To determine how best to combine the data from each exposure, in the embodiments described herein, an additional "quality parameter" is evaluated for each point in each set of image data. Figure 3 shows an example of a point cloud matrix 301 used in one embodiment. In contrast to the conventional point cloud matrix of Figure 1, each element of the point cloud matrix of Figure 3 is evaluated with an additional value q ij The value of the quality parameter q ij is the {x ij ,y ij ,z ij In some embodiments, the value of the quality parameter reflects the degree of uncertainty in the {x ij ,y ij ,z ij} provides a best estimate of the expected error in the value.
[0042] The quality parameter may be defined in one of several different ways, depending on the particular technique used to capture the image data. In one example, discussed in more detail below, a structured illumination approach is employed to collect image data from which the output point cloud is computed. In this case, the value of the quality parameter for a given point may be derived, for example, from the degree of contrast seen at that point on the surface of the object as bright and dark fringes in the structured illumination pattern are projected thereon. In other words, the quality parameter may relate to or be derived from the difference between the amount of light detected from that point when illuminated by the bright fringes and the dark fringes in the illumination pattern. In some embodiments, a series of sinusoidally modulated intensity patterns may be projected onto the object, with the intensity patterns in successive images being phase shifted relative to each other, where the quality parameter may relate to or be derived from the intensity measured at each point as each of the intensity patterns is projected onto the object.
[0043] In another example, when using a time-of-flight system for 3D imaging, the amplitude of the restored signal at a given point may be compared to a measurement of the ambient light in the system to derive a value of the quality parameter at that point. In a further example, when using a stereo imaging system, the quality parameter may be obtained by comparing and matching features identified in an image captured on a first camera with features identified in an image captured on a second camera and determining a score for the quality of the match. The score may be calculated, for example, as the sum of absolute differences (SAD):
[0044]
number
[0045] where (r,c) are the image pixel coordinates, I1, I2 are the stereo image pair, (x,y) are the disparities for which the SAD is evaluated, and (A,B) are the dimensions (in the form of number of pixels) of the arrays on which the matching is performed. (x',y')=argmin x,y Using SAD(r,c,x,y), a disparity estimate (x',y') for pixel (r,c) is found. Depending on the estimated SAD value, this can then be further interpreted as a measure of uncertainty in the 3D coordinates at each point in the reconstructed point cloud based on the stereo images. If desired, the SAD can be normalized to, for example, regional image intensity to improve performance.
[0046] Regardless of which particular imaging method is used to collect the image data, the quality parameters can be used to determine the extent to which data is taken from each exposure to be taken into account when deriving coordinate values for points in the final 3D image. In some embodiments, each set of image data can be used to construct a respective input point cloud matrix. Then, for a given element {i,j} in the point cloud matrix used to construct the final 3D image, the respective weight is assigned to the coordinate value {x ij y ij ,z ij} to find the coordinate value {x ij y ij ,z ij} can be derived. The weight added to each matrix is the quality parameter q ij In this way, the values {x ij y ij ,z ij} may be biased towards values obtained from image datasets that have a higher SNR at that point.
[0047] In another example, we can use the input point cloud matrix to find the ij y ij ,z ij} value, the coordinate value {x ij y ij ,z ij} may be chosen. The input point cloud matrix from which a value is selected may be the input point cloud matrix that has the highest q value at that point compared to the other input matrices.
[0048] In another example, for a given element {i,j} in the output point cloud matrix, the coordinate value for that point {x ij y ij ,z ij} is the set of coordinates {x ij y ij ,z ij}, i.e., the output point cloud matrix values x ij is the value x at that point {i,j} in each input point cloud matrix. ij and the output point cloud matrix values y ij is the value y at point {i,j} in each input point cloud matrix. ij and the values z ij is the value z at point {i,j} in each input point cloud matrix. ij The averaging may include an average of each of the q values. The averaging may be subject to a threshold criterion, such that only values whose associated q value exceeds a threshold are used in calculating the average. Thus, for example, the values {x ij y ij ,z ij} is accompanied by a q-value that is less than the threshold, the value {x ij y ij ,z ij} are the values {x ij y ij ,z ij} can be discarded to compute
[0049] It will be appreciated that the thresholding criterion may also be applied to other scenarios discussed above. For example, if the coordinate value {x ij y ij ,z ij}, the value {x ij y ij ,z ij}, the coordinate values {x ij y ij ,z ij} may be given a zero weight.
[0050] It will be appreciated that in order to construct a point cloud for the final 3D image, it is not necessary to construct a respective point cloud for each set of collected image data; rather, in some embodiments, by use of a suitable algorithm, it may be possible to infer the value of the quality parameter for a given pixel by considering the actual amount of signal received at that pixel over the course of an exposure, without the need to actually construct a point cloud from that data. To generate the point cloud for the final 3D image, the sets of image data acquired from each exposure may be merged to generate a single merged set of image data. The value for each pixel in the merged set of data may then be based on the values for that pixel in one or more of the collected sets of image data, with a bias for the set of image data for which that pixel is associated with a higher value of the quality parameter.
[0051] Embodiments of the invention will now be described, by way of example only, with reference to the accompanying drawings, in which: [Brief description of the drawings]
[0052] [Figure 1] FIG. 1 is a diagram illustrating an example of a conventional point cloud matrix. [Diagram 2]1A-1C show examples of how multiple point cloud matrices can be merged or combined to form a single point cloud matrix. [Diagram 3] FIG. 2 illustrates an example of a point cloud matrix in one embodiment described herein. [Figure 4] 1 is a flow diagram of steps in one embodiment described herein. [Figure 5A] FIG. 1 is a schematic diagram of a structured illumination imaging system in one embodiment. [Figure 5B] FIG. 5B is a schematic diagram of the geometry of the structured illumination imaging system of FIG. 5A. [Figure 6] FIG. 13 shows an example of how the standard deviation of GCPS values can vary as a function of signal amplitude in an embodiment in which a combination of Gray coding and phase shifting is used to recover 3D spatial information from an object. [Figure 7] FIG. 6 is a schematic diagram illustrating how multiple input point cloud matrices can be acquired using the imaging system of FIG. 5 and used to generate a single output point cloud matrix. [Figure 8] FIG. 2 illustrates an example of a point cloud matrix in one embodiment described herein. [Figure 9] 1 is a graph illustrating how depth noise for a point in a 3D image varies depending on the contrast acquired in the sequence of images used to generate the 3D image. [Figure 10] 1 is a histogram of several highly exposed pixels in an image for different candidate exposure times. [Figure 11] 1 is a histogram of several highly exposed pixels in an image for different candidate exposure times. [Figure 12] FIG. 2 shows a sequence of images representing several well-exposed pixels obtained from cumulative exposures of an image sensor. [Figure 13] FIG. 1 illustrates pseudocode for implementing an algorithm according to one embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0053] 4 shows a flow diagram of the steps performed in the embodiments described herein. In a first step S401, multiple sets of 3D image data are collected by an imaging system. Each set of 3D image data can be used to calculate a respective input point cloud that defines the 3D coordinates of different points on the surface of the object being imaged, together with a value for the intensity or brightness level of the surface at each point.
[0054] In step S402, a value of a quality parameter for data associated with each point in each of the respective input point clouds is evaluated. As discussed above, the value of the quality parameter includes a measure of uncertainty in the three-dimensional coordinates at each point. The quality parameter may be calculated as a function of the collected intensity values used to calculate the spatial coordinates at each point in the respective input point clouds.
[0055] In step S403, a single output set of 3D image data is calculated based on the image data contained in each of the collected sets of image data. Similar to the collected image data set, the output set of image data defines values for the 3D coordinates of different points on the surface of the object being imaged, along with the intensity or brightness level of the surface at each point. Here, the values for the 3D coordinates are calculated by weighting the values for the 3D coordinates specified in each input point cloud according to their respective quality parameter values. The output image data set may then be used to render a 3D image of the object (step S404).
[0056] An exemplary embodiment of using a structured light illumination method to collect 3D image data will now be described with reference to FIGS.
[0057] Referring to FIG. 5A, a schematic diagram of a system suitable for capturing a 3D image of an object 501 using structured light illumination is shown. The system comprises a projector 503 and a camera 505. The projector is used to project a spatially varying 2D illumination pattern onto the object 501. The pattern itself includes a series of bright and dark fringes 507, 509. The pattern may be generated, for example, using a spatial light modulator. The camera 505 is used to collect a 2D image of the object as it is illuminated by the projector.
[0058] Figure 5B shows a simplified diagram of the system geometry. The camera and projector are located a distance B apart. A point on the object a distance D away from the camera and projector is θ c and from the projector at an angle of θ p Due to the angle between the camera and the projector, any change in the surface topology of the object will distort the pattern of bright and dark fringes detected by the camera. Essentially, the distortions in the pattern will encode information about the 3D surface of the object, which can be used to infer its surface topology. The 3D information can be recovered by capturing successive images in which the object is illuminated with different patterns of light, and comparing the intensities measured for each pixel across the sequence of images.
[0059] In this embodiment, a phase shifting technique is used to obtain 3D information. Phase shifting is a well-known technique in which a sequence of sinusoidally modulated intensity patterns is projected onto an object, with each pattern being phase shifted relative to the previous one. A 2D image of the illuminated object is captured each time the intensity pattern changes. Changes in the surface topology of the object will change the phase of the intensity pattern seen by the camera at different points across the surface. By comparing the light intensity in the same pixel across a sequence of 2D images, the phase at each point can be calculated, which can be used to obtain depth information about the object. The data is output as a 2D array, where each element is mapped to a respective pixel of the camera and defines the 3D spatial coordinates of the point seen within that pixel.
[0060] It will be appreciated that other techniques than phase shifting may be used to recover the 3D spatial information. For example, in some embodiments, a gray coding technique may be used, or a combination of gray coding and phase shifting may be used. The exact algorithms used to decode the 3D spatial information from a sequence of 2D images will vary depending on the particular illumination patterns and the manner in which those patterns change across the sequence of images. More information on algorithms for recovering depth information using these and other techniques is available in "Structured light projection for accurate 3D shape determination," O. Skotheim and F. Couweleers, ICEM12 - 12th International Conference on Experimental Mechanics, August 29-September 2, 2004, Politecnico di Bari, Italy. In each case, the 3D spatial information within the object is calculated by considering the change in intensity at each point on the object as the illumination pattern changes and the point is exposed to light and dark areas of the pattern.
[0061] In one example where a combination of Gray coding and phase shifting is used to recover 3D spatial information, a projector is used to project a series of both Gray code and phase shift patterns onto an object. (Further details of such methods can be found, for example, in the article by Giovanna Sansoni, Matteo Carocci, and Roberto Rodella entitled "Three-dimensional vision based on a combination of Gray-code and phase-shift light protection: analysis and compensation of the systematic errors", Applied Optics 38, 6565-6573, 1999.) Now, for each point on the object, two corresponding pixel locations can be defined: (i) projector pixel coordinates, i.e., the pixel location in the projector from which the light incident on that point on the object originates, and (ii) camera pixel coordinates, i.e., the pixel location in the camera from which the light reflected by that point on the object is captured. Using suitable algorithms, the images captured at the camera can be processed to determine the corresponding projector coordinates for each camera pixel, taking into account the relative positions of the camera and projector (which are easily determined using standard calibration measurements). In effect, a determination can be made as to which projector pixel a particular camera pixel is "looking at." Furthermore, by combining the image data received from the Gray code pattern and the phase shift pattern, projector pixel coordinates can be determined with a higher resolution than the projector pixel itself.
[0062] The above method can be understood as follows. First, select N Gray code patterns and set the number of sinusoidal fringes in the phase shift pattern to 2. NBy setting the fringes to 0, the fringes may be aligned with binary transitions in a sequence of N Gray code patterns. The resulting Gray code words GC(i,j) and the values obtained for phase Φ(i,j) may be combined to form a set of “GCPS” values that describe the absolute fringe positions within each location in the field of view. The GCPS values are then scaled to the width (w p ) and height (h p ) can be used to determine projector pixel coordinates. In practice, it is possible to measure "fringe displacement" by estimating the phase of the sine pattern in every pixel in the camera.
[0063] Then it is possible to define:
[0064]
number
[0065] In the formula, GC v (i,j) is the result of the Gray code measurement, and Φ v (i,j) are the results of the phase shift measurements, both of which are performed using the vertical fringes. (As before, the indices i,j refer to pixel elements of the image sensor.) From the above equations, it is possible to calculate the originating sub-pixel projector column for each pixel in the camera image.
[0066]
number
[0067] In the formula, α max and α min are the maximum and minimum values for the GCPS code for the vertical fringes. Similarly, when using the horizontal fringes to obtain GCPS values, it is possible to define:
[0068]
number
[0069] Then, using the equation for β(i,j), it is possible to calculate the origin subpixel projector row for each pixel in the camera image by:
[0070]
number
[0071] In the formula, β max and β min are the maximum and minimum values for the GCPS code for the horizontal fringes.
[0072] The column and row coordinates P of the subpixel projector c (i,j), P r Now we have obtained the coordinates (i,j) of a point on the object being imaged, which can be used to obtain the {x,y,z} coordinates of the point on the object being imaged. In particular, for a given camera pixel p that is established to receive light from a point on the projector g, the coordinates {x,y,z} can be obtained by using known triangulation methods, similar to those used for stereoscopic vision, taking into account the lens parameters, the distance between the camera and the projector, etc. ij y ij ,z ij}, a position estimate E of a point on the object can be derived.
[0073] The uncertainty in each GCPS value will be influenced to a greater extent by the amplitude of the restored signal and, to a lesser extent, by the presence of ambient light. Experiments performed by the inventors have shown that the uncertainty of the GCPS value is generally approximately constant until the amplitude of the received signal falls below a certain level, after which the uncertainty grows approximately exponentially. This means that the measured amplitude and ambient light can be converted into an expected measurement uncertainty of the GCPS value through a pre-established model. An example showing the standard deviation (std) of the GCPS value as a function of intensity is provided in FIG. 6. Using the same calculations used above to obtain the position estimate E, but now by utilizing the projector pixel positions at g+Δg, where Δg is derived from the standard deviation of the GCPS value in the detected signal amplitude, it is possible to obtain a new position estimate E'. An estimated standard deviation ΔE in the position measurement E can then be derived by assuming ΔE=|(|E-E'|)|. The estimated ΔE can then be used to define a quality parameter.
[0074] Regardless of the exact algorithm used for structured illumination, it becomes clear that in order to calculate 3D spatial information with high accuracy, it will be desirable to measure the change in intensity at each point on the object with maximum signal-to-noise. This requires that the contrast in intensity seen as a particular point is alternately exposed to light and dark areas of the illumination pattern be as high as possible. In line with the previous discussion, the range of contrast at a particular point may be limited by the need to accommodate a large range of intensities across the surface of the object. If the surface of the object itself includes several light and dark areas, given the finite dynamic range of the camera, it may not be possible to maximize the signal recovered from the dark areas of the object's surface without also creating a saturation effect where the brighter areas of the object are exposed to the brighter fringes in the illumination pattern. Thus, it may not be possible to optimize the contrast in intensity seen at all points on the surface within a single exposure (it should be understood that in this context, "exposure" refers to the capture of a sequence of 2D images from which respective point clouds can be calculated).
[0075] To address the above issues, the embodiments described herein utilize several exposures using different settings. An example of this is shown diagrammatically in FIG. 7. Initially, a first sequence of 2D images is captured using a first set of exposure settings. In this example, the exposure settings are varied by changing the relative size of the camera aperture, but it will be appreciated that the exposure settings may be varied by adjusting one or more other parameters of the imaging system. Circle 701 indicates the relative size of the camera aperture used to capture the first sequence of images, and patterns 703a, 703b, 703c indicate the illumination patterns projected onto the subject when capturing each image in the first sequence of images. As can be seen, each illumination pattern 703a, 703b, 703c includes a sinusoidally modulated intensity pattern in which a series of alternating bright and dark fringes are projected onto the subject. The phase of the patterns across the field of view is indicated diagrammatically by a wave beneath each pattern, with the dark fringes in the pattern corresponding to the troughs of the wave and the bright fringes corresponding to the peaks of the wave. As can be seen, the illumination patterns in successive images are phase shifted relative to each other, and the positions of the bright and dark fringes in each pattern are transformed relative to each other. A suitable algorithm is used to compare the light intensity in the same pixel across a sequence of 2D images to calculate the depth information at that point. The three-dimensional spatial information is then stored in a first point cloud matrix 705, along with the intensity values for each point.
[0076] In addition to the three-dimensional coordinates {x,y,z} and the intensity value I, each element in the point cloud matrix 705 includes a value of a quality parameter q, as previously shown in Figure 3. In this embodiment, the value q represents the observed intensity I within each camera pixel as the three different illumination patterns 703a, 703b, 703c are projected onto the object. nThe quality parameter is determined based on the standard deviation of (i,j). In another embodiment, the value q is determined based on the difference between the maximum and minimum intensity values in that pixel across the sequence of images. In this way, the quality parameter defines a measure of the contrast seen in each pixel across the sequence of images.
[0077] In the next stage of the process, the exposure setting is adjusted by expanding the size of the camera aperture, as reflected by circle 707, thereby increasing the amount of light from the subject that will reach the camera. A second sequence of 2D images is captured with the illumination patterns 709a, 709b, 709c again different for each image in the sequence. The second sequence of 2D images is used to calculate a second input 3D image point cloud 711 that records depth information for each element along with intensity values and a value for a quality parameter q. Due to the exposure difference between the first sequence of images and the second sequence of images, the degree of contrast seen within a given pixel element {i,j} may be different across the two sequences of images. For example, the difference between the maximum and minimum intensity levels found within a given pixel element {i,j} will be different between the two sets of images. Thus, the value of q recorded in the first point cloud matrix 705 may be different from the value of q in the second point cloud matrix 711.
[0078] In a further step, the exposure settings are further adjusted by expanding the size of the camera aperture as indicated by circle 713. With illumination patterns 715a, 715b, 715c projected onto the subject, a third sequence of 2D images is captured on the camera. The third sequence of 2D images is used to calculate a third input 3D image point cloud matrix 717 that again records depth information for each element, along with intensity values and a value for a quality parameter q. As in the case of the first and second point cloud matrices, the exposure difference between the second and third sequences of images means that the degree of contrast seen within the same pixel element {i,j} may be different across the second and third sequences of images. It follows that the value of q recorded in the third point cloud matrix 717 may be different from both the first point cloud matrix 705 and the second point cloud matrix 711.
[0079] Having calculated a point cloud matrix for each sequence of images, the method proceeds by compiling a single output point cloud matrix 719 using the data in each point cloud matrix, which may then be used to render a 3D representation of the subject.
[0080] While the example shown in FIG. 7 includes a total of three input image sequences, it will be understood that this is by way of example only and that the output point cloud matrix may be calculated by capturing any number N of point clouds, where N>=2. Similarly, in the above example, the exposure was varied by increasing the size of the camera aperture between capturing each sequence of images, but it will be readily understood that this is not the only means by which exposure may be varied. Other examples of ways in which exposure may be varied include increasing the illumination intensity, increasing the camera integration time, increasing the camera sensitivity or gain, or changing the strength of a neutral density filter placed in the optical path between the subject and the camera.
[0081] 7 includes generating an individual point cloud for each set of collected image data, as previously discussed, it will be understood that this is by way of example only and that in some embodiments the output point cloud may be generated without having to construct an individual point cloud for each set of image data. In such cases, values for quality parameters associated with pixels in each respective set of image data may be inferred from the signal levels detected within those pixels, as well as the amount of ambient light incident on the image sensor.
[0082] By capturing multiple sets of image data using different groups of exposure settings and combining the data from those data sets to provide a single output 3D image, the embodiments described herein can help compensate for the limited dynamic range of the camera and provide an improved signal-to-noise ratio for both darker and brighter points on the object's surface. In doing so, the embodiments can help ensure that usable data is captured from regions that would be either completely missing from the final image or otherwise dominated by noise if conventional 3D surface imaging methods were used, and can ensure that the object's surface topography is mapped with greater accuracy compared to such conventional 3D surface imaging methods.
[0083] In the embodiments described above, it is assumed that the camera sensor is imaging in grayscale; that is, a single intensity value related to the total incident light level on the sensor is measured within each pixel. However, it will be appreciated that the embodiments described herein are equally applicable to color imaging scenarios. For example, in some embodiments, the camera sensor may comprise an RGB sensor where a Bayer mask is used to resolve the incident light into red, green, and blue channels, while in other embodiments the camera sensor may comprise a 3CCD device where three separate CCD sensors are used to collect red, blue, and green light signals, respectively. In this case, the point cloud matrix would be collected in the same way as in the embodiments described above, but within each matrix element, the intensity value Iij is the three color intensities r ij , g ij , b ij FIG. 8 shows an example of such a point cloud matrix.
[0084] As previously discussed, the data used to render the final 3D image may be obtained by changing one or more exposure settings, including, for example, illumination, camera integration time (exposure time), increasing the sensitivity or gain of the camera, or by varying the strength of a neutral density filter placed in the optical path between the subject and the camera. Thus, there will be many combinations of different settings that can be used for any one exposure. Some of these combinations may provide a more optimal solution than others. For example, in some cases it may be desirable to choose a group of exposure settings that will minimize the exposure time, while in other cases there may be additional or different considerations to be made, such as the need to keep the size of the aperture constant to avoid changes in depth of field, which will impose other constraints in terms of which parameters of the imaging system are changed and to what extent. In general, when collecting images, it is usually necessary to strike a balance between achieving an acceptable SNR in the image (specifically, the noise in the depth values for each point on the surface of the subject being imaged) and achieving one or more exposure requirements, such as (i) the total duration of collection, (ii) the illumination required for collection, and (iii) the aperture size.
[0085] It should be noted that 3D measurement systems employing active illumination (e.g., structured light, laser triangulation, time-of-flight) differ from regular cameras in that the amount of ambient light has a large effect on the dynamic range of the system if a desired maximum noise level is to be achieved. This means that traditional auto-exposure algorithms cannot be used directly. While these algorithms are typically optimized solely for a sufficient total signal level, active 3D measurement systems must instead adhere to a precise ratio between emitted light and ambient light while simultaneously avoiding saturation. On the other hand, manually finding a valid exposure set is a complex endeavor, as it requires the user to have a perfect mental model of the camera and its parameter set.
[0086] It is therefore desirable to provide a means for determining which exposure settings should be changed and by how much to optimize data quality for a given scene or subject. As discussed above with respect to FIG. 5B, noise improvement for any given pixel will decrease as contrast increases beyond a certain point. The embodiments described herein may use this fact to determine a suitable set of exposure settings given one or more imaging constraints (e.g., maximum total exposure time, maximum aperture size, etc.).
[0087] More specifically, targets can be specified in terms of values of quality parameters to be obtained for pixels in the final 3D image, where the targets shall be achieved subject to one or more imaging constraints. The objective is to try to optimize the final 3D image in terms of noise while imposing one or more constraints ("costs"), such as, for example, a maximum total exposure time, or a total illumination power.
[0088] We define an exposure cost E, which is a function of one or more exposure settings of the system. Cost We can start by defining exposure setting = exposure_setting_1, where exposure setting 1 is the exposure setting that defines the amount of light incident on the camera. ECost = f1 (exposure time) + f2 (aperture size) + f3 (neutral density filter strength) + f4 (illuminance) + ...
[0089] Here, functions {f1, f2, … f n} defines how the cost of exposure changes with changes in each respective parameter. These functions can be user-defined, effectively defining the "downside" to the user in changing each parameter. As an example, if it is desired to capture a 3D image in a very short spatial time, the user may allocate a high cost to exposure time. In another example, if the user is not time-constrained but wishes to minimize overall power usage, the user may allocate a high cost to illumination. E Cost The value of can be used to provide a constraint in determining the optimal exposure setting for a given acquisition.
[0090] E is the exposure value that indicates how much light reaches the sensor, both in terms of ambient light and light from the sensor system itself. value It is also possible to define E value = e1 (exposure time) e2 (aperture size) e3 (neutral density filter strength) e4 (illuminance) …
[0091] E value helps incorporate a number of influences that will affect how much signal is received by the camera and therefore the quality of each individual pixel in the system. value E value There are two exposure values: the exposure value for ambient light (E ambient ) and the exposure value (E amplitude ).
[0092] Function {e1, e2, … e n} may also be determined, where the relationship between each pair of functions defines the degree to which modifying one parameter will change the amount of incident light on the camera versus modifying the other parameter. As an example, doubling the exposure time may be equivalent to doubling the aperture size in terms of increasing the amount of incident light on the camera. In another example, doubling the aperture size may be equivalent to reducing the neutral density filter strength by a factor of four in terms of increasing the amount of incident light on the camera. The functions {e1, e2, ... e n} may be defined to take these relationships into account.
[0093] Function {e1, e2, … e n}, and the relationships between them can be determined empirically, off-line, by experimentation. Knowledge of the relationships between these functions is useful because they may allow one to translate changes in the value of one parameter into changes in the value of the other. For example, E value A specific value must be achieved for E value If it is possible to determine the change in exposure time that will achieve E, then it is possible to convert the change in exposure time into a change in the size of the aperture, which is then E value As will become apparent below, this is advantageous because it simplifies the determination of exposure settings to be collected by focusing only on one parameter (typically exposure time), and then translates changes in exposure time into values for other parameters according to the user's particular needs (e.g., wanting to minimize aperture size versus wanting to minimize overall exposure time, etc.).
[0094] In the following, the value E value , E Cost , and q min We provide two examples of how exposure settings can be determined based on min is the minimum acceptable value of the quality parameter.
[0095] In a first example, we set a target to find exposure settings for each image, where the sum of the exposure costs over a sequence of exposures is less than a predefined maximum cost.
number
number
number
number
[0096] As a second example, suppose a threshold number of pixels in the final 3D image satisfies a quality parameter value q(p)>q min
[0046] We set a target to find a set of exposure settings for each image that will minimize the sum of the exposure costs over the sequence of exposures while ensuring that we have
[0097] Below we will discuss different strategies for meeting these targets. In each case, the targets can be limited to use specific 2D or 3D regions of a given scene.
[0098] For most 3D measurement systems, there is a linear relationship between the received contrast / signal, c, and the noise, approximately
number
[0099] Referring again to the diagram of FIG. 5B, the distance uncertainty σ of the point being measured by the structured light system SL can be estimated using the following formula:
[0100]
number
[0101] where D is the distance to the point being measured, B is the camera-projector baseline, and θ c and θ p are the camera angle and projector angle, respectively. The value FOV is the field of view of the projector, and Φ maxis the number of projected sine waves, A is the amplitude of the observed sine waves, and C is the DC level of the received fixed light signal, i.e., ambient illumination + radiated sine waves (Bouquet, G. et al., "Design tool for TOF and SL based 3D cameras," Opt. Express 25, 27758-27769 (2017), the contents of which are incorporated herein by reference).
[0102] However, there is also the issue of sensor saturation. The maximum signal level that the sensor can handle without saturating is S. max If you define A+C>S, max If σ SL will degrade immediately. First, there will be a degradation in quality when only the sine wave portion can be restored, then further degradation when only the Gray code portion of the code can be decoded, and subsequently, when the sensor is fully saturated, information will be lost entirely. It will be appreciated that when capturing data using multiple exposure settings, it is not necessary to always capture both the Gray code and the phase image / sine wave. Since Gray code is more robust to saturation, one can capture the Gray code using one exposure setting and use that Gray code in combination with phase images captured using multiple exposure settings. This saves time as it saves the effort of having to recapture multiple Gray codes.
[0103] To take into account the possibility of saturation, the previous equation can be rewritten as:
[0104]
number
[0105] Note that in most cases, B and FOV can be considered to be constants.
[0106] As mentioned above, the exposure cost E for a particular image acquisition Costis modeled according to different parameters, and the amount of incident light on the camera affects each parameter. For simplicity, in the following we use E for the exposure length t. COST We will assume that each of these parameters is held constant, except for the exposure time, so that =t.
[0107] In that case, the following equation can be devised: A'(p)=tA * (p) C'(p)=tC * (p) where A'(p) denotes the measured amplitude of point p in the image, C'(p) denotes the measured ambient light of point p, and A * (p) is the amplitude-independent exposure time of point p, and C * (p) is the ambient light independent exposure time of point p.
[0108] A * and C * can be thought of as a normalized value of some unit of time. * and C * can be explained using the following: A * (p)=v + (p)-v - (p) C*(p)=v - (p) where for each pixel p, v + (p) is the pixel signal level with the system's active illumination on (projector for structured light on), and v - (p) is the pixel signal level with system illumination off.
[0109] The predicted noise of the system is then given by:
[0110]
number
[0111] In terms of pixels p, this yields:
[0112]
number
[0113] σ SL The expression for can be simplified to:
[0114]
number
[0115] To establish the value of the exposure time, we use the quality parameter q
number
[0116] Excess saturation (σ SL ≫σ max t defines the maximum time that a pixel can be exposed without experiencing saturation due to the effect of max In practice, this would require signal levels A and C that would saturate the sensor, so for some pixels, σ max >σ SLmeans that it becomes impossible to satisfy. This can occur, for example, when the sensor system is used outdoors under sunlight or in the presence of a strong light source. In such a case, it may not be possible for the integrated lighting to cancel out the ambient light.
[0117] t for each pixel min and t max The values of can be determined in several ways. In the first example, these values are predicted from the captured image as follows.
[0118] Using the exposure time t0, assume that the image I(t0) is captured by including the pixel p0 having v + (p0) and v - (p0). Then, based on the assumption that p0 is not completely over - saturated or under - saturated, the following equations are used to predict t min and t max for that pixel:
[0119]
Equation
[0120] At the same time,
[0121]
Equation
[0122] results in the following.
[0123]
Equation
[0124] If the captured pixel is over - saturated, it is not possible to make a prediction other than t max < t0. If the captured pixel is under - saturated (for example, the contrast is close to zero), t minIn the case of undersaturation, however, (σ SL (p0)>σ max At the same time, min and t max There will be a range of responses over which predictions can be made.
[0125] Interestingly, t min >t max There will be situations where the active illumination of the system is too weak in practice to overcome the ambient illumination and allow imaging with a sufficiently high quality, which may not be the case for the entire scene, for example the surface normal of the imaged object and its specular characteristics may be unfavourable due to the setup.
[0126] In the first example, different exposure settings i init ={E1,E2,…E n} a set of candidate images I init ={I1,I2,…I n}. In this case, we only consider the exposure time, so E init ={t1,t2,…t n}. Then the total time spent is T max Determine which subset of these images, and therefore the candidate exposure times, provides the best results in terms of maximizing the number of pixels in the final 3D image with an acceptable noise level while satisfying the constraint that E Cost is expressed solely in terms of exposure time, so in this example the value T max teeth,
number
[0127] In this example, I InitUsing the exposure times from the candidate exposures from, the following histogram-based approach can be used to quickly assess the amount of overexposed pixels: Init Given one or more images from, for each pixel, t min and t max Calculate t for most of the pixels. min and t max So that we can estimate the value of I Init We assume that covers the dynamic range of the entire scene (typically around 7 stops). For each pixel p, t min and t max Some images give a valid estimate for
number
number
number
[0128]
number
[0129] Also, T * is the maximum time bin under consideration.
[0130] Each value H(x,y) is min is in the xth row, and tmax Define x as the number of pixels in the yth column: H(xy)={p i |t min (p)∈[xΔ t ,(x+1)Δ t ] and t max (p)∈[yΔ t ,(y+1)Δ t ]
[0131] The histogram is a function of the time t for a given exposure time t ∈ [0,T * ], pixel N will be fully exposed using t by good This allows us to estimate the number of
[0132]
number
[0133] where idx(t) is the row / column number corresponding to t:
[0134]
number
[0135] This is I init Not only for exposure times of [0,T * Note that this can be applied to any exposure time within .
[0136] An example of such a histogram is shown in FIG. 10, where the exposure time is Δ t The time is given in milliseconds, where t = 5 ms. min ∈(0,500), and the y-axis represents t max ∈(0,500). The region labeled R1 represents the N good That is, the region R1 contains the values of H that are to be summed to obtain t min (p)∈(0,t') and t max (p)∈(t',T* ), for all pixels p such that
[0137] Given a set of exposure times E'={t0,t1,…} in ascending order, we can deduce:
[0138]
number
[0139] This is shown in Figure 11. Because rectangles R1 and R2 overlap, the above formula avoids counting a histogram bin twice.
[0140] To further accelerate this process, we use the cumulative histogram H c Can be used:
[0141]
number
[0142] Then, N for the complete set of exposures is calculated using the following formula: good can be readily calculated (i.e., to calculate the cumulative sum of pixels having an acceptable noise level in at least one of the captured images):
[0143]
number
[0144] This method uses only the sum of four numbers to find N good This allows (t) to be calculated immediately.
[0145] The sum of the exposure times over the images is T max The first target is met, i.e., the quality parameter value q(p)>q in the final image. minto find a set of exposure times that maximizes the number of pixels with good The above techniques may be used to determine
[0146] In one example, a "greedy" algorithm can be used as follows:
[0147] E opt Set =Φ. E opt The total time is T max At the same time less than: Exposure E opt For the set ∪E, q(p)>q min E∈E that maximizes the number of pixels with init This can be readily found by using the histogram technique discussed above; E to E opt In addition, E init Excluding E.
[0148] Figure 12 shows the results after each round in the algorithm, where each row represents one cycle in the while loop. After each step, we estimate G, where G is the number of iterations of E opt G is a binary image that is true for all pixels whose exposure is higher than the exposure contained within G. The initial image for G (Image 1) is completely black because there are no pixels known to have an acceptable value of q. Then, the maximum number of valid pixels (q>q min ) is selected (see image 2). This gives an updated G (image 3). In the next round, the updated image G is used as a basis and a new E is selected as the E that provides the maximum amount of well-exposed additional pixels. This then results in a further updated image for G.
[0149] If the cost is defined as a function of multiple imaging parameters (not just exposure time), then the sum of the costs for a selected set of exposures is T maxnot beyond, [Number] until it exceeds, this algorithm will be executed. More specifically, determine a candidate group of exposure settings and determine the value E for that group of exposure settings value and use this to convert / generate one or more alternative candidate groups of exposure settings where the value of E value is the same. This can be achieved, as discussed previously, by considering the function {e1, e2, … e n}. As an example, a candidate group of exposure settings where the exposure time is t1 and the aperture size is a1 can be identified. This can then be converted into a second group of exposure settings, where in this case the exposure time is t2 and the aperture size is a2, where t2 > t1 and a2 < a1. These two alternatives will, in that case, provide similar results in terms of the light collected from bright and dark regions of the subject being imaged. However, these two alternatives can have very different costs depending on whether the user places more emphasis on keeping the total exposure time as short as possible or on keeping the aperture size as small as possible (in other words, depending on the functions f1 and f2 of E cost ).
[0150] The greedy algorithm can be easily modified to find a set of exposure settings for each image that satisfies the second target while ensuring that the number of pixels at the threshold value in the final 3D image has a quality parameter value q(p) > q min and to minimize the total exposure cost over the sequence of exposures.
[0151] Set E opt = Φ. E opt where the total number of pixels for q(p) > q min is less than N min and at the same time: For the set of exposures E opt ∪E where q(p) > qmin maximize the number of pixels with init This can be readily found by using the histogram technique as discussed above; E to E opt In addition, E init Excluding E.
[0152] When selecting the optimal set of exposure times, the initial sequence of captured images I init It will be appreciated that the present invention need not be limited to selecting exposure times that were within the range of exposure times t i ∈E init This allows us to utilize the exposure times corresponding to each respective time bin in the histogram rather than just H(x,y). Given H(x,y), we can readily compute a set of optimal exposure times (where the sum of the exposure times over the image is T max The quality parameter value q(p)>q in the final image, provided that min (This leads us back to the first target, which is to maximize the number of pixels with
[0153]
number
number
number
number
number
number
number
number
[0154] To determine the optimal set of exposure times, proceed as follows: - T max have a total time of less than
number
number
number
number
number
[0155] When using a cost-based approach, the basis for the algorithm is
number
number
number
[0156] For each integer k, this partition is
number
number
number
[0157] A bottom-up approach using dynamic programming is simple, and one can obtain precomputed tables of all such partitionings up to some large value of k.
[0158] - Each subset
number
number
number
[0159] The previous example is discussed in the context of capturing an initial sequence of images and then determining the best exposure settings for image acquisition. Thus, in the previous example, the candidate group of exposure settings is determined offline, such that each group of exposure settings to be used in the 3D image acquisition process is determined prior to acquiring each set of image data. In the following, an online approach for selecting optimal exposure settings is described. Here, it is assumed that an initial exposure is given and the objective is to predict the next best exposure. Thus, in this online approach, the 3D image acquisition process is performed iteratively by determining, at each iteration, the next best group of exposure settings to use and then capturing a set of image data with that group of exposure settings. At each iteration of the method, one or more new candidate groups of exposure settings are considered, and one of the candidate groups of exposure settings is selected for use to acquire the next set of image data.
[0160] Consider an image I(t) taken with exposure time t. For each pixel p i For I(t), let v be the pixel signal level when the active lighting of the system is on (projector on for structured light). + (p i ) and v denotes the pixel signal level with the system's active illumination off. - (p i ) is indicated.
[0161] We have some constraints (below, the numbers given are based on the camera sensor being an 8-bit sensor with 256 intensity levels): -v bad , the maximum value for "black" underexposed pixels (e.g., 10). If you increase the exposure time, v <v bad cannot guarantee an increase in value. This would include pixels in shadow areas, for example. -v min, the minimum acceptable pixel value (e.g., 50). If the exposure time is increased, v ∈ [v bad ,v min ] increases its value. -v max , the maximum allowed pixel value (for example, 230). max are overexposed pixels and it is impossible to estimate an appropriate exposure time for those pixels. - σ max , the maximum measurement uncertainty that gives an acceptable 3D quality. For each pixel, we estimate σ as follows:
[0162]
number
[0163] Given p with σ(p) and time t, we can estimate σ'(p) using time t'=αt as follows:
[0164]
number
[0165] The algorithm proceeds as follows: - Start with the initial exposure time and collect images - Next, calculate a candidate set of possible exposure times - Calculate the expected number of overexposed pixels for each candidate exposure time - Choose the exposure time that has the maximum expected number of exposed pixels - Repeat until a condition is met (e.g. the maximum allowed time is reached)
[0166] A probabilistic model can be used to estimate the number of highly exposed pixels given a set of exposures that have already occurred. Consider an image I(t) taken with exposure time t. For each pixel p i ∈I(t) can be classified into one of the following categories:
[0167] - Case 1 (impossible): The pixel is properly exposed, but σ(p) is too high, i.e. v - (p)∈(v min ,v max ), v + (p)∈(v min ,v max ), σ(p)>σ max
[0168] Since it is not possible to obtain adequate measurements from pixels in this category by increasing the exposure time, these pixels are removed from consideration.
[0169] - Case 2 (Acceptable Quality): The pixel is properly exposed and the 3D measurements at this point have acceptable quality, i.e. v + (p)∈(v min ,v max ), σ(p)<σ max
[0170] These pixels may be eliminated from further consideration.
[0171] - Case 3 (Overexposure). Here, there are actually two possibilities: (a) The pixel is overexposed even with the projector off, i.e., v - (p)>v max In this case, there is nothing that can be done other than reducing the ambient light. For each exposure stop
number
number
[0172]
number
[0173] (b) The pixel is overexposed in the projector-on state, but properly exposed in the projector-off state, i.e., v + (p)>v max and v - (p) <v min In this case the aim should be to reduce the exposure time. As before: Pr(σ(p)<σ max |t>t i )=0
[0174] If p is properly exposed, then p will have some probability α 2b By |v + (p)-v - Suppose we say that we have (p)|=R. This means that with probability α 2b . + (p)=v max This means that.
[0175]
number
[0176] This can be written as follows:
[0177]
number
[0178] Then, σ(p)<σ max It is possible to estimate the range in which t' can be reduced while still maintaining t"=β"t', then
number
[0179] this is, Pr(σ(p)<σ max|t∈(β′′t i ,β't i ))=α 2b , Pr(σ(p)<σ max |t<β''t i )=0 Gives.
[0180] - Case 4 (Underexposure). Here, there are three possibilities: (a) A pixel is underexposed with the projector on, i.e., v + (p) <v bad Now we need to increase the exposure time. Again, for each stop we increase the α 3α increases the probability, but α' 3α is less than.
[0181]
number
[0182] (b) The pixel is underexposed with the projector on, i.e., σ(p)>σ max However, v + (p)∈(v bad ,v min ) σ(p)<σ max However, v + (p) <v max To maintain the same, v + By increasing the exposure time to obtain (p), we can estimate the probability as follows:
[0183]
number
[0184] (c) The pixel is underexposed with the projector off, i.e. v - (p) <v min , v + (p)∈(v min ,v max ), σ(p)>σ maxThis case is similar to case 4(b) above.
[0185] FIG. 13 presents a general framework for the above-mentioned online approach with an example of a possible algorithm for choosing an estimated expected number of pixels with tolerances for candidate time and quality parameters. Let E={(t i ,I(t i )}. The algorithm uses ComputeExpectedNum to check whether σ(p)<σ max The algorithm uses the probability model described above and uses machine learning techniques to estimate the expected number of pixels for which α i This procedure can be extended by estimating the probability for each pixel based on which of the above categories it falls into (e.g., by using the values of neighboring pixels using a convolutional neural network CNN). - I(t i ) for any t i Based on the premise -t i Candidate times in a logarithmic grid with 1 / 6 stops between <t min ,>t max Calculate - For each candidate time, calculate the expected number of valid pixels (using above). - Choose the time with the highest expected number of valid pixels.
[0186] As before, if the cost is defined as a function of multiple imaging parameters (not just exposure time), the algorithm can be max The sum of the costs for a set of exposures is
number
[0187] It will be appreciated that in the above example, the change in exposure settings for each image acquisition is limited to exposure time only, with the understanding that other parameters (aperture size, illuminance, neutral density filter strength, etc.) remain constant over the sequence of image acquisitions. However, as previously discussed, it is possible to infer from the change in exposure time the extent to which other parameters need to be changed to achieve the same signal to noise. This can be achieved by using the respective functions {e1, e2, ..., e n} for a particular acquisition. n}, which is calculated by multiplying their respective costs {f1,f2,…,f n} while minimizing each function {e1, e2, …, e n}, the aperture sizes {a1, a2, …, a n} or a combination of both of these parameters {{t1,a1},{t2,a2}…,{t n ,a n}}. Thus, although the algorithms described herein focus on exposure time, determining different exposure times can act as a proxy for determining settings for other parameters that affect the amount of incident light on the camera.
[0188] For adjusted projector brightness, this will primarily affect the amplitude of the received signal, with much less effect on the ambient light intensity. This can be easily incorporated into the "greedy" algorithm described above by simply including images captured at different projector brightnesses in a set, e.g., a discrete set of different brightnesses.
[0189] It will be further appreciated that although the particular examples described herein relate to structured illumination systems, the methods described herein may be readily extended to other forms of 3D imaging by taking into account how the signal to noise in the final image varies as a function of the received optical signal. As an example, for an active time-of-flight system, the depth noise σ TOF The following relations exist for:
[0190]
number
[0191] In the formula, N ph is the total received signal level (sum of amplitude A and ambient light C), and τ response is the time response of the system, c is the speed of light, and m is the number of samples taken.
number
[0192] It will be appreciated that the above algorithms can be easily constrained to operate only on relevant regions within an image. This region can be specified in 2D (in pixel coordinates) or in 3D (in world XYZ coordinates). Pixels that are determined not to fall within the specified region of interest, either in 2D or 3D, can then be excluded from further consideration of the algorithm.
[0193] In summary, the embodiments described herein provide a means for rendering high SNR 3D images of a scene or subject. It is possible to accommodate large variations in the amount of light available from different points in a scene by performing multiple image acquisitions with different exposure settings and merging the data from those sets of image data to form a single point cloud in which the signal-to-noise ratio is maximized for each point. Additionally, the embodiments provide a means for determining a set of exposure settings to use in acquiring each image in a manner that will maximize the signal-to-noise ratio in the final 3D image while satisfying one or more constraints on time, depth of focus, illumination power, etc.
[0194] It will be appreciated that implementations of the subject matter and operations described herein may be realized in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or in any combination of one or more of them. Implementations of the subject matter described herein may be realized using one or more computer programs, i.e., one or more modules of computer program instructions, encoded on a computer storage medium for execution by or to control the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal generated to encode information for transmission to a receiver apparatus suitable for execution by a data processing apparatus. The computer storage medium may be or may be included in a computer readable storage device, a computer readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Furthermore, the computer storage medium is not a propagating signal, but the computer storage medium may be the source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium may also be, or be contained within, one or more separate physical components or media (eg, multiple CDs, disks, or other storage devices).
[0195] Although several embodiments have been described, these embodiments are presented by way of example only and are not intended to limit the scope of the present invention. In fact, the novel methods, devices, and systems described herein may be embodied in various forms. Furthermore, various omissions, substitutions, and changes may be made to the forms of the methods and systems described herein without departing from the spirit of the present invention. The appended claims and their equivalents are intended to cover such forms or modifications that fall within the scope and spirit of the present invention. [Explanation of symbols]
[0196] 101 2D point cloud matrix 201a Point cloud matrix 201b point cloud matrix 201c point cloud matrix 203 Point cloud matrix 301 point cloud matrix 501 Subject 503 Projector 505 Camera 507 Fringe 509 Fringe 701 yen 703a Lighting Pattern 703b Lighting Pattern 703c Lighting Pattern 705 First point cloud matrix 707 yen 709a Lighting Pattern 709b Lighting Pattern 709c Lighting Pattern 711 Second Point Cloud Matrix 713 yen 715a Lighting Pattern 715b Lighting Pattern 715c Lighting Pattern 717 Third input 3D image point cloud matrix 719 Output point cloud matrix
Claims
1. 1. A method for determining one or more groups of exposure settings for use in a 3D image acquisition process performed with an imaging system, the imaging system comprising an image sensor, the 3D image acquisition process comprising capturing one or more sets of image data with the image sensor using respective groups of exposure settings, the one or more sets of image data enabling the generation of one or more 3D point clouds defining three dimensional coordinates of points on a surface of one or more objects being imaged, each group of exposure settings including an exposure time used when capturing a respective one of the sets of image data, the method comprising: (i) identifying one or more candidate groups of exposure settings using image data captured by the image sensor; (ii) for each candidate group of exposure settings, determining an amount of signal that may be received at different pixels of the image sensor when the candidate group of exposure settings is used to capture a set of image data for use in the 3D image acquisition process; determining whether each pixel would be an exposed pixel when using the candidate group of exposure settings based on the amount of signal likely to be received at the different pixels, an exposed pixel being a pixel for which a value of a quality parameter associated with the pixel exceeds a threshold, the value of the quality parameter for the pixel reflecting a degree of uncertainty that would exist in the three-dimensional coordinates of a point in the point cloud associated with the pixel if a point cloud were generated using the set of image data captured with the candidate group of exposure settings; (iii) selecting one or more groups of exposure settings to be used for the 3D image acquisition process from the one or more candidate groups of exposure settings, said selection being made so as to maximize the number of pixels belonging to a set N, such that a pixel belongs to set N if there is at least one selected group of exposure settings for which the pixel is determined to be a highly exposed pixel; Including, A method wherein the selection is performed under the constraint that the sum of the exposure times of the selected group of exposure settings is less than a predefined threshold.
2. 2. The method of claim 1, wherein for one or more of the candidate groups of exposure settings, determining an amount of signal that may be received at different pixels of the image sensor when the candidate group of exposure settings is used to capture a set of image data comprises capturing the set of image data using the candidate group of exposure settings.
3. The method of claim 2 , wherein the set of image data captured when using the one or more candidate groups of exposure settings is used to identify one or more other candidate groups of exposure settings.
4. Steps (i) through (iii) are repeated through one or more iterations, and for each iteration: One of the candidate groups of exposure settings identified in the iterations is selected; 2. The method of claim 1, wherein the number of pixels in the set N is maximized such that the sum of exposure times for the group of exposure settings selected in the current iteration and the groups of exposure settings selected in all previous iterations is less than the predefined threshold.
5. for each iteration, the selected group of exposure settings is used to capture a set of imaging data with the imaging system; The method of claim 4 , wherein for each iteration after the second iteration, a set of image data captured in a previous iteration is used in determining the candidate group of exposure settings for the current iteration.
6. 6. The method of claim 5, wherein the step of determining whether each pixel will be an exposed pixel when using the candidate group of exposure settings comprises the step of determining a probability that the each pixel will be exposed, the probability being determined based on an amount of signal received at the each pixel in a previous iteration of the method.
7. and / or wherein identifying one or more candidate groups of exposure settings comprises determining, for one or more pixels of the image sensor, a range of exposure times over which the pixel is likely to be an overexposed pixel; The method of claim 1 , wherein the value of the quality parameter associated with a pixel is determined based on an amount of ambient light in a scene being imaged.
8. Each group of exposure settings is the size of an aperture stop in the path between the object and the sensor; the intensity of the light used to illuminate the object; and The strength of an ND filter placed in the optical path between the subject and the sensor 8. The method of claim 1 , comprising one or more of the following:
9. the imaging system is an optical imaging system comprising one or more optical sensors; 9. The method of any one of claims 1 to 8, optionally wherein the imaging system includes one or more light sources used to illuminate the object being imaged.
10. 10. The method of claim 1, wherein the image data in each set of image data comprises one or more 2D images of the object captured by the sensor, and / or each set of image data comprises color information.
11. 11. The method of claim 1, wherein the imaging system is an imaging system that uses structured illumination to acquire each set of image data.
12. each set of image data includes a sequence of 2D images of the object captured by a photosensor, each 2D image in the sequence being captured using a different illumination pattern; Optionally, each set of image data comprises a sequence of gray coded images and a sequence of phase shifted images.
13. 1. A method for generating a 3D image of one or more objects using an imaging system having an image sensor, comprising: capturing one or more sets of image data with the image sensor using respective groups of exposure settings, the sets of image data enabling generation of one or more 3D point clouds defining three-dimensional coordinates of points on a surface of the one or more objects, each group of exposure settings including an exposure time used when capturing a respective one of the sets of image data; constructing a 3D point cloud using data from one or more of said captured sets of image data; Including, A method, wherein the exposure settings used to capture each set of image data are determined using a method according to any one of claims 1 to 12.
14. A computer readable storage medium having stored thereon computer executable code which, when executed by a computer, causes the computer to perform the method of any one of claims 1 to 13.
15. 14. An imaging system for performing a 3D image collection process by capturing one or more sets of image data using one or more groups of exposure settings, the one or more sets of image data such that one or more 3D point clouds defining three-dimensional coordinates of points on a surface of one or more objects being imaged are generated, the imaging system comprising an image sensor for capturing the one or more sets of image data, and the imaging system is configured to determine the one or more groups of exposure settings to use for the 3D image collection process by performing a method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Method and system for determining optimal exposure of structured light based 3D camera
US20090185800A1
Method and system for determining optimal exposure time and number of exposures in structured light-based 3D camera
US20170118456A1
Structured-light-based exposure control method and exposure control apparatus
US20180176440A1